Field OS

Check it before you say it.

Five steps between an idea and a message. Each step gets a stamp: go or stop. If any step says stop, you fix it before anyone hears it.

It checks that every fact you mark names a source, that no new marked facts sneak in along the way, and that at least one review disagrees. It can't tell you whether a claim is true. That part is still yours.

  1. 1Find the factsWhen you need to know what's true before building on it.analysis-by-arcs3 checks, all visibleGo
  2. 2Check the factsWhen you need to tell a solid fact from a shaky one.assessment-by-arcs3 checks, all visibleGo
  3. 3Poke holesWhen you need to find the weak spot before someone else does.evaluation-by-arcs4 checks, all visibleGo
  4. 4Make it shortWhen you need it short without adding anything unchecked.synthesis-by-arcs5 checks, all visibleStop
  5. 5Write the noteWhen you need a note that's ready for a person to send.recapitulation-by-arcs4 checks, all visibleReady for a person

Example run: steps 1 to 3 said go. Step 4 said stop, because the short version added "trusted by 50+ teams" and nobody had checked it.

The problem it solves

What we say tends to run ahead of what we checked. The line that slips through is usually the one that sounds best, like "trusted by 50+ teams" in the run above.

Use it when

An update has numbers in it

A board, investor or team update. Mark each figure with where it came from; step 1 stops on any that names no source.

A claim about a customer

A case study or a launch post. Mark a big claim as high stakes and it needs two sources from different sites.

You summarize an analysis

Load a Jupyter notebook. Each sentence in its notes that has a number becomes a fact, matched against the code output that printed that number.

An AI drafted it for you

Mark the facts you would stand behind before you forward it. Any without a source stops at step 1.

Your docs live in a repo

Run it before each commit or in a GitHub check, so an unsourced marked fact fails in the pull request, not in front of a reader.

Marking a fact means wrapping the sentence in a tag that names its source: <span class="claim" data-src="https://...">The fact.</span>. Anything you don't mark isn't checked.

What you can rely on it for, and what you can't

Rely on it for

  • The same answer every time. The same input always gets the same stamp.
  • Checks you can read. Every rule is listed below, with its code.
  • Privacy. It runs in this page. Nothing you paste is uploaded.
  • A person sends. The last step can only say "ready."

Not for

  • Knowing what's true. It never looks; it checks that you did the steps that would let you find out.
  • Facts you didn't mark. Plain prose gets no check.
  • Judgment calls. Each step has a scorecard it leaves blank for you.
  • Objectivity. The rules are fixed, but a person chose them, and the scorecard floors aren't calibrated yet.

Three questions

Is this the AI's chain of thought?
No. It checks the finished work, not the thinking. Each step saves a file you can open later.
What do I need to run it?
This page, or Python. The checks run here in your browser. For a terminal, code, GitHub or a pre-commit hook, clone the open-source repo: Python 3.9+, standard library only, nothing to install.
Who decides the parts a machine can't?
A person. Each step has a short scorecard. The tool leaves it blank for you.

Try it on your own work

Pick a step, paste your work, press Check it. It runs in this page. Nothing you paste is uploaded.

Press Check it to see the stamp. Start with "Load a broken one" to watch a check catch it.

Inside each step

Pick a step to see its checks and copy its code. Step 4 is showing because that's where the example stopped.

4Make it shortSynthesis by ARCSStop

Machine checks

Any one of these fails, the step says stop.

  • No new marked facts snuck in"Trusted by 50+ teams" was never checkedfailed
  • Required fixes were madethe 40-customer number still looks certainfailed
  • Writing has no fillercleanpassed
  • No long dashes0passed
  • Hidden from search until readyyespassed

Person checks

Scored 0 to 2 by a person. The tool never scores these itself.

  • Plain words?
  • Clear structure?
  • Easy rhythm?
  • Specific, not vague?
  • About five times shorter?

Copy it

tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json --out prints/synthesis.json

For builders

The code is MIT-licensed at github.com/flashesofbrilliance/check-it-before-you-say-it. The snippets on this page assume it sits at tools/arcs-verb-eval/ in your repo.

All five steps as one GitHub workflow
# .github/workflows/check-before-you-say-it.yml
name: check-before-you-say-it
on: [pull_request]
jobs:
  five-steps:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: git clone --depth 1 https://github.com/flashesofbrilliance/check-it-before-you-say-it tools/arcs-verb-eval
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - name: Self-test (the checks check themselves)
        run: tools/arcs-verb-eval/arcs-verb-eval --self-test
      - name: All five steps
        run: tools/arcs-verb-eval/arcs-verb-eval pipeline --config arc.json
      - name: Report
        if: always()
        run: |
          tools/arcs-verb-eval/arcs-verb-eval export prints --format md >> "$GITHUB_STEP_SUMMARY"
          tools/arcs-verb-eval/arcs-verb-eval export prints --format junit --out prints/junit.xml
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: five-step-prints
          path: prints/
What each step saves

One format for all five. The next step reads it.

{
  "verb": "synthesis",
  "fate": "hold",
  "upstream": "evaluation.json#sha256:9f1c",
  "gates": [
    {
      "id": "no-new-claims",
      "ok": false,
      "evidence": "line 212"
    }
  ],
  "rubric": [
    {
      "row": "D-specificity",
      "score": 2,
      "judge": "human"
    }
  ],
  "sycophancy": {
    "agreement": 4,
    "dissent": 3,
    "verdict": "OK",
    "blocking": false
  }
}