ARCSField OS
Check it before you say it.
Five steps between an idea and a message. Each step gets a stamp: go or stop. If any step says stop, you fix it before anyone hears it.
It checks that every fact you mark names a source, that no new marked facts sneak in along the way, and that at least one review disagrees. It can't tell you whether a claim is true. That part is still yours.
- 1Find the factsWhen you need to know what's true before building on it.
analysis-by-arcs3 checks, all visibleGo - 2Check the factsWhen you need to tell a solid fact from a shaky one.
assessment-by-arcs3 checks, all visibleGo - 3Poke holesWhen you need to find the weak spot before someone else does.
evaluation-by-arcs4 checks, all visibleGo - 4Make it shortWhen you need it short without adding anything unchecked.
synthesis-by-arcs5 checks, all visibleStop - 5Write the noteWhen you need a note that's ready for a person to send.
recapitulation-by-arcs4 checks, all visibleReady for a person
Example run: steps 1 to 3 said go. Step 4 said stop, because the short version added "trusted by 50+ teams" and nobody had checked it.
The problem it solves
What we say tends to run ahead of what we checked. The line that slips through is usually the one that sounds best, like "trusted by 50+ teams" in the run above.
Use it when
An update has numbers in it
A board, investor or team update. Mark each figure with where it came from; step 1 stops on any that names no source.
A claim about a customer
A case study or a launch post. Mark a big claim as high stakes and it needs two sources from different sites.
You summarize an analysis
Load a Jupyter notebook. Each sentence in its notes that has a number becomes a fact, matched against the code output that printed that number.
An AI drafted it for you
Mark the facts you would stand behind before you forward it. Any without a source stops at step 1.
Your docs live in a repo
Run it before each commit or in a GitHub check, so an unsourced marked fact fails in the pull request, not in front of a reader.
Marking a fact means wrapping the sentence in a tag that names its source: <span class="claim" data-src="https://...">The fact.</span>. Anything you don't mark isn't checked.
What you can rely on it for, and what you can't
Rely on it for
- The same answer every time. The same input always gets the same stamp.
- Checks you can read. Every rule is listed below, with its code.
- Privacy. It runs in this page. Nothing you paste is uploaded.
- A person sends. The last step can only say "ready."
Not for
- Knowing what's true. It never looks; it checks that you did the steps that would let you find out.
- Facts you didn't mark. Plain prose gets no check.
- Judgment calls. Each step has a scorecard it leaves blank for you.
- Objectivity. The rules are fixed, but a person chose them, and the scorecard floors aren't calibrated yet.
Three questions
- Is this the AI's chain of thought?
- No. It checks the finished work, not the thinking. Each step saves a file you can open later.
- What do I need to run it?
- This page, or Python. The checks run here in your browser. For a terminal, code, GitHub or a pre-commit hook, clone the open-source repo: Python 3.9+, standard library only, nothing to install.
- Who decides the parts a machine can't?
- A person. Each step has a short scorecard. The tool leaves it blank for you.
Try it on your own work
Pick a step, paste your work, press Check it. It runs in this page. Nothing you paste is uploaded.
Press Check it to see the stamp. Start with "Load a broken one" to watch a check catch it.
Inside each step
Pick a step to see its checks and copy its code. Step 4 is showing because that's where the example stopped.
Machine checks
Any one of these fails, the step says stop.
- Every marked fact names a source41 of 41passed
- Facts marked big name two sources6 big claims, all with 2+passed
- Every source is on the evidence list0 missingpassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Did we look in the obvious places?
- Did we use the original, not a summary?
- Did we keep the question open?
- Did we log the weird stuff?
Copy it
tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl --out prints/analysis.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate, load_ledger
upstream = None
report = evaluate("analysis", open("teardown.html").read(), upstream=upstream, ledger=load_ledger("evidence.jsonl"))
json.dump(report, open("prints/analysis.json", "w"), indent=2)
print(report["fate"]) # sourced
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Analysis by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl --out prints/analysis.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: analysis-print
path: prints/analysis.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: analysis-by-arcs
name: Analysis by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl
language: system
files: '\.html$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- Every fact has a verdict and a size41 of 41passed
- A fact marked true was found in step 136 of 36passed
- The totals match the list36 true, 2 gaps, 3 unsurepassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Is each problem sized right?
- Are the gaps said as plainly as the wins?
- No hedging to sound safe?
- Are the objections written down?
trueSetup takes under 10 minutes on the demo.
gapThe deck says 40 customers. The case page lists 12.
unsureNo security report (SOC 2) was found.
Shaky facts look shaky. Only true facts render sharp.
Copy it
tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json --out prints/assessment.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/analysis.json"))
report = evaluate("assessment", open("report.html").read(), upstream=upstream)
json.dump(report, open("prints/assessment.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Assessment by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json --out prints/assessment.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: assessment-print
path: prints/assessment.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: assessment-by-arcs
name: Assessment by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json
language: system
files: '\.html$'
pass_filenames: falseToo-agreeable check: fine. 9 agreements, 5 disagreements.
If reviewers only agree, this step stops. A test nothing could fail is not a test.
Machine checks
Any one of these fails, the step says stop.
- At least 4 reviewers ran6 ranpassed
- Each reviewer gave a reason6 of 6passed
- At least one review disagrees2 disagreed: the risk reviewer and the skepticpassed
- The result follows from the reviewspass, with 2 fixes requiredpassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Was the best objection a real one?
- Did we argue the product works?
- Did we argue we can prove it?
- Are the required fixes concrete?
Copy it
tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4 --out prints/evaluation.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/assessment.json"))
report = evaluate("evaluation", open("lenses.json").read(), upstream=upstream, min_lenses=4)
json.dump(report, open("prints/evaluation.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Evaluation by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4 --out prints/evaluation.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: evaluation-print
path: prints/evaluation.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: evaluation-by-arcs
name: Evaluation by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4
language: system
files: 'lenses.*\.json$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- No new marked facts snuck in"Trusted by 50+ teams" was never checkedfailed
- Required fixes were madethe 40-customer number still looks certainfailed
- Writing has no fillercleanpassed
- No long dashes0passed
- Hidden from search until readyyespassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Plain words?
- Clear structure?
- Easy rhythm?
- Specific, not vague?
- About five times shorter?
Copy it
tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json --out prints/synthesis.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/evaluation.json"))
report = evaluate("synthesis", open("surface.html").read(), upstream=upstream)
json.dump(report, open("prints/synthesis.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Synthesis by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json --out prints/synthesis.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: synthesis-print
path: prints/synthesis.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: synthesis-by-arcs
name: Synthesis by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json
language: system
files: '\.html$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- Still marked as a draftDRAFTpassed
- Our prediction was written down firstfiled before the draftpassed
- The note says nothing the short version didn't5 of 5passed
- Not too agreeableborderline: 13 agreements, 2 disagreementspassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Does it sound right for this person?
- Honest about what we dropped and kept?
- Is the ask clear?
- Did we note to follow up?
Copy it
tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json --out prints/recapitulation.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/synthesis.json"))
report = evaluate("recapitulation", open("note.md").read(), upstream=upstream)
json.dump(report, open("prints/recapitulation.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Recapitulation by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json --out prints/recapitulation.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: recapitulation-print
path: prints/recapitulation.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: recapitulation-by-arcs
name: Recapitulation by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json
language: system
files: 'note.*\.md$'
pass_filenames: falseFor builders
The code is MIT-licensed at github.com/flashesofbrilliance/check-it-before-you-say-it. The snippets on this page assume it sits at tools/arcs-verb-eval/ in your repo.
All five steps as one GitHub workflow
# .github/workflows/check-before-you-say-it.yml
name: check-before-you-say-it
on: [pull_request]
jobs:
five-steps:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: git clone --depth 1 https://github.com/flashesofbrilliance/check-it-before-you-say-it tools/arcs-verb-eval
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Self-test (the checks check themselves)
run: tools/arcs-verb-eval/arcs-verb-eval --self-test
- name: All five steps
run: tools/arcs-verb-eval/arcs-verb-eval pipeline --config arc.json
- name: Report
if: always()
run: |
tools/arcs-verb-eval/arcs-verb-eval export prints --format md >> "$GITHUB_STEP_SUMMARY"
tools/arcs-verb-eval/arcs-verb-eval export prints --format junit --out prints/junit.xml
- uses: actions/upload-artifact@v4
if: always()
with:
name: five-step-prints
path: prints/What each step saves
One format for all five. The next step reads it.
{
"verb": "synthesis",
"fate": "hold",
"upstream": "evaluation.json#sha256:9f1c",
"gates": [
{
"id": "no-new-claims",
"ok": false,
"evidence": "line 212"
}
],
"rubric": [
{
"row": "D-specificity",
"score": 2,
"judge": "human"
}
],
"sycophancy": {
"agreement": 4,
"dissent": 3,
"verdict": "OK",
"blocking": false
}
}