ARCSField OS
Check it before you say it.
Five steps between an idea and a message. Each step gets a stamp: go or stop. If any step says stop, you fix it before anyone hears it.
It checks that every fact says where it came from, that nothing new sneaks in along the way, and that someone really pushed back. It can't tell you whether a claim is true. That part is still yours.
- 1Find the factsWhen you need to know what's true before building on it.
analysis-by-arcs3 checks, all visibleGo - 2Check the factsWhen you need to tell a solid fact from a shaky one.
assessment-by-arcs3 checks, all visibleGo - 3Poke holesWhen you need to find the weak spot before someone else does.
evaluation-by-arcs4 checks, all visibleGo - 4Make it shortWhen you need it short without adding anything unchecked.
synthesis-by-arcs5 checks, all visibleStop - 5Write the noteWhen you need a note that's ready for a person to send.
recapitulation-by-arcs4 checks, all visibleReady for a person
Example run: steps 1 to 3 said go. Step 4 said stop, because the short version added "trusted by 50+ teams" and nobody had checked it.
Why bother
Nothing sneaks in
Each step can only use what the step before it approved. A fact that skipped a check gets caught.
No yes-men
If every reviewer just agrees, the check fails. Someone has to have tried to break it.
A person hits send
The last step can only say "ready." It never sends anything for you.
Three questions
- Is this the AI's chain of thought?
- No. It checks the finished work, not the thinking. Each step saves a file you can open later.
- What do I need to run it?
- Just Python. Nothing to install. It runs in a terminal, in code, in GitHub, or before each commit.
- Who decides the parts a machine can't?
- A person. Each step has a short scorecard. The tool leaves it blank for you.
Try it on your own work
Pick a step, paste your work, press Check it. It runs in this page. Nothing you paste is uploaded.
Press Check it to see the stamp. Start with "Load a broken one" to watch a check catch it.
Inside each step
Pick a step to see its checks and copy its code. Step 4 is showing because that's where the example stopped.
Machine checks
Any one of these fails, the step says stop.
- Every marked fact names a source41 of 41passed
- Facts marked big name two sources6 big claims, all with 2+passed
- Every source is on the evidence list0 missingpassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Did we look in the obvious places?
- Did we use the original, not a summary?
- Did we keep the question open?
- Did we log the weird stuff?
Copy it
tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl --out prints/analysis.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate, load_ledger
upstream = None
report = evaluate("analysis", open("teardown.html").read(), upstream=upstream, ledger=load_ledger("evidence.jsonl"))
json.dump(report, open("prints/analysis.json", "w"), indent=2)
print(report["fate"]) # sourced
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Analysis by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl --out prints/analysis.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: analysis-print
path: prints/analysis.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: analysis-by-arcs
name: Analysis by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval analysis teardown.html --ledger evidence.jsonl
language: system
files: '\.html$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- Every fact has a verdict and a size41 of 41passed
- A fact marked true was found in step 136 of 36passed
- The totals match the list36 true, 2 gaps, 3 unsurepassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Is each problem sized right?
- Are the gaps said as plainly as the wins?
- No hedging to sound safe?
- Are the objections written down?
trueSetup takes under 10 minutes on the demo.
gapThe deck says 40 customers. The case page lists 12.
unsureNo security report (SOC 2) was found.
Shaky facts look shaky. Only true facts render sharp.
Copy it
tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json --out prints/assessment.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/analysis.json"))
report = evaluate("assessment", open("report.html").read(), upstream=upstream)
json.dump(report, open("prints/assessment.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Assessment by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json --out prints/assessment.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: assessment-print
path: prints/assessment.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: assessment-by-arcs
name: Assessment by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval assessment report.html --upstream prints/analysis.json
language: system
files: '\.html$'
pass_filenames: falseToo-agreeable check: fine. 9 agreements, 5 disagreements.
If reviewers only agree, this step stops. A test nothing could fail is not a test.
Machine checks
Any one of these fails, the step says stop.
- At least 4 reviewers ran6 ranpassed
- Each reviewer gave a reason6 of 6passed
- At least one review disagrees2 disagreed: the risk reviewer and the skepticpassed
- The result follows from the reviewspass, with 2 fixes requiredpassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Was the best objection a real one?
- Did we argue the product works?
- Did we argue we can prove it?
- Are the required fixes concrete?
Copy it
tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4 --out prints/evaluation.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/assessment.json"))
report = evaluate("evaluation", open("lenses.json").read(), upstream=upstream, min_lenses=4)
json.dump(report, open("prints/evaluation.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Evaluation by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4 --out prints/evaluation.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: evaluation-print
path: prints/evaluation.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: evaluation-by-arcs
name: Evaluation by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval evaluation lenses.json --upstream prints/assessment.json --set min_lenses=4
language: system
files: 'lenses.*\.json$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- No new marked facts snuck in"Trusted by 50+ teams" was never checkedfailed
- Required fixes were madethe 40-customer number still looks certainfailed
- Writing has no fillercleanpassed
- No long dashes0passed
- Hidden from search until readyyespassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Plain words?
- Clear structure?
- Easy rhythm?
- Specific, not vague?
- About five times shorter?
Copy it
tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json --out prints/synthesis.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/evaluation.json"))
report = evaluate("synthesis", open("surface.html").read(), upstream=upstream)
json.dump(report, open("prints/synthesis.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Synthesis by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json --out prints/synthesis.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: synthesis-print
path: prints/synthesis.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: synthesis-by-arcs
name: Synthesis by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval synthesis surface.html --upstream prints/evaluation.json
language: system
files: '\.html$'
pass_filenames: falseMachine checks
Any one of these fails, the step says stop.
- Still marked as a draftDRAFTpassed
- Our prediction was written down firstfiled before the draftpassed
- The note says nothing the short version didn't5 of 5passed
- Not too agreeableborderline: 13 agreements, 2 disagreementspassed
Person checks
Scored 0 to 2 by a person. The tool never scores these itself.
- Does it sound right for this person?
- Honest about what we dropped and kept?
- Is the ask clear?
- Did we note to follow up?
Copy it
tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json --out prints/recapitulation.jsonimport json, sys
sys.path.insert(0, "tools/arcs-verb-eval")
from arcs_verb_eval import evaluate
upstream = json.load(open("prints/synthesis.json"))
report = evaluate("recapitulation", open("note.md").read(), upstream=upstream)
json.dump(report, open("prints/recapitulation.json", "w"), indent=2)
print(report["fate"]) # the step's go/stop stamp
if not report["ok"]:
sys.exit("stopped by: " + ", ".join(report["failed"]))- name: Recapitulation by ARCS
run: tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json --out prints/recapitulation.json
- name: Keep the paper
if: always()
uses: actions/upload-artifact@v4
with:
name: recapitulation-print
path: prints/recapitulation.json# .pre-commit-config.yaml
repos:
- repo: local
hooks:
- id: recapitulation-by-arcs
name: Recapitulation by ARCS
entry: tools/arcs-verb-eval/arcs-verb-eval recapitulation note.md --upstream prints/synthesis.json
language: system
files: 'note.*\.md$'
pass_filenames: falseFor builders
All five steps as one GitHub workflow
# .github/workflows/check-before-you-say-it.yml
name: check-before-you-say-it
on: [pull_request]
jobs:
five-steps:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Self-test (the checks check themselves)
run: tools/arcs-verb-eval/arcs-verb-eval --self-test
- name: All five steps
run: tools/arcs-verb-eval/arcs-verb-eval pipeline --config arc.json
- name: Report
if: always()
run: |
tools/arcs-verb-eval/arcs-verb-eval export prints --format md >> "$GITHUB_STEP_SUMMARY"
tools/arcs-verb-eval/arcs-verb-eval export prints --format junit --out prints/junit.xml
- uses: actions/upload-artifact@v4
if: always()
with:
name: five-step-prints
path: prints/What each step saves
One format for all five. The next step reads it.
{
"verb": "synthesis",
"fate": "hold",
"upstream": "evaluation.json#sha256:9f1c",
"gates": [
{
"id": "no-new-claims",
"ok": false,
"evidence": "line 212"
}
],
"rubric": [
{
"row": "D-specificity",
"score": 2,
"judge": "human"
}
],
"sycophancy": {
"agreement": 4,
"dissent": 3,
"verdict": "OK",
"blocking": false
}
}