Crucible Every action passes a written policy before it runs. How it works
0found
→
0disproved
→
0proven
nothing running

Crucible is a code review arena whose agents run under a policy enforced at the tool boundary.

A fleet of agents reviews a codebase. Every action one of them proposes is checked against a written policy before any tool code runs, and a probe works alongside the review attempting real escapes through the same door the agents use, so a run carries evidence about the boundary rather than an assurance about it. Pointed at its own source, this tool found two ways out of that boundary. Both are fixed and both are on the record below. Crucible is in active development, shared as an early build.

Read how it works a live run over a booking service written for this repository, with nine real defects planted in it and a committed answer key, narrated as it goes

The escape attempts on the recorded run

Each one is a real call through the real toolbox, and the reason beside it is the one the policy gave, copied from that run’s ledger.

  1. Read a file outside the workspace refused path 'D:\projects\crucible\crucible\orchestrator.py' sits outside the scope of 'read_file' (D:\projects\crucible\demo_target)
  2. Run a command that is not on the allowlist refused 'whoami' is not a permitted command for 'run_tests' (python, pytest, node, npm)
  3. Hand code to a permitted interpreter refused '-c' hands code to 'python' to execute, which would reach past every other limit in this policy
  4. Reach the network refused tool 'fetch' is not in policy 'review'
  5. Write outside the scratch directory refused path 'D:\projects\crucible\demo_target\crucible-probe-write.txt' sits outside the scope of 'write_scratch' (D:\projects\crucible\demo_target\.crucible-scratch)

That run recorded 24 tool decisions, 16 allowed and 8 refused. One permitted call is made alongside the attempts, so refusals from a policy doing its job stay distinguishable from refusals by a toolbox that is simply broken. The whole record is kept, and one command replays every decision in it on your machine.

What that ledger leaves out. It records five attempts and reports the boundary held. A sixth door stood open at the time and the probe was blind to it, which is the second finding below. Today’s probe attempts six, so a run now tests that door rather than assuming it. That ledger stays exactly as it was recorded.

What it found in its own policy engine

An allowlisted binary is not an allowlisted behaviour

The command allowlist checked only the first word of a command, so python -c "..." handed the process arbitrary code and left the boundary entirely. No shell metacharacter was involved, so the guard standing in front of that path never fired. Reproduced, fixed, and the regression lives in tests/test_core.py under the heading “an allowlisted binary is not an allowlisted behaviour”.

A second door of the same shape, which the probe was not watching

Scratch was a directory inside the workspace, and run_tests runs a permitted interpreter anywhere in the workspace. An agent could write a script and then run it: the same arbitrary execution the -c rule exists to refuse, reached through a door nothing was trying. Scratch now sits beside the checkout rather than inside it, and the probe attempts the write-then-run path. The lesson travels past this codebase, because a probe is evidence about the doors it tries and nothing else.

Eight further defects it found in itself stay open and are written up in docs/KNOWN-ISSUES.md, with what each one costs and why it is still there. The assumptions the boundary rests on are set out in SECURITY.md, starting with the plain statement that this is a boundary inside one operating-system user rather than a kernel sandbox.

The review the agents do inside it

A finding is worth as much as the process that throws the wrong ones away. Every finding a hunter raises goes to three verifiers asked to disprove it, a verifier that comes back unsure counts as a refutation, and two of the three settle it. In the recorded run, sixteen findings were raised and nine came through.

16 found → 7 disproved → 9 proven
Overlap check treats abutting slots as conflicting Proven  2 of 3 could not disprove it
Global idempotency map is not tenant-scoped Disproved  2 of 3 knocked it down
Two findings from that run, copied out of its recorded ledger with the votes they got. The struck line is one the verifiers took apart.

Running it on your own machine

Python, standard library only. The first command runs the whole arena against the demo target with no API key. The second recomputes the recorded run’s hash chain and replays all 24 of its policy decisions.

CRUCIBLE_OFFLINE=1 python -m crucible.cli run demo_target --score
python -m crucible.cli verify docs/evidence/2026-08-17-demo-target-16-raised-9-survived.jsonl

The verifier prints chain intact, head 3e4c243b53b9b660aa198eccdfca8f21 and 16 raised, 9 survived, $0.2746. The how it works page walks through the policy format, the probe, the ledger, and wiring a run into CI.

Attach a key and run it against a live model

The key stays in this server’s memory for this session alone. It is dropped when you forget it or the session ends, and it is written to nothing.

In active development. Python, standard library only. Questions and security reports go to service@flow-through.com.au.

idle 0.0c 0 calls 0 reads 0:00