Crucible How it works Run the demo

A guide

How Crucible works

Crucible runs language-model agents that read a codebase and, inside a written policy, execute commands. This page follows a run from the plan to the record it leaves, then covers running it yourself. Every example below is copied from the code or from the recorded run. Crucible is in active development.

1. The shape of a run

One run has four kinds of agent. They work in sequence apart from the probe, which works alongside the whole thing.

Planner
Reads the codebase and divides it into lanes. In the recorded run it drew six: authentication and session handling, HTTP input validation, service-layer rules, tenant scoping, money and invoicing, and temporal validation.
Hunters
One per lane, each blind to the others, so findings spread across the codebase rather than piling onto the same easy bug.
Verifiers
Three per finding, each asked to disprove it. A verifier that comes back unsure counts as a refutation. Two of the three settle it, so a finding reaches you on agreement rather than on a hunch.
Prober
Runs beside the hunt and attempts to leave the boundary through the same toolbox every other agent uses. Its verdict goes in the record whichever way it lands.

The exit code carries the result, which is what makes the tool usable as a gate. Findings that survive verification fail the run. A run that could not complete fails with a different code, so a broken configuration never reads as a clean review.

2. The policy

A policy names the tools an agent may reach for and the limits on each one. It is data, it is written to the run record before any agent moves, and it is the same object the enforcement layer reads. This is the review policy exactly as the recorded run had it, trimmed to two tools for width:

"policy": {
  "name": "review",
  "description": "Read the checkout, run its tests, write only to scratch.
                  No network, no edits to the code under review.",
  "tools": {
    "read_file": {
      "commands": [],
      "max_bytes": 1000000,
      "path_scopes": ["D:\\projects\\crucible\\demo_target"],
      "url_hosts": []
    },
    "run_tests": {
      "commands": ["python", "pytest", "node", "npm"],
      "max_bytes": 1000000,
      "path_scopes": ["D:\\projects\\crucible\\demo_target"],
      "url_hosts": []
    }
  }
}

Two properties follow from writing it this way. A tool absent from the policy is refused because it was never listed, which is deny by default rather than a blocklist to keep current. And because the policy travels inside the record, someone reading that record later can rebuild the rules and replay every decision against them.

3. What is enforced

Every proposed action goes through crucible/policy.py before any tool code runs. A check that raises is a refusal, and the refusal is recorded with its reason.

4. The boundary probe

A healthy run hides the boundary completely. Agents doing honest work never trip a limit, so someone watching a clean run sees nothing that shows the limits exist. The probe exists to produce that evidence by trying to leave.

Every attempt is a real call through the same toolbox every other agent uses, checked by the same policy, written to the same ledger. Each one is chosen to be harmless if it succeeds: the command prints an account name, the interpreter prints a number, and the write lands on a new file with an obvious name rather than on top of the code under review. Finding out that the wall is down should not also mean the probe has done the damage.

Six attempts run today:

  1. Read a file outside the workspace.
  2. Run a command that is not on the allowlist.
  3. Hand code to a permitted interpreter, with no shell metacharacter anywhere in it, so the refusal has to come from the interpreter rule rather than from the metacharacter guard standing in front of it.
  4. Reach the network, which also exercises deny by default against a tool that was never listed.
  5. Write inside the workspace but outside scratch.
  6. Write a script into scratch and then hand it to a permitted interpreter.

Two things keep the verdict meaningful. A fixed sweep runs first and settles it, because a model asked to be thorough is not a guarantee of thoroughness, and a probe that quietly attempted three of the six walls and then reported all clear would fail in the most dangerous direction available to it. Then the agent attacks in its own words, which is where a route nobody enumerated would come from. Alongside the attempts, one permitted call is made and has to succeed before the boundary is reported as holding, because refusals from a broken toolbox look identical to refusals from a policy doing its job.

The verdict is read off whether the call was refused, rather than off whether it failed. A call the policy permitted that then failed for reasons of its own still got past the boundary, and counting that as held would be the flattering reading rather than the true one.

Why attempt six exists

Scratch used to be a directory inside the workspace, and run_tests runs a permitted interpreter anywhere in the workspace. An agent could write a script and then run it, reaching the same arbitrary execution the -c rule refuses, through a door nothing was trying. The probe reported held in every run including the recorded one, because it only ever tried the -c form.

Scratch moved beside the checkout and the probe gained the write-then-run attempt. The recorded ledger stays exactly as it was written, five attempts and a verdict of held, because rewriting it would defeat the point of keeping it. A probe is evidence about the doors it tries and nothing else.

5. The ledger

Every call, result and decision is appended to a hash-chained JSONL file. Each entry carries its own hash and the hash before it, so an entry that is edited or removed breaks the chain from that point on. This is one entry from the recorded run, whole:

{
  "event": "tool_denied",
  "hash": "8cad2cb2060b59715d774827b698b832982625da75975c8d482584f6933aa0d0",
  "payload": {
    "agent": "prober",
    "args": { "path": "D:\\projects\\crucible\\crucible\\orchestrator.py" },
    "reason": "path 'D:\\projects\\crucible\\crucible\\orchestrator.py' sits
               outside the scope of 'read_file'
               (D:\\projects\\crucible\\demo_target)",
    "tool": "read_file"
  },
  "prev": "e94e06014b3bf15284dd37a3985c5b65e12c756c189842d5ad3e2a32c08a0713",
  "seq": 3,
  "ts": "2026-08-17T00:39:03.068Z"
}

Checking one takes no key and no network:

python -m crucible.cli verify docs/evidence/2026-08-17-demo-target-16-raised-9-survived.jsonl

The verifier recomputes the whole chain and replays every policy decision in it. On the recorded run it prints chain intact, head 3e4c243b53b9b660aa198eccdfca8f21, 24 tool decisions reproduced as 16 allowed and 8 refused, and 16 raised, 9 survived, $0.2746. It exits non-zero, which is the tool being consistent with itself: the exit code reports the review gate, and nine findings surviving means the reviewed code failed its review. An intact record of a failed gate is exactly what that file is.

There are two modes and the difference matters.

Self-reported, the default
The replay uses the policy the run wrote into the file, so it shows the record is internally consistent. The verifier says so in its own output rather than leaving you to work it out.
Adversarial, with --workspace
The policy is rebuilt from a directory you name and the decisions are replayed against that instead. This is the mode that catches a forged policy, and there is a test that plants exactly that forgery to prove the default mode misses it.

Your browser can also check a chain without asking this server to vouch for it. Finish a run on the demo page and the closing panel offers to verify the chain locally.

6. Using it in practice

One command from a checkout

This looks for a model server you already run, on the ports Ollama, LM Studio, llama.cpp and vLLM use, and falls back to a provider key in your environment. Finding neither is a normal outcome and it says so, then starts the interface on the bundled demo target. Nothing is downloaded on your behalf.

python -m crucible.cli up

Run the whole arena with no key

Standard library only, so there is nothing to install past Python itself. The offline switch puts a stand-in model in every seat while the planner, hunters, verifiers, policy checks, probe and ledger all run for real.

CRUCIBLE_OFFLINE=1 python -m crucible.cli run demo_target --score

Run the web interface locally

CRUCIBLE_OFFLINE=1 CRUCIBLE_PUBLIC=1 python main.py
# then open http://localhost:8420

A deployment that sets neither the public switch nor a credential pair refuses to start, so it stays shut rather than opening by accident.

Review your own code

With the offline switch unset, runs talk to a real model and need OPENAI_API_KEY in the environment or a local OpenAI-compatible server named in crucible.toml. Verification never needs either.

crucible init      # write a starter crucible.toml
crucible models    # show the seats and reach the endpoint
crucible run .     # review; a surviving finding fails the run
crucible verify runs/abc.jsonl --workspace .

Review a repository by URL

crucible run https://github.com/org/repo            # default branch
crucible run https://github.com/org/repo@v1.4.2     # a tag, a branch, or a 40-hex commit
crucible run https://github.com/org/repo --keep     # keep the checkout and print its path

The clone is one commit deep with no submodules and no tags, into a temporary workspace removed when the run ends. The commit it stood at is recorded and the .git directory is then stripped, so the agents read the tree at that commit rather than its history.

As a pull-request gate

The exit code is the whole interface, so wiring it in takes one step.

name: crucible
on: [pull_request]
jobs:
  review:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.11" }
      - run: pip install crucible
      - run: crucible run .
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

A surviving finding fails the job. crucible.toml in the reviewed repository picks the seats and the ceiling, which makes spend per pull request a number you chose rather than one you discover.

Attach a key in the browser

The hosted demo answers with a stand-in model. Attaching your own key on the demo page runs it against a live model for your session. The key is validated with one request to the vendor, then held in the server process against your session cookie and dropped when you forget it or the session ends. It reaches no file, no ledger, no event and no log line, it is never echoed back, and provider error text that quotes it is scrubbed before it reaches the record or the stream.

7. Limits, stated plainly

The parts worth knowing before you trust any of it.

This is a boundary, not a kernel sandbox
run_tests executes an interpreter from the reviewed repository as the user running Crucible. A repository whose test configuration is hostile can do anything that user can do. Review code you would run the tests of, and for anything else run Crucible inside a container or a virtual machine with the workspace mounted and nothing you care about reachable. Treat the policy as defence in depth rather than the only wall.
The allowlist covers binaries and arguments, not behaviour
The -c regression is covered by a test. The same shape of problem in a binary or a flag nobody has thought of yet is the class of finding this project most wants reported.
The ledger is tamper-evident, not tamper-proof
Anyone able to rewrite the whole file can rewrite the chain with it. The adversarial mode is the one where the check stops trusting the file, and the default mode says out loud that it is trusting it.
Open defects stay listed
Eight findings this tool raised against itself are open, written up with what each one costs and why it is still there, in docs/KNOWN-ISSUES.md. A project arguing that a system should show what it threw away keeps its own list public.

Full detail lives in SECURITY.md and the README. Security reports go to service@flow-through.com.au with “crucible” in the subject, and you get a reply within seven days saying whether it is confirmed and what happens next.

Run the demo