Evidence

Benchmark: AI reviewers finding planted bugs with fact tools or source

Do machine-checked program facts help an AI reviewer find planted defects faster or more accurately than reading source? We ran a paired review experiment to find out.

Each round generates a program in Bosatsu, a small total language with declared effects, with planted defects and a hidden answer key. Eleven rounds are recorded (000–008 and 010; round 5 ran twice) on codebases from 2,028 to 46,922 lines. The analysis depends on Bosatsu's typed intermediate representation.

Two fresh AI reviewers (“probes”) then hunt the defects. Each round's tables call them by their mode. The tooling probe is briefed to work from the fact tools, yichus's static-analysis queries, which answer questions about the compiled program (who writes this field, what reaches this output, does this declared rule hold?) with machine-checked facts instead of source text. The baseline probe reads the source. Neither probe sees the key, and both are scored against it.

One point needs clarifying, because the numbers below would otherwise contradict the setup: the tooling probe is not blocked from opening files. It is briefed to answer from facts, and how much source it read is measured and published as the “% source read” column. On the tooling rows that figure ranges from 0% to 27% across rounds. On the early, smaller rounds the baseline rows read all of the source, or nearly all of it: every baseline row through round 005a is at 100% except round 004's haiku-tier baseline, which read 4,800 of 4,832 lines (99.34%). Once the codebases outgrew the token budget the baseline rows fall to between 7% and 50% (rounds 005b–008; the capped baseline in round 008 managed 7.03%, and the highest is round 007's haiku-tier baseline at 50.4%). The column records what happened, and nothing in the setup capped it. The baseline probe is given no fact tools.

Read the numbers as what they are: a self-run benchmark. The same project generates the challenges and builds the tools. Each table cell is one run, so nothing here is an average. And between rounds the tools get improved against the previous round's misses, which is the point of the loop. Two things keep it honest. Failed rounds stay on this page with their reasons, and every score renders from committed result artifacts. Each round below links to its raw scores. The per-round ./tooling-evals/*.json links are those same committed files, copied to this site at deploy time; the GitHub tree is where they live in the repository. Model labels are harness tiers (default, sonnet, haiku), not exact model IDs, so a round cannot be re-run under identical conditions.

Table fields, contract cases, and pass/fail rules

Each round below renders one table, one row per reviewer run, then a verdict line. Here is a real row read out in full, the first row of round 008:

tooling-default-uncapped · tooling · default · 96.2% · 80.0% · +5 · 0 · 137.0k · 21.5 · 0.07%
A run named tooling-default-uncapped, which used the fact tools (mode tooling) on the frontier-tier model with no spend cap. Against the 46,922-line program it found 96.2% of the 26 planted defects (25 of 26) and 80.0% of the 5 defects tagged subtle, i.e. 4 of 5. It scored +5 on calibration (it never claimed a certainty the checker could not grant), made 0 false accusations, spent 136,989 tokens over 21.5 minutes, and read 0.07% of the program, about 33 lines out of 46,922.

Probe: the name of one run, written mode-model plus a variant tag where a round ran the same pair more than once (-uncapped and -capped60k in round 008, -budget in rounds 006 and 007). Mode: which reviewer this was. tooling is the probe briefed to work from the fact tools, and baseline is the probe that reads source. These two words are the main axis of the whole benchmark. Model: Anthropic Claude tiers as run by the coding harness. “default” is the harness's default (frontier-tier) model at the time of the round, and “sonnet” and “haiku” are the mid and small tiers. Rounds record the tier, not the exact model ID, so a round cannot be re-run under identical conditions. Recall: how many of the planted defects the reviewer found (a partially-correct answer counts half). Subtle: recall restricted to the defects designed to survive a first pass: buried behind indirection, spread across files, or invisible to static reading. “n/a” means the round planted none. Subtle denominators are small (2–5 defects), so one half-point moves the number a lot. Calibration: did the reviewer overclaim? Scored over the round's explicit verdict questions (“does this declared rule hold?”), not per defect: +1 for a correct verdict or a correct “can’t be sure, because…”; −1 for declaring safe what the ground-truth checker could not confirm. The number of scored questions varies by round, so compare within a round, not across rounds. A positive score is printed with a + and shown in green, a negative one with a − in red, and a zero is printed bare and uncoloured. The sign carries the meaning, and the colour only reinforces it. No round has yet produced a negative score. A zero means the round's verdict questions cancelled out or scored nothing. It says nothing bad about the reviewer. False +: findings the reviewer asserted that the answer key contradicts, i.e. wrong accusations. Lower is better, and 0 is clean. Tokens: the AI agent's total token usage, which is the cost of the review. Minutes: wall-clock time the run took. Recorded for context only. It plays no part in the pass/fail rule, and it is not comparable across rounds run on different hardware. % source read: lines of the generated program the reviewer read, over its total line count. Re-reads count, so a reviewer that revisits files can exceed 100%.

Planned budget, in the grey line under each round title from round 001 on (the pilot had none): the tokens budgeted for the round as a whole before it ran, kept so plan can be compared against actual. It is bookkeeping. It constrains no probe, and the rows beneath it routinely sum to less. Round 008 did test real caps, and reports honestly that they failed to bind.

What “contract accuracy” is, and why it is not Recall. The two accuracy figures in the verdict line are not scored over the round's whole defect set. An artifact-contract case is one small program added to a standing corpus, carrying a ground-truth label: is this binding's result fabricated, does this declared rule survive every interleaving, is this permission check real? Each round contributes two such cases, one drawn from each defect family, and the gate asks both reviewers to classify them; contract accuracy is the fraction they get right, so with two cases it can only ever read 0, 0.5 or 1. (Round 010 is the exception: it scores 14 cases, its whole conformance census.) Recall, by contrast, is scored over every planted defect. The two can and do disagree: rounds 006, 007 and 008 all pass with contract accuracy tied at 1.0 while the Recall column plainly shows the tooling probe finding fewer defects than the source reader. That is a real weakness of the gate. A two-case accuracy test is too coarse to detect the recall gap its own table displays. Read the verdict as “cheaper at no cost on the contract cases”, and read the Recall column for who found more.

Pass/fail: after applying the contract-case definition above, a round passes if the fact-using reviewer did not lose contract accuracy and was cheaper. Formally, that is contract accuracy Δ ≥ 0 and tokenLift above 1.000, where Δ is the fact reviewer's contract accuracy minus the source reader's, and tokenLift = source-reader tokens ÷ fact-reviewer tokens, so above 1 means the facts were cheaper. Each verdict compares exactly one pair of probes. Where a round ran several, it is the frontier-tier (“default”) pair. Round 008 ran two frontier-tier pairs, and its verdict uses the uncapped one. Rounds that ran a single pair are judged on that pair, whatever its tier: round 010's verdict is its sonnet pair. Round 005a is the one round with no verdict, because round 5's official comparison was made in round 005b.

Each round's raw scores JSON, linked below the round, itemizes the exact per-item scoring: every planted defect's found/missed status and every calibration event. Scoring is adjudicated by this project's own harness against the generator's answer key. Round numbering: round 009 is absent because iteration 9 only built tooling and ran no measurement round, so there was nothing to score.

Evaluate the change an agent would make

Future decision evaluations use a frozen protocol with separate quality floors for finding defects, preserving required behavior, and recognizing insufficient evidence. Gains in one task cannot conceal a regression in another. Unsafe edits fail the gate, and the same pair must also be cheaper on its declared cost basis. Read the protocol and its limits. The historical outcomes below retain the contract-case definition under which they were measured.

Recorded results from committed artifacts

Best result under nominal caps that did not bind: on a 46,922-line codebase with 26 planted defects, an AI reviewer working from the fact tools found 25 of 26, and a reviewer reading source found 23. (Given an unlimited budget, that same source reader found all 26, at 1.85× the cost. Round 008 has both figures.)

Loading result artifacts…

The record so far: 6 of 10 judged rounds passed the gate, 4 rounds failed, and 1 round carries no verdict. Of the failures, all lost on cost, and 1 round lost on contract accuracy (1 on both). The gate does not measure defects found. Counting by those instead, the source reader matched or beat the fact tools in 8 rounds of 11, comparing the same pair each verdict was computed from. The tools' advantage is cost and scale, not raw finding power.

Early UI fixtures and experiments live in the archive.

Where this fits

What this is about: Fresh-agent testing