Evidence

Evidence for the tools themselves

Does the extra visibility help the next decision?

A fact about a program can be correct while the edit it suggests is wrong. A field with no extraction in Bosatsu can still be part of a public response. A result built from literals can vary when an input selects a branch. A function absent from a sampled reference set can still be needed by a registered handler.

So we test what a person or agent does with the evidence. Do they find the real problem? Do they keep the behavior the program must keep? When the evidence is not enough, do they say what must be checked before changing the program?

Test the decisions separately

A prospective evaluation fixes its questions, public requirements, accepted answers, unsafe edits, and quality floors before readers begin. Independent readers receive either source code or generated artifacts, with the same questions and public premises. An artifact cannot earn credit from information withheld from the source reader.

Defect detection, preservation, and insufficient-evidence decisions are scored separately. Better preservation cannot compensate for missing more defects. Deferring is correct when the case requires more evidence; deferring on every case does not satisfy defect detection. Missing, repeated, or extra answers invalidate the run.

Require useful accuracy and a real cost comparison

The artifact arm must meet every frozen quality floor, match or beat the source arm within each task type, and recommend no edit the key identifies as unsafe. It must also be cheaper than that same source run.

The gate distinguishes reported runtime tokens from input bytes. A byte comparison counts the common questions and each arm's material. It measures the size of that reading packet, not model reasoning, output, tool overhead, or the token cost of an image. Runtime token comparisons require preserved runtime receipts for both readers. Neither basis measures the entire cost of generating and maintaining the artifact.

Materials and protocols are checked by their hashes. This prevents an unnoticed substitution during replay; it does not certify an impartial answer key. The project's own cases remain a small, selected sample. Keep failures, disagreements, and the stated limits alongside a pass.

A boundary decision pilot, including its failed first round

We tested whether selected tool evidence helps an agent judge proposed edits to API fields, transaction reads, calculation rules, and UI state updates. Independent readers received either complete user-package source or an operator-selected packet of tool records and compiler-located excerpts. The trial used ten proposed edits across six programs, with identical questions and public requirements for both readers. The key and quality floors were frozen before dispatch.

In the first round, the source reader answered all ten correctly and the artifact reader answered nine. It deferred on replacing a clock result with zero: a “may influence output” edge did not establish the value returned. That safe abstention failed the preservation floor. We kept the failure, added the compiler-located definition, froze the revised packet, and used fresh independent readers. Both then answered all ten correctly.

Recorded results for these selected proposed edits
PacketArtifact correctSource correctGate
Initial9 / 1010 / 10fail
With the located definition10 / 1010 / 10pass
Fresh CLI readers, same revised packet10 / 1010 / 10pass

For the revised round, the common questions plus artifact contained 9,960 UTF-8 bytes; the common questions plus source contained 30,485. These are input bytes, not measured model tokens. The comparison excludes reasoning, output, tool overhead, and producing or maintaining the packet.

A separate fresh pair then judged the same revised evidence and source using the read-only Codex CLI, with identical instructions to read their complete inline packet without tools. We retained the runtime event streams. Both answered all ten correctly; reported input plus output tokens were 17,952 for the artifact reader and 23,099 for the source reader. These costs include the CLI's input context and generated output. They exclude producing and maintaining the evidence packet.

Frozen runtime-token protocol; runtime trial verdict; artifact usage receipt; source usage receipt; artifact event stream; source event stream.

This is one reader pair per round on selected tasks. It does not establish open-ended defect discovery, general product effectiveness, or the usefulness of agenda cards alone. The packet includes selected source excerpts.

Method, reproduction and limitations; common questions; initial protocol and material hashes; initial failing verdict; revised protocol and material hashes; revised passing verdict; revised evidence packet; artifact reader answers; source reader answers.

Keep the historical record intact

The historical measurements used a narrower contract-case gate. Some passes coexist with lower defect recall. Their original scores and verdicts remain intact. A new decision gate does not retroactively validate those experiments or establish broad product effectiveness.

The CLI selects the prospective protocol explicitly:

yichus eval-gate artifact-run.json baseline-run.json \
  --decision-protocol protocol.json --out verdict.json

Use the organization lenses to inspect a proposed change, and the relevant checks to test the property it should preserve. A legible diagram and a passing check answer different parts of that decision.

Where this fits

What this is about: Fact Tooling