Idea

How the project tests itself

Attack your own tools with fresh agents

An author who tests their own tool or page fills every gap from memory. We hand our analyzers, our docs pages, and our own Bosatsu to agents that start with none of the project's context, and turn what those agents get wrong into a list of defects to fix.

See the probe results round by round →Read the docs-probe protocol →

Credit: usability testing with new participants, red teaming, cold-read review, LLM-as-judge evaluation, answer keys fixed before the test, and control arms are all established practice. Yichus applies them to its own analyzers and docs pages.

Give the agent only what a user would have

A docs probe gets the page as a browser renders it and nothing else. It answers a set of questions and must write "NOT ANSWERABLE FROM PAGE" instead of filling a gap from general knowledge. The answer key is written from code and recorded runs before the page is touched, so the page's wording cannot shape the key.

A tooling probe hunts defects planted in a freshly generated Bosatsu program. It never sees the answer key, the brief that produced the program, or the other probes' reports. One probe works from the analysis tools; a second reads the source. The project's policy is that each kind of tool output agents read ships with this comparison: yichus eval-gate passes it only if the tool arm is at least as accurate as the source arm and cheaper in tokens.

Adversarial agents get one checker each and a single job: find an input where the checker reports safe while the guarantee was never established. They report and fix nothing.

Turn each failure into a named deficit

Probe reports become closed-form entries: a term used before its definition, a claim with no route to evidence, a dead struct field no query could surface. Each entry goes in a committed register with the round that found it and the change that closed it, such as the tooling register and the docs register.

A fix is checked by new agents that never saw the old version. When prose changes, fresh readers read both versions without knowing another exists, and two scorers grade the mixed answers blind. The edited page is kept only if its readers believe no more false things and answer about as well, and the page is shorter or within a word allowance set in advance.

Check the finding before acting on it

No finding from a probe, adversarial agent, or review bot is accepted or rejected on its say-so. The mechanism is re-derived from source and the counterexample is run where that is feasible. Rejecting one as a known limitation requires quoting where that limitation is disclosed. A confirmed soundness fix ships with a test that was first seen failing on the unfixed code.

Bosatsu written by the project's own agents gets a fresh auditor before merge. Those agents can read the Scala interpreters underneath, so the auditor also records each place a user without that access would have been stuck. The audit records keep those entries. In one, the auditor found that a benchmark's question packet gave away answers before any reader was dispatched.

What a passed probe does not show

A pass means these readers answered these questions. It says nothing about human readers, other models, or questions nobody wrote. The project writes the programs, the keys, and the tools, and improves the tools against the previous round's misses, so every result is a self-run benchmark with small samples.

Dig deeper: where the limits are recorded

Freezing the questions and hashing the materials prevents an unnoticed swap during replay. It does not make the answer key impartial. Probe model labels are harness tiers rather than exact model IDs, so a round cannot be re-run under identical conditions. A probe that fails to find a defect is evidence about that probe and that program, not proof the defect class is absent.

The decision-evaluation protocol covers how questions and floors are frozen, and the evidence index keeps failed rounds beside passing ones.

Where this fits

Measure it