# Boundary decision pilot

This pilot asks independent readers to judge ten proposed edits across six complete programs, using either full user-package source or an operator-selected packet of existing tool records and compiler-located source excerpts. The questions, public premises, key, material hashes, and per-kind quality floors were frozen before dispatch. It tests these selected decisions, not open-ended defect discovery.

## Preserved first round

`protocol.json` and the first results are unchanged. The source reader answered all ten correctly. The artifact reader answered nine correctly and deferred on q3: a may-influence edge did not establish the actual returned clock value. That safe abstention failed the frozen preservation floor. `results/gate.json` records the failure.

## Revised packet and fresh readers

Round two adds the compiler-located `consumed` definition returned by the existing Explorer path tool. It leaves the source inputs, questions, and key unchanged. `protocol-v2.json` was committed in 4b4d0a5b before either fresh reader began. Both readers answered all ten correctly; `results/gate-v2.json` records the pass. Raw reader responses, dispatch timestamps, wrapped runs, and both protocols are retained.

The common sheet plus revised artifact contains 9,960 UTF-8 bytes; the common sheet plus source contains 30,485. The exact costLift is 3.0607429718875503, based only on materialBytes. Runtime token accounting was unavailable for these readers. This does not measure reasoning, output, tool overhead, artifact generation or maintenance; it is not a measured token reduction. One reader pair per round and selected cases do not establish general product effectiveness.

`node eval/corpus/boundary-decisions/reproduce.mjs` generates the first packet from the real CLI; `node eval/corpus/boundary-decisions/refine-read-packet.mjs` adds the saved compiler-located definition. See generation.json, generation-v2.json, source-provenance.json and key-rationale.md for selection, commands and grounding. The raw files are preserved so the recorded experiment can be replayed without silently replacing inputs.

The six frozen source inputs were independently audited in docs/audits/2026-09-14-boundary-corpus-bosatsu-audit.md. The Forum copy intentionally retains a discovered inherited JSON-escaping defect; that defect is recorded separately and is outside these selected edit questions.

Replay each recorded verdict with the strict gate:

```sh
yichus eval-gate results/artifact-run.json results/baseline-run.json --decision-protocol protocol.json --out results/gate.json
yichus eval-gate results/artifact-v2-run.json results/baseline-v2-run.json --decision-protocol protocol-v2.json --out results/gate-v2.json
```

The first command is expected to exit nonzero with a produced failing verdict.

## Separate fresh trial with runtime receipts

After freezing protocol-tokens.json in commit 070bb337, fresh independent readers
ran through Codex CLI 0.153.4 with read-only ephemeral sessions, the same CLI
default settings, complete assigned packets embedded in their prompts, and no
tool calls. The answer keys and arm material bytes stayed unchanged; cli-reader.md
is a shared instruction to consume the inline packet without tools. Both readers
answered every case correctly. This is a separate pair, not token accounting
retroactively attached to the earlier answers.

The recorded artifact cost is 17,952 tokens and source cost is 23,099; tokenLift is 1.2867090017825311. Both the prospective gate (results/tokens/gate.json) and original token/accuracy gate (results/tokens/token-gate.json) pass.

Cost is the sum of the completed CLI turn's reported input_tokens and
output_tokens. Cached-input and reasoning-output breakdowns are retained and
are not added a second time. The receipts preserve the complete usage object,
run id, protocol hash and raw event-stream hash. The events contain only start,
final-answer and completion events: no commands or external tools ran. Runtime
input includes the CLI's instructions as well as the evaluation packet, so this
is a different cost basis from UTF-8 packet size. Artifact generation and
maintenance remain outside the comparison. The runner version and flags are
recorded; the default model was not overridden or named by the JSONL stream.

Every exact prompt, dispatch record, final answer, raw event stream, receipt and
gate result is preserved under results/tokens/. Replay with:

```sh
yichus eval-gate results/tokens/artifact-run.json results/tokens/baseline-run.json --decision-protocol protocol-tokens.json --out results/tokens/gate.json
yichus eval-gate results/tokens/artifact-scores.json results/tokens/baseline-scores.json --out results/tokens/token-gate.json
```

Accounting references: [Codex JSONL completion usage](https://learn.chatgpt.com/docs/non-interactive-mode#make-output-machine-readable) and [output usage including reasoning](https://developers.openai.com/api/docs/guides/reasoning#managing-the-context-window).
