# Iteration 001 — the first number (2026-09-05)

## Design (written before any arm ran)

- **Arms and runs.** Per task, three fresh sonnet agents per arm
  (general-purpose, isolated workspaces under a scratch directory that
  holds only the brief, the property list, the arm's surface file, and
  for arm A a copy of `demos/service/forum.bosatsu` and the jar; the
  instruction forbids reading anything else). Tasks run one at a time,
  six agents at once. Every agent ends with the same six-line report
  (deliverable, compiles / runs, verify, claimed property ids,
  iterations, notes); the claimed ids feed calibration.
- **Scoring.** `harness/score-workspace.sh` deploys or loads each
  deliverable and writes that task's classifications; a deliverable
  that does not compile, verify, deploy, or load gets every property
  of the task not-established with the reason (did-not-compile is a
  number, not an exclusion). `harness/merge-runs.mjs` joins run index
  k of the four tasks into arm file k, so each arm has three
  classifications files covering the seventeen properties and the
  gate is run three times, k against k, with `yichus eval-score` then
  `eval-gate --min-lift 1.0`. Tokens: each agent's `subagent_tokens`
  from its completion notification, summed over the four tasks of run
  k, into the arm-A file as `sourceTokens` (arm B's total) and
  `toolTokens` (arm A's total), so `tokenLift = sourceTokens /
  toolTokens` reads "TypeScript tokens per Bosatsu token".
- **Gate (the README's).** accuracy(A) ≥ accuracy(B) and tokenLift > 1
  on each of the three pairs; the iteration passes when all three
  pass. A failing gate is published as such.
- **Extra columns**, per run: did-not-compile (arm A files that fail
  `build` or `verify`; arm B modules that fail to load), iterations
  (the agent's own count), calibration (claimed ids the harness did
  not establish, and established ids not claimed).
- **Prediction.** Arm A's accuracy is at or above arm B's on tasks 1,
  2, and 4 (the properties are plain handler logic and `api verify`
  catches route and access slips); task 3 is the open question on
  both arms (arm B must use `cas` correctly under the interleaver; arm
  A must write a dist world that reaches holds). Tokens: arm A reads a
  500-line starting file and runs a JVM tool, so its token count is
  not obviously lower; a lift below 1 is a real possibility and is
  the number this iteration exists to get.

## Results (appended after the runs; the design above is unchanged)

**The gate fails on all three run pairs.** Arm A (Bosatsu with the
tools) established 16 of 17 properties on every run; arm B (TypeScript
from scratch) established 16, 17, and 17. Arm A spent roughly twice
the tokens. Published as such: the first number is a loss on both
axes, and it says where.

| pair | verdict | accuracy A | accuracy B | tokens A | tokens B | tokenLift (B/A) |
|---|---|---|---|---|---|---|
| run 1 | fail | 16/17 (0.941) | 16/17 (0.941) | 527,348 | 286,849 | 0.544 |
| run 2 | fail | 16/17 (0.941) | 17/17 (1.000) | 482,343 | 242,753 | 0.503 |
| run 3 | fail | 16/17 (0.941) | 17/17 (1.000) | 540,819 | 238,418 | 0.441 |

Per task (established / properties, then the agent's own iteration count):

| task | A run 1 | A run 2 | A run 3 | B run 1 | B run 2 | B run 3 |
|---|---|---|---|---|---|---|
| 1 edit window | 5/5 (9) | 5/5 (9) | 5/5 (12) | 5/5 (1) | 5/5 (1) | 5/5 (1) |
| 2 locked thread | 5/5 (4) | 5/5 (9) | 5/5 (18) | 5/5 (1) | 5/5 (1) | 5/5 (2) |
| 3 reply-cap race | 1/2 (34) | 1/2 (84) | 1/2 (68) | 1/2 (7) | 2/2 (4) | 2/2 (5) |
| 4 pinned then recent | 5/5 (10) | 5/5 (19) | 5/5 (13) | 5/5 (1) | 5/5 (1) | 5/5 (5) |

Tokens per task (the agent's `subagent_tokens`, copied from its
completion notification):

| task | A run 1 | A run 2 | A run 3 | B run 1 | B run 2 | B run 3 |
|---|---|---|---|---|---|---|
| 1 | 107,986 | 79,603 | 93,516 | 57,095 | 58,110 | 56,902 |
| 2 | 80,527 | 94,230 | 103,962 | 55,826 | 58,542 | 61,977 |
| 3 | 224,238 | 180,075 | 224,935 | 112,214 | 69,770 | 58,202 |
| 4 | 114,597 | 128,435 | 118,406 | 61,714 | 56,331 | 61,337 |

Did-not-compile: 0 of 12 on arm A (every file built and was proven
under `--instances 1`); 0 of 12 modules failed to load on arm B.
Calibration: every agent claimed every property of its task; the only
over-claims are the four task-3 race properties the harness did not
establish (three on arm A, one on arm B), listed below.

### The one property arm A never established, and why

`vp3-race-at-most-one` on arm A is read off the agent's dist world:
complete `holds` on `never-over-cap` and `one-landed` with
`a-reply-landed` covered. All three arm-A runs reached complete
`holds` on both invariants and none had the coverage: each declared
`Sometimes("a-reply-landed", ...)` as a standalone binding, which the
extractor does not discover (it discovers `CoverageSpec([Sometimes(...)])`
by type, as `demos/service/forum.dist.bosatsu` does). All three agents
said so in their notes: "no printed docs for Yichus/Dist's shapes",
"every constructor's field order had to be reverse-engineered from
deliberate type-mismatch errors", "I could not find any wiring point
that accepts a Sometimes value". The brief said `Sometimes` and not
`CoverageSpec`; the tool has no `--help` that lists the Dist package's
types. That is VP-4 in the register: a brief and tool-doc gap, charged
here to arm A as the pre-registration requires, and fixed before
iteration 002 (the brief names `CoverageSpec`; `yichus dist
--check-only` lists coverage specs; the Dist package gets the guide
treatment the service framework has in `api_guide`).

Task 3 on arm A also cost the most tokens of any cell (180k–225k,
34–84 iterations): the three agents spent their runs reverse-engineering
the Dist DSL, and all three left `post_reply` unchanged, correctly
observing that the single-instance topology already serializes each
request's IO plan — the property the harness could not see for want
of one coverage line.

### The one property arm B did not establish

`vp3-race-at-most-one` on B run 1: the schedule walk was truncated at
the scored budget (4,000 schedules, 12 operations per lane) with no
failing schedule found; that adapter makes 35 store calls per request
path. A diagnostic rerun at 200,000 schedules and 40 operations per
lane exhausted the space with no failure. The scored number stands
(VP-3); iteration 002 raises the budget and reports the schedules each
deliverable needed.

### What the tokens went to

Arm A's runs read a 500-line starting file, learned the language's
surface by trial builds ("infix `+` does not parse; `add`/`sub` are
functions", "the export list must precede the definitions", "every
positional pattern match on `Thread` had to change"), and ran
`api verify` and `why` between edits: 4–19 tool invocations on tasks
1, 2, and 4 against 1–5 node runs on arm B. What the tools bought on
those tasks is visible in the notes and not in the score: agents
confirmed boundaries with `why` on seeded rows, and `api verify` caught
a missing `WritePerm("threads")` when a reply began to bump the
thread's activity (two of three task-4 runs report it). Arm B's agents
wrote their own stores and tests and, on task 3, their own
interleaving schedulers (two of three built a DFS driver; one said its
scheduler "silently truncated a schedule when it mis-detected a stall"
before it fixed it).

### Against the prediction

The prediction said accuracy(A) ≥ accuracy(B) on tasks 1, 2, and 4: it
held (15/15 on every run, both arms). It said task 3 was the open
question on both arms: on arm A it was a documentation gap, on arm B a
harness budget. It said a lift below 1 was "a real possibility and the
number this iteration exists to get": it is 0.44–0.54.

### Register updates

- VP-3 (budget) and VP-4 (the Dist DSL has no printed docs, and the
  brief named `Sometimes` without `CoverageSpec`) opened; VP-2 stands.
- The agent-era plan's tranche 4 gate is failed at iteration 001; the
  evidence page (item 18) publishes the fail.
