# Iteration 009 — the reflection attack: the down-tier security win, finally measured

Three security iterations ended with the same honest confession: the
checker's win was real but **unmeasured**, because the bug shipped was
reading-legible. it4 (replay) and it8 (cleartext) had one-line flaws a careful
reader spots; it6 (forgery) was legible once the threat model was stated. Each
deferred the discriminating down-tier measurement "to protocol scale." This
iteration builds that task — a **reflection attack** on a mutual-authentication
handshake — and runs the probe. The delta materialized, cleanly.

## What shipped — task 010, reflection-auth (committed `0f1d8cd`)

A two-party mutual-auth handshake over direction-tagged symbolic MACs
(register DD26): `Mac(0, ·)` is the initiator's proof, `Mac(1, ·)` the
responder's. The server is the responder under test. The reflection
`AdversaryNode` (register DD24) MITMs all client↔server traffic, opens its own
session 2, and can **reflect** the server's own challenge nonce into a session
3 — using the server as a MAC oracle for a value the keyless attacker cannot
compute.

| variant | server answers / accepts | `mutual-authentication` |
|---|---|---|
| `reference/solution.bosatsu` | answers `Mac(1,·)`, accepts `Mac(0,·)` | **holds** (complete, 13641 schedules) + covered |
| `reference/broken-reflectable.bosatsu` | answers `Mac(0,·)`, accepts `Mac(0,·)` | **violated** (reflection oracle) |

The broken variant's counterexample is textbook: the attacker opens session 3
challenging with the server's own nonce (10), the server answers `Mac(0/10)`,
and the attacker replays it to finish session 2 — `server=srv(ok=[2,])` while
`client=client(fin=[])`. The bug is a single shared direction tag; it exists
**only across two interleaved sessions**. Fault budget 0 keeps the secure
reference inside an honest exhaustive cap; JVM and Node engines agree
verdict + 13641-count (`dist_differential_smoke.sh` task-010 block).

## The measurement — the down-tier delta, at last

A win-condition probe, three cells, all **haiku**, one trial each. Cell A is
the fair baseline the it5/it6 discipline demands: an active-attacker threat
model is stated (that context is realistic and unavoidable for an auth
review), but the **reflection attack itself is withheld** — the agent must
discover the cross-session oracle on its own.

| cell | arm | input | verdict | truth | tokens |
|---|---|---|---|---|---|
| **A** | baseline, no tool | protocol + **reflectable** server source | **SECURE** | vulnerable | 19,969 |
| **B** | checker | `yichus dist` verdict + counterexample trace | **VULNERABLE** | vulnerable | 19,877 |
| **C** | baseline control | protocol + **secure** server source | **SECURE** | secure | 19,971 |

**The result.** The source-reading baseline returned **SECURE on both**
servers — it is *blind to the reflection defect*, not trigger-happy (Cell C
correctly clears the secure server, citing the direction-tag binding). So it
**false-certified the reflectable server** (Cell A). The checker arm, handed
the counterexample, caught it and reconstructed the cross-session oracle in
its own words (Cell B). **False-certification rate: baseline 1/1 defective,
checker arm 0/1** — the first clean down-tier security false certification the
loop has measured, closing the it4/it6/it8 deferral.

**The killer detail.** Cell A *noticed the anomaly* — "the server sends
`Mac(0, n)` (initiator's tag) instead of `Mac(1, n)`" — and then reasoned
itself back out of it: "However, this doesn't create a security
vulnerability." It enumerated where the attacker could obtain `Mac(0, pn)` and
concluded the only source was the honest client, **never considering that the
attacker can open a fresh session whose challenge nonce is `pn` and let the
server compute it**. The smoking gun was in its hands and it cleared the
server anyway. That is what reading-resistant means, and it is exactly the
regime iterations 4/6/8 could not reach with a legible bug.

**On tokens — flat, and that is the point.** All three cells cost ≈19.9k
tokens; there is no token-lift story here. The win is **accuracy**, precisely
the it6 finding at protocol scale: the checker's value is exhaustive
certainty over a move space beyond a reader's reach, not cheapness. The
numbers above are the agents' real reported `subagent_tokens` (captured from
the run notifications, per DD4), not eyeballed. One trial per cell, haiku
tier — a single measured false certification, not a rate over a corpus;
scaling it into the dist-verdict eval-gate corpus is the natural follow-up.

## Regression guard

`DistBenchTaskTest` ×2 (23 total green): the direction-bound server holds +
covered over the complete 13641-schedule run; the untagged server is violated
with an adversary-move counterexample whose end state has the forged session
(2) in `ok_sessions`. `dist_differential_smoke.sh` task-010 block: JVM/Node
verdict + count parity (13641) on the secure reference, both engines gate the
reflectable variant by exit code. Full smoke: **PASS**.

## Register deltas

- **DD21 (trap legibility)** gains its strongest data point: at protocol scale,
  a reading-resistant defect false-certified a cheap reader *that had already
  spotted the anomaly*; the checker arm caught it. The down-tier security win
  the register had marked "gated on scale" is now measured.

## Budget

Probe spend: **59,817 tokens** (three haiku cells, 19,969 + 19,877 + 19,971).
Main-agent work (task files, tests, differential block, this record) is not
probe spend.

## Next iteration (proposed / sequenced backlog)

- **DD25 general half — noninterference via self-composition** (`--self-compose`
  product mode): the largest remaining *capability* gap and the only near-term
  candidate needing real engine work; now has a protocol-scale world to be
  measured against.
- **Scale this into the eval-gate**: fold the reflection pair into the
  `dist-verdict` corpus (it7) so the false certification becomes a standing,
  scored regression rather than a single haiku trial.
- **DD7 — generator reproducibility**: codify the adversarial-mutant protocol
  as a committed generator-brief template.
