Six stages, computed from the ledger — never hand-maintained. * marks a run whose record was reconstructed afterwards from its log; those runs never stated a falsifier, so Proposed stays empty for all of them.
| run | Proposed | Trained | Diagnosed | Anchored | Ranked | Human-tested | iters | promos | verdict |
|---|---|---|---|---|---|---|---|---|---|
| v28 | ✓ | – | – | – | – | – | –/– | – | |
| v27 | ✓ | ✓ | ✓ | ✓ | ✓ | – | 40/40 | 1 | REGRESSED against its own baseline: -93 Elo [-117, -71]. |
| v25 * | – | ✓ | ✓ | ✓ | ✓ | – | 100/100 | 12 | RECONSTRUCTED — 100 iterations, 12 promotions, all gated on 20 games (41% false-pass). No compar |
| v24 * | – | ✓ | ✓ | ✓ | ✓ | – | 80/80 | 10 | RECONSTRUCTED — 80 iterations, 10 promotions, all gated on 20 games (41% false-pass). No compari |
| v23 * | – | ✓ | ✓ | – | – | – | 67/70 | 1 | RECONSTRUCTED — 67 iterations, 1 promotions, all gated on 20 games (41% false-pass). No comparis |
| v22 * | – | – | ✓ | – | ✓ | – | 80/110 | 8 | RECONSTRUCTED — incomplete (80/110 iterations). No baseline comparison was ever run. |
| v20 * | – | ✓ | ✓ | ✓ | ✓ | – | 54/60 | 5 | RECONSTRUCTED — 54 iterations, 5 promotions, all gated on 20 games (41% false-pass). No comparis |
| v20cont * | – | – | ✓ | ✓ | – | – | 26/80 | 2 | RECONSTRUCTED — incomplete (26/80 iterations). No baseline comparison was ever run. |
| v19 * | – | – | ✓ | – | – | – | 20/50 | 2 | RECONSTRUCTED — incomplete (20/50 iterations). No baseline comparison was ever run. |
| v18 * | – | – | ✓ | – | – | – | 38/50 | 3 | RECONSTRUCTED — incomplete (38/50 iterations). No baseline comparison was ever run. |
| v17d * | – | – | ✓ | – | – | – | 42/50 | 0 | RECONSTRUCTED — incomplete (42/50 iterations). No baseline comparison was ever run. |
| v17b * | – | – | ✓ | – | – | – | 1/50 | 0 | RECONSTRUCTED — incomplete (1/50 iterations). No baseline comparison was ever run. |
| v17 * | – | – | ✓ | – | – | – | 12/50 | 1 | RECONSTRUCTED — incomplete (12/50 iterations). No baseline comparison was ever run. |
| v17c * | – | – | ✓ | – | – | – | 16/50 | 2 | RECONSTRUCTED — incomplete (16/50 iterations). No baseline comparison was ever run. |
| v16 * | – | – | ✓ | – | – | – | 46/100 | 7 | RECONSTRUCTED — incomplete (46/100 iterations). No baseline comparison was ever run. |
| v15 * | – | – | ✓ | – | – | – | 31/50 | 1 | RECONSTRUCTED — incomplete (31/50 iterations). No baseline comparison was ever run. |
| v14 * | – | ✓ | ✓ | – | – | – | 49/50 | 8 | RECONSTRUCTED — 49 iterations, 8 promotions, all gated on 20 games (41% false-pass). No comparis |
| v13 * | – | – | ✓ | – | – | – | 33/50 | 5 | RECONSTRUCTED — incomplete (33/50 iterations). No baseline comparison was ever run. |
| v12 * | – | – | ✓ | – | – | – | 32/50 | 5 | RECONSTRUCTED — incomplete (32/50 iterations). No baseline comparison was ever run. |
Everything quoted against one fixed, deterministic anchor, so the rows are comparable. ab-d3 is reproducible forever; a checkpoint is not.
The first within-run curve measured against a fixed anchor — every historical eval used a moving opponent. ~820 Elo of real learning, then flat: 46 further iterations in v20cont added nothing.