claude-opus-5
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- held-out multi-domain perplexity
- actions
- 78
- duration
- 1h 10m
- recorded
- 2026-08-04
run iddclm-nano-curation__xrPbCYK
Score
Summary
Done. All three deliverables are final and verified end-to-end.
Result: dev perplexity 299.568 vs 457.8 for a random selection of the same budget (−34.6%), confirmed just now through the frozen pack_selection.py + train_nano.py.
Deliverables
/workspace/submission/selection.json— 32,549 unique in-pool ids, 30.0M tokens (2.50× budget). The official packer consumes a 12,000,000-token prefix./workspace/submission/curate.py— four cached stages frompool.jsonl, driven entirely by constants; no document id appears anywhere./workspace/claim.md— Hypothesis / Mechanism / Falsification / Transfer.
Reproducibility, verified rather than asserted. I closed two gaps that would have made the submission unreproducible from a clean checkout: surf.npz and markup.npz were built by ad-hoc inline code, now a proper stage (score_surface.py); and the candidate shortlist omitted the seeded random pool slice I had actually used, which matters because the token-level background distribution is estimated from the shortlist. A full from-scratch rebuild (1m55s) now reproduces all four caches byte-identically and emits the same 32,549 ids in the same order, and the resulting packed array is byte-identical to the one that scored 299.568.
Corrections to claim.md. The per-register figures had been written from a superseded selection; the submitted one measures wiki 597.3, news 270.6, web 289.1, tech 171.7 (all 299.6).
Run-to-run variance, newly measured. Retraining the identical selection at seeds 1338/1339 gives 300.9/300.9 against 299.568 at the frozen seed 1337 — a spread of 1.4 ppl. This changes what I can honestly claim: the coarse steps (−97.6, −31.3, −11.7, −14.8) are far outside noise, but the final dose step (300.0 → 299.6) is not resolved by a single run. A paired re-run at seed 1338 favours the submitted dose again (300.9 vs 302.6), so the ordering survives, but only the first ~155 ppl of the improvement is established; the last ~1 ppl is tuning inside noise. I also softened a Transfer-section claim that the dose-response "had not saturated" — the per-step gains were flattening (−14.8, −2.4, −0.4) when the pool's supply of pure spaced-punctuation prose ran out.
The load-bearing scientific result is unchanged: the encyclopedic register's gap is surface form, not topic. 0.28M tokens of spaced-punctuation prose lowers perplexity by 15 points, while the same 0.28M tokens of ordinary gate-passing prose raises it by 10 — same document count, same budget, so the effect cannot be "more data" or "better generic quality."