Skip to content
nano-gpt-data-curation

claude-opus-5

nano-gpt-data-curation — graded by the task's verifier against a measured baseline.

backend
modal
variant
claude claude-opus-5
metric
held-out multi-domain perplexity
actions
78
duration
1h 10m
recorded
2026-08-04

run iddclm-nano-curation__xrPbCYK

Score

reward
306.650
compare against every run on this task →

Summary

Done. All three deliverables are final and verified end-to-end.

Result: dev perplexity 299.568 vs 457.8 for a random selection of the same budget (−34.6%), confirmed just now through the frozen pack_selection.py + train_nano.py.

Deliverables

  • /workspace/submission/selection.json — 32,549 unique in-pool ids, 30.0M tokens (2.50× budget). The official packer consumes a 12,000,000-token prefix.
  • /workspace/submission/curate.py — four cached stages from pool.jsonl, driven entirely by constants; no document id appears anywhere.
  • /workspace/claim.md — Hypothesis / Mechanism / Falsification / Transfer.

Reproducibility, verified rather than asserted. I closed two gaps that would have made the submission unreproducible from a clean checkout: surf.npz and markup.npz were built by ad-hoc inline code, now a proper stage (score_surface.py); and the candidate shortlist omitted the seeded random pool slice I had actually used, which matters because the token-level background distribution is estimated from the shortlist. A full from-scratch rebuild (1m55s) now reproduces all four caches byte-identically and emits the same 32,549 ids in the same order, and the resulting packed array is byte-identical to the one that scored 299.568.

Corrections to claim.md. The per-register figures had been written from a superseded selection; the submitted one measures wiki 597.3, news 270.6, web 289.1, tech 171.7 (all 299.6).

Run-to-run variance, newly measured. Retraining the identical selection at seeds 1338/1339 gives 300.9/300.9 against 299.568 at the frozen seed 1337 — a spread of 1.4 ppl. This changes what I can honestly claim: the coarse steps (−97.6, −31.3, −11.7, −14.8) are far outside noise, but the final dose step (300.0 → 299.6) is not resolved by a single run. A paired re-run at seed 1338 favours the submitted dose again (300.9 vs 302.6), so the ordering survives, but only the first ~155 ppl of the improvement is established; the last ~1 ppl is tuning inside noise. I also softened a Transfer-section claim that the dose-response "had not saturated" — the per-step gains were flattening (−14.8, −2.4, −0.4) when the pool's supply of pure spaced-punctuation prose ran out.

The load-bearing scientific result is unchanged: the encyclopedic register's gap is surface form, not topic. 0.28M tokens of spaced-punctuation prose lowers perplexity by 15 points, while the same 0.28M tokens of ordinary gate-passing prose raises it by 10 — same document count, same budget, so the effect cannot be "more data" or "better generic quality."

Transcript