nano-gpt-data-curation
claude-opus-5
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- held-out multi-domain perplexity
- actions
- 99
- duration
- 1h 3m
- recorded
- 2026-08-03
run iddclm-nano-curation__4gETduU
Score
reward
348.460
Summary
Submission is final and verified.
Deliverables
| file | state |
|---|---|
submission/selection.json | 22,747 ids, all in pool, no duplicates, priority-ordered; official pack_selection.py fills exactly 12,000,000 tokens from the first ~10k |
submission/curate.py | rerun just now to /tmp/rerun.json → byte-identical to the submitted file; criterion-driven end to end (features → target references → register classifier → token-balanced interleave) |
claim.md | hypothesis / mechanism / 6 non-perplexity observables / falsification list / 22-row ablation table / transfer recipe + a new "Reproducing" section pointing at the ablation scripts |
Score: HELDOUT_PPL 341.255 through the frozen trainer on the exact submitted file, versus 469.9 for random selection (−27%).
The last two knobs, now documented rather than adopted:
- ranker trained 3000 steps instead of 1200 → 340.5 (0.8 better than the submission, inside the ~15 PPL band that same-token-set seed retrains showed; not worth swapping a verified submission for)
--min_chars 1400→ 345.3 (worse)
Both are in the ablation table. Nothing outstanding.