gpt-5.6-sol
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- held-out multi-domain perplexity
- actions
- 138
- duration
- 47m
- recorded
- 2026-08-04
run iddclm-nano-curation__W6EguaR
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗The revised prefix now looks materially cleaner: factual explainers and reference prose, conventional reporting, substantial web essays, and math/programming explanations, with no giant PDF corruption or foreign pages in the audited top set. I’m packing the exact 12M-token prefix and running the frozen dev evaluation now.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({\"cmd\":\"bash dev_eval.sh submission/selection.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(r.output);\nif (r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=2771encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":2771,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":5000});text(rest);text(output)\n"
}exec result
Script running with cell ID 75
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "75",
"yield_time_ms": 30000,
"max_tokens": 5000
}wait result
Script completed
Wall time 15.1 seconds
Output:
{"chunk_id":"0d0b00","wall_time_seconds":30.001275513,"session_id":2771,"original_token_count":46}Token indices sequence length is longer than the specified maximum sequence length for this model (2040 > 1024). Running this sequence through the model will result in indexing errors
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":2771,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":5000});text(rest);text(output)\n"
}exec result
Script running with cell ID 76
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "76",
"yield_time_ms": 30000,
"max_tokens": 5000
}wait result
Script completed
Wall time 16.4 seconds
Output:
{"chunk_id":"4f30d2","wall_time_seconds":30.002049245,"session_id":2771,"original_token_count":15}packed 12000000 tokens from 13476 ids -> /tmp/dev_train.npy
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":2771,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":5000});text(rest);text(output)\n"
}exec result
Script running with cell ID 77
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "77",
"yield_time_ms": 30000,
"max_tokens": 5000
}wait result
Script completed
Wall time 9.4 seconds
Output:
{"chunk_id":"453cfe","wall_time_seconds":27.151608069,"exit_code":0,"original_token_count":38}HELDOUT_PPL 446.268 loss 6.1009 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({\"cmd\":\"python - <<'PY'\\nimport random,json\\nx=list(range(182016));random.Random(1337).shuffle(x)\\njson.dump(x[:30000],open('/tmp/random.json','w'))\\nPY\\nbash dev_eval.sh /tmp/random.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":3000});\ntext(r.output);\nif (r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=80792encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":80792,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":3000});text(rest);text(output)\n"
}exec result
Script running with cell ID 79
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "79",
"yield_time_ms": 30000,
"max_tokens": 3000
}wait result
Script completed
Wall time 7.6 seconds
Output:
{"chunk_id":"f43406","wall_time_seconds":30.002432753,"session_id":80792,"original_token_count":46}Token indices sequence length is longer than the specified maximum sequence length for this model (2126 > 1024). Running this sequence through the model will result in indexing errors
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":80792,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":3000});text(rest);text(output)\n"
}exec result
Script running with cell ID 80
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "80",
"yield_time_ms": 30000,
"max_tokens": 3000
}wait result
Script completed
Wall time 13.8 seconds
Output:
{"chunk_id":"48d9d7","wall_time_seconds":30.002187665,"session_id":80792,"original_token_count":15}packed 12000000 tokens from 30000 ids -> /tmp/dev_train.npy
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let{output,...rest}=await tools.write_stdin({\"session_id\":80792,\"yield_time_ms\":30000,\"chars\":\"\",\"max_output_tokens\":3000});text(rest);text(output)\n"
}exec result
Script running with cell ID 81
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "81",
"yield_time_ms": 30000,
"max_tokens": 3000
}wait result
Script completed
Wall time 2.5 seconds
Output:
{"chunk_id":"8f15ed","wall_time_seconds":27.081556009,"exit_code":0,"original_token_count":38}HELDOUT_PPL 476.714 loss 6.1669 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
The curated run reached dev perplexity 446.27 versus 476.71 for a deterministic random 12M-token control—about a 6.4% reduction under the identical frozen recipe. I’m finalizing the falsifiable claim and doing structural validity/reproducibility checks on the deliverables.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/curate.py\\n@@\\n tok=GPT2TokenizerFast.from_pretrained('gpt2',local_files_only=True)\\n+ tok.model_max_length=10**9 # documents are encoded, never passed to GPT-2\\n@@\\n- # GPT-2 English token count is close enough for balancing here; the\\n- # published list deliberately has a 35% coverage margin.\\n+ # Exact GPT-2 count is used for prefix balancing; the published list\\n+ # deliberately has a 35% coverage margin.\\n*** Add File: /workspace/claim.md\\n+# Hypothesis\\n+\\n+At a severe token bottleneck, a token-balanced mixture of clean, substantive English documents representing encyclopedia-like exposition, high-quality general prose, reporting, and technical explanation will train a better broad language model than random web data. Hard rejection of corruption, non-English text, tiny pages, giant threads, and templated boilerplate should matter more than maximizing raw document count.\\n+\\n+# Mechanism\\n+\\n+The filter raises the fraction of sentence-bearing English prose and lowers navigation, repeated-line, URL/policy, mojibake, and first-person-chatter rates. Positive lexical/structural evidence assigns eligible documents to four disclosed registers, and weighted round-robin ordering keeps each near 25% of the consumed token prefix. An observable prediction independent of final perplexity is that the selected 12M-token prefix has fewer repeated/short lines and corrupted characters than a random prefix, while no register differs from 25% token share by more than one boundary document. In the provided diagnostic, the resulting dev perplexity was 446.27 versus 476.71 for a seed-1337 random control under the same frozen training recipe.\\n+\\n+# Falsification\\n+\\n+Reject the hypothesis if (a) a random 12M-token sample does not have worse held-out loss across repeated training seeds, (b) ablating the cleanliness gates and balance constraint does not worsen loss, or (c) the predicted corpus statistics above fail on the actually consumed prefix. A particularly useful test is four-way evaluation: improvement confined to one target quarter while the others regress would falsify the broad-mixture mechanism even if aggregate noise produces a small win.\\n+\\n+# Transfer\\n+\\n+The criterion uses only document text, aggregate frequencies from the disclosed development distribution, and generic web-quality signals—never hand-picked ids—so it transfers to a newly crawled pool with the same target registers. For a different language or target mixture, replace the English function-word test and vocabulary reference, then set round-robin shares to the new evaluation mixture; the corruption, repetition, size-diversity, and prefix-balancing principles remain unchanged.\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\n"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({\"cmd\":\"chmod +x submission/curate.py && python -m py_compile submission/curate.py && python - <<'PY'\\nimport json\\np='submission/selection.json'; x=json.load(open(p))\\nassert isinstance(x,list) and x\\nassert len(x)==len(set(x))\\nassert all(type(i) is int and 0<=i<182016 for i in x)\\nprint('valid ids:',len(x),'unique:',len(set(x)),'range:',min(x),max(x))\\nPY\\nls -lh submission/selection.json submission/curate.py claim.md\\ncat /tmp/dev.json\\ncat .timer/remaining_secs\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.2 seconds
Output:
valid ids: 13476 unique: 13476 range: 39 181924
-rw-r--r-- 1 root root 2.2K Aug 4 01:50 claim.md
-rwxr-xr-x 1 root root 6.4K Aug 4 01:50 submission/curate.py
-rw-r--r-- 1 root root 93K Aug 4 01:46 submission/selection.json
{"train_npy": "/tmp/dev_train.npy", "heldout_ppl": 476.7139227949391, "mean_loss": 6.1669165684025655, "max_iters": 3000, "seed": 1337, "train_s": 56.2}2957
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({\"cmd\":\"python - <<'PY'\\nimport json,collections\\nx=json.load(open('submission/selection.json'))\\nc=collections.Counter(min(9,i*10//182016) for i in x)\\nprint(c)\\nPY\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":1000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
Counter({2: 2156, 4: 2137, 0: 2107, 5: 2095, 1: 2093, 3: 2072, 6: 678, 9: 52, 7: 43, 8: 43})
Completed all deliverables:
- selection.json: 13,476 unique valid IDs, covering over 12M tokens.
- curate.py: reproducible quality filtering and domain-balancing pipeline.
- claim.md: hypothesis, mechanism, falsification, and transfer.
Frozen dev evaluation:
- Curated: 446.27 perplexity
- Random control: 476.71 perplexity
- Improvement: 6.4% lower perplexity.
The required artifacts were already finalized; I’m only rechecking the authoritative timer and exact paths after the interruption.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({\"cmd\":\"cat .timer/remaining_secs && test -s submission/selection.json && test -s submission/curate.py && test -s claim.md && python - <<'PY'\\nimport json\\nx=json.load(open('/workspace/submission/selection.json'))\\nassert len(x)==len(set(x)) and all(type(i) is int and 0<=i<182016 for i in x)\\nprint(len(x), 'valid unique ids')\\nPY\\nls -lh /workspace/submission/selection.json /workspace/submission/curate.py /workspace/claim.md\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":2000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.2 seconds
Output:
2887
13476 valid unique ids
-rw-r--r-- 1 root root 2.2K Aug 4 01:50 /workspace/claim.md
-rwxr-xr-x 1 root root 6.4K Aug 4 01:50 /workspace/submission/curate.py
-rw-r--r-- 1 root root 93K Aug 4 01:46 /workspace/submission/selection.json
Rechecked successfully. All required deliverables remain finalized at the exact paths, with 13,476 valid unique IDs.