gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 195
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "195",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 27.0 seconds
Output:
Processed prompts: 43%|▍| 36/84 [14:28<14:27, 18.08s/it, est. speed input: 4.5
Processed prompts: 88%|▉| 84/96 [14:52<01:35, 7.99s/it, est. speed input: 11.
Processed prompts: 38%|▍| 36/96 [15:15<26:08, 26.14s/it, est. speed input: 5.1encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 196
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "196",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 23.5 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 197
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "197",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 21.4 seconds
Output:
Processed prompts: 57%|▌| 48/84 [16:27<08:54, 14.85s/it, est. speed input: 5.1
Processed prompts: 100%|█| 96/96 [16:36<00:00, 8.20s/it, est. speed input: 15.
Processed prompts: 100%|█| 96/96 [16:36<00:00, 8.20s/it, est. speed input: 15.
Processed prompts: 100%|█| 96/96 [16:36<00:00, 10.38s/it, est. speed input: 15.
Processed prompts: 100%|█| 84/84 [16:53<00:00, 12.44s/it, est. speed input: 8.6
Processed prompts: 100%|█| 84/84 [16:53<00:00, 12.44s/it, est. speed input: 8.6
Processed prompts: 100%|█| 84/84 [16:53<00:00, 12.06s/it, est. speed input: 8.6
Processed prompts: 71%|▋| 60/84 [16:54<04:07, 10.30s/it, est. speed input: 6.0encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 198
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "198",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 34.1 seconds
Output:
Processed prompts: 50%|▌| 48/96 [17:38<16:24, 20.52s/it, est. speed input: 6.4encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 199
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "199",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 33.7 seconds
Output:
Processed prompts: 62%|▋| 60/96 [18:50<09:10, 15.28s/it, est. speed input: 8.1
Processed prompts: 86%|▊| 72/84 [18:52<02:01, 10.15s/it, est. speed input: 7.7
Processed prompts: 100%|█| 84/84 [18:56<00:00, 6.93s/it, est. speed input: 8.7
Processed prompts: 100%|█| 84/84 [18:56<00:00, 6.93s/it, est. speed input: 8.7
Processed prompts: 100%|█| 84/84 [18:56<00:00, 13.53s/it, est. speed input: 8.7
Processed prompts: 75%|▊| 72/96 [19:05<04:11, 10.50s/it, est. speed input: 9.2encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:400});\ntext(r.output);"
}exec result
Script running with cell ID 200
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "200",
"yield_time_ms": 50000,
"max_tokens": 400
}wait result
Script completed
Wall time 26.8 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:400});\ntext(r.output);"
}exec result
Script running with cell ID 201
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "201",
"yield_time_ms": 50000,
"max_tokens": 400
}wait result
Script completed
Wall time 24.2 seconds
Output:
Processed prompts: 88%|▉| 84/96 [21:08<02:05, 10.44s/it, est. speed input: 9.1encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}exec result
Script running with cell ID 202
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "202",
"yield_time_ms": 50000,
"max_tokens": 500
}wait result
Script completed
Wall time 11.9 seconds
Output:
Processed prompts: 100%|█| 96/96 [21:55<00:00, 8.35s/it, est. speed input: 9.9
Processed prompts: 100%|█| 96/96 [21:55<00:00, 8.35s/it, est. speed input: 9.9
Processed prompts: 100%|█| 96/96 [21:55<00:00, 13.70s/it, est. speed input: 9.9
FINAL avg@12 accuracy: 54.44% (196/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/frontier_v2_rep2_full.json
v2 repeat scored 54.44%, giving a 55.00 two-run mean versus v1’s 55.14; the uniform full-response change does not help, so I’m reverting it. I have time for one higher-leverage curriculum test: upweight direct-answer AoPS problems to roughly balance them with the broader olympiad pool. This preserves the verified early-weighted OPSD loss and changes only which allowed dataset examples receive the fixed 3,200-example budget.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/repo/opsd_trainer.py\\n@@\\n- # The correctness anchor covers the whole successful trajectory,\\n- # including late final-answer and boxing tokens. Unlike the dense\\n- # teacher KL, it should not discount those late tokens.\\n- correct_weights = correct_mask.to(token_jsd.dtype)\\n+ correct_weights = correct_mask.to(token_jsd.dtype) * position_weights.unsqueeze(0)\\n*** Update File: /workspace/repo/opsd_train.py\\n@@\\n- frontier_indices = [\\n+ olympiad_indices = [\\n i\\n for i, (source, answer) in enumerate(zip(sources, answers))\\n- if source in {\\\"olympiads\\\", \\\"aops_forum\\\"} and is_aime_style_answer(answer)\\n+ if source == \\\"olympiads\\\" and is_aime_style_answer(answer)\\n ]\\n+ aops_indices = [\\n+ i\\n+ for i, (source, answer) in enumerate(zip(sources, answers))\\n+ if source == \\\"aops_forum\\\" and is_aime_style_answer(answer)\\n+ ]\\n+ # AoPS direct-answer problems are the closest in-domain frontier examples but\\n+ # only 4% of the raw corpus. Repeating their indices gives them about 44% of\\n+ # sampling mass; each visit still gets a fresh on-policy rollout.\\n+ frontier_indices = olympiad_indices + aops_indices * 5\\n train_dataset = train_dataset.select(frontier_indices)\\n- print(f\\\"Verified frontier curriculum: selected {len(train_dataset)} / {len(dataset['train'])} examples\\\")\\n+ print(\\n+ f\\\"Verified frontier curriculum: {len(olympiad_indices)} olympiad + \\\"\\n+ f\\\"5x{len(aops_indices)} AoPS views = {len(train_dataset)} examples\\\"\\n+ )\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"python -m py_compile opsd_train.py opsd_trainer.py && git diff --check -- opsd_train.py opsd_trainer.py data_collator.py\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}exec result
Script completed
Wall time 0.2 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/frontier_v3_aops PORT=12954 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":350,\"tty\":true});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"f5ee9b","wall_time_seconds":1.002237377,"session_id":42659,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 205
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "205",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 32.6 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 12
Warning: truncated output (original token count: 8890)
Total output lines: 94
<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.
================================================================================
RUN CONFIGURATION
================================================================================
WandB Run Name: capped_seed42_lr5e-6_bs32
Output Directory: /workspace/runs/frontier_v3_aops/capped_seed42
===…27 tokens truncated…00:56, 144.32 examples/s]
Tokenizing train dataset: 40%|████ | 5502/13662 [00:36<00:55, 148.19 examples/s]
Tokenizing train dataset: 40%|████ | 5518/13662 [00:36<00:54, 148.47 examples/s]
Tokenizing train dataset: 41%|████ | 5537/13662 [00:37<00:51, 157.08 examples/s]
Tokenizing train dataset: 41%|████ | 5555/13662 [00:37<00:50, 161.13 examples/s]
Tokenizing train dataset: 41%|████ | 5577/13662 [00:37<00:52, 154.15 examples/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 206
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "206",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 31.0 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 8
Warning: truncated output (original token count: 12871)
Total output lines: 60
Tokenizing train dataset: 41%|████ | 5596/13662 [00:37<00:49, 162.11 examples/s]
Tokenizing train dataset: 41%|████ | 5620/13662 [00:37<00:51, 156.05 examples/s]
Tokenizing train dataset: 41%|████▏ | 5637/13662 [00:37<00:50, 159.02 examples/s]
Tokenizing train dataset: 41%|…27 tokens truncated…ected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
Loading safetensors checkpoint shards: 0% Completed | 0/2 [00:00<?, ?it/s]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 207
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "207",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 25.5 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 1853)
Total output lines: 54
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00, 2.21it/s]
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00, 2.21it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/5 [00:00<?, ?it/s]
Capturing…27 tokens truncated…: 765.4 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 2981, avg length: 745.2, speed: 814.1 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3634, avg length: 908.5, speed: 976.2 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3191, avg length: 797.8, speed: 862.9 tok/s