gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 208
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "208",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 18.2 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 2468)
Total output lines: 156
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3558, avg length: 889.5, speed: 969.8 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2701, avg length: 675.2, speed: 735.6 tok/s
vLLM generation done - elapsed time: 3.75s, prompt…27 tokens truncated…910.4 tok/s
vLLM generation done - elapsed time: 3.60s, prompts: 4, total tokens: 2752, avg length: 688.0, speed: 764.3 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3385, avg length: 846.2, speed: 908.8 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4022, avg length: 1005.5, speed: 1067.8 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 209
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "209",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 20.8 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 3129)
Total output lines: 185
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3346, avg length: 836.5, speed: 910.5 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3193, avg length: 798.2, speed: 865.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompt…27 tokens truncated…0s, prompts: 4, total tokens: 3290, avg length: 822.5, speed: 888.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3288, avg length: 822.0, speed: 887.3 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1085.3 tok/s
17%|█▋ | 17/100 [02:32<12:22, 8.95s/it]The AoPS-balanced curriculum is training normally. It contains 7,612 direct-answer olympiad views plus five passes over 1,210 direct-answer AoPS examples, producing roughly 44% AoPS sampling mass while preserving fresh on-policy rollouts on every revisit.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 210
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "210",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 33.8 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 9
Warning: truncated output (original token count: 2763)
Total output lines: 205
vLLM generation done - elapsed time: 2.57s, prompts: 4, total tokens: 2333, avg length: 583.2, speed: 906.2 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3004, avg length: 751.0, speed: 815.2 tok/s
vLLM generation done - elapsed time: 3.73s, prompt…27 tokens truncated…ens: 4096, avg length: 1024.0, speed: 1090.5 tok/s
24%|██▍ | 24/100 [03:34<11:21, 8.97s/it]
{'loss': 0.0038, 'grad_norm': 0.07585375756025314, 'learning_rate': 3.85e-06, 'on_policy_loss': 0.0038, 'rollout_accuracy': 0.3125, 'epoch': 0.06}
24%|██▍ | 24/100 [03:35<11:21, 8.97s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 211
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "211",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 32.6 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 1851)
Total output lines: 61
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3367, avg length: 841.8, speed: 901.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3545, avg length: 886.2, speed: 947.7 tok/s
vLLM generation done - elapsed time: 3.73s, prompts…27 tokens truncated… 779.4 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3450, avg length: 862.5, speed: 928.4 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3516, avg length: 879.0, speed: 946.2 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3801, avg length: 950.2, speed: 1020.7 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 212
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "212",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 28.4 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 2060)
Total output lines: 71
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3469, avg length: 867.2, speed: 932.4 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.6 tok/s
vLLM generation done - elapsed time: 3.74s, promp…27 tokens truncated…031.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3908, avg length: 977.0, speed: 1045.1 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3885, avg length: 971.2, speed: 1042.2 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3691, avg length: 922.8, speed: 986.3 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 213
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "213",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 33.7 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 1875)
Total output lines: 62
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3524, avg length: 881.0, speed: 949.2 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1099.1 tok/s
vLLM generation done - elapsed time: 3.75s, promp…27 tokens truncated…rompts: 4, total tokens: 3221, avg length: 805.2, speed: 877.2 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3095, avg length: 773.8, speed: 833.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3852, avg length: 963.0, speed: 1029.7 tok/s
44%|████▍ | 44/100 [06:33<08:22, 8.97s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 214
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "214",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 31.1 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 8
Warning: truncated output (original token count: 2049)
Total output lines: 66
{'loss': 0.0001, 'grad_norm': 0.05258685722947121, 'learning_rate': 2.85e-06, 'on_policy_loss': 0.0001, 'rollout_accuracy': 0.1875, 'epoch': 0.1}
44%|████▍ | 44/100 [06:33<08:22, 8.97s/it]vLLM generation done - el…27 tokens truncated… 900.6 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3690, avg length: 922.5, speed: 983.7 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3484, avg length: 871.0, speed: 928.9 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3853, avg length: 963.2, speed: 1022.6 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 215
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "215",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 30.7 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 13
Warning: truncated output (original token count: 2833)
Total output lines: 264
================================================================================
Saved 40 generation outputs to:
/workspace/runs/frontier_v3_aops/capped_seed42/generations/generations_step_50.json
============================================================================…27 tokens truncated…6.2 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 2774, avg length: 693.5, speed: 753.5 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1099.8 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1087.2 tok/s
The AoPS-balanced run is past halfway and stable. Its rollout correctness remains in the intended frontier range rather than saturating, so the added AoPS mass is supplying nontrivial privileged-teacher signal rather than merely repeating easy solved items.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 216
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "216",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 31.5 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 2066)
Total output lines: 66
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1093.6 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3989, avg length: 997.2, speed: 1063.8 tok/s
vLLM generation done - elapsed time: 3.77s, prom…27 tokens truncated…3.3 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1096.5 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3736, avg length: 934.0, speed: 991.3 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1085.1 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 217
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "217",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 30.7 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 14
Warning: truncated output (original token count: 1966)
Total output lines: 68
vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3052, avg length: 763.0, speed: 838.1 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3524, avg length: 881.0, speed: 949.5 tok/s
vLLM generation done - elapsed time: 3.73s, prompts…27 tokens truncated…9 tok/s
================================================================================
Saved 40 generation outputs to:
/workspace/runs/frontier_v3_aops/capped_seed42/generations/generations_step_70.json
================================================================================
71%|███████ | 71/100 [10:35<04:19, 8.94s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 218
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "218",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 33.2 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 2194)
Total output lines: 70
vLLM generation done - elapsed time: 3.55s, prompts: 4, total tokens: 2417, avg length: 604.2, speed: 681.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3545, avg length: 886.2, speed: 954.6 tok/s
vLLM generation done - elapsed time: 3.72s, prompts…27 tokens truncated…2.9 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3996, avg length: 999.0, speed: 1069.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3886, avg length: 971.5, speed: 1039.2 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1086.2 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:42659,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 219
Wall time 11.0 seconds
Output: