gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 627
- duration
- 6h 30m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__RqGcDty
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 297
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "297",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 976)
Total output lines: 31
vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 5224, avg le…26 tokens truncated…tion done - elapsed time: 5.06s, prompts: 4, total tokens: 5260, avg length: 1315.0, speed: 1040.4 tok/s
25%|██▌ | 25/100 [04:54<14:45, 11.81s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 298
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "298",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 10.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1165)
Total output lines: 41
vLLM generation done - elapsed time: 3.44s, prompts: 4, total tokens: 3148, avg l…27 tokens truncated…tion done - elapsed time: 5.14s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1194.5 tok/s
29%|██▉ | 29/100 [05:42<13:58, 11.81s/it]The third run is near step 30 and continues to mirror the first two. Once it finishes, one final full evaluation will give a three-run local mean to compare with the official five-run baseline.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 299
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "299",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 3.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1169)
Total output lines: 41
vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4360, avg l…27 tokens truncated…on done - elapsed time: 5.11s, prompts: 4, total tokens: 5754, avg length: 1438.5, speed: 1126.1 tok/s
33%|███▎ | 33/100 [06:29<13:11, 11.81s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 300
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "300",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5
Warning: truncated output (original token count: 956)
Total output lines: 33
vLLM generation done - elapsed time: 4.70s, prompts: 4, total tokens: 4093, avg le…26 tokens truncated…█▌ | 36/100 [07:05<12:36, 11.82s/it]vLLM generation done - elapsed time: 4.82s, prompts: 4, total tokens: 3796, avg length: 949.0, speed: 788.1 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 301
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "301",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 2.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 884)
Total output lines: 28
vLLM generation done - elapsed time: 4.98s, prompts: 4, total tokens: 4688, avg le…26 tokens truncated…63, avg length: 1265.8, speed: 1006.6 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1193.5 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 302
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "302",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 9.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1036)
Total output lines: 36
vLLM generation done - elapsed time: 4.98s, prompts: 4, total tokens: 4609, avg l…27 tokens truncated…56, avg length: 1289.0, speed: 1015.6 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1191.9 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 303
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "303",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 10.7 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5
Warning: truncated output (original token count: 942)
Total output lines: 33
43%|████▎ | 43/100 [08:28<11:17, 11.88s/it]vLLM generation done - …26 tokens truncated…05659085512161255, 'learning_rate': 2.7500000000000004e-06, 'on_policy_loss': -0.0054, 'epoch': 0.4}
46%|████▌ | 46/100 [09:03<10:41, 11.88s/it]At step 46, the third run’s clipped objective (−0.0054) and gradient norm (0.057) remain in the same range as the validated runs. Training is on schedule.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 304
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "304",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 12.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 799)
Total output lines: 26
vLLM generation done - elapsed time: 5.00s, prompts: 4, total tokens: 4367, avg le…26 tokens truncated… done - elapsed time: 5.10s, prompts: 4, total tokens: 6031, avg length: 1507.8, speed: 1182.6 tok/s
49%|████▉ | 49/100 [09:39<10:06, 11.89s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 305
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "305",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 11.5 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1074)
Total output lines: 37
vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 4712, avg l…27 tokens truncated…144, avg length: 1536.0, speed: 1197.2 tok/s
vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4662, avg length: 1165.5, speed: 938.9 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 306
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "306",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 2.2 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 2516)
Total output lines: 170
vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5572, avg …27 tokens truncated…d_norm': 0.056033145636320114, 'learning_rate': 2.25e-06, 'on_policy_loss': -0.008, 'epoch': 0.49}
56%|█████▌ | 56/100 [11:03<08:46, 11.97s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 307
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "307",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 7.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 857)
Total output lines: 27
vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 5143, avg le…26 tokens truncated…4169, avg length: 1042.2, speed: 843.7 tok/s
vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4395, avg length: 1098.8, speed: 884.2 tok/s
The third replicate is past halfway and remains stable. No further method changes are planned; the remaining work is replication, final integrity checks, and concise handoff.