gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "257",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 29.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 2982)
Total output lines: 192
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2918, avg length: 729.5, speed: 795.2 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3052, avg length: 763.0, speed: 824.1 tok/s
vLLM gener…27 tokens truncated…9.7 tok/s
24%|██▍ | 24/100 [03:33<11:15, 8.89s/it]
{'loss': -0.017, 'grad_norm': 0.053049515932798386, 'learning_rate': 3.85e-06, 'on_policy_loss': -0.017, 'rollout_accuracy': 0.1875, 'epoch': 0.09}
24%|██▍ | 24/100 [03:33<11:15, 8.89s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 258
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "258",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 29.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 1854)
Total output lines: 61
vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2876, avg length: 719.0, speed: 793.7 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3075, avg length: 768.8, speed: 830.0 tok/s
vLLM genera…27 tokens truncated…d time: 3.73s, prompts: 4, total tokens: 3153, avg length: 788.2, speed: 844.8 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4084, avg length: 1021.0, speed: 1080.2 tok/s
vLLM generation done - elapsed time: 3.80s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1078.5 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 259
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "259",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 20.1 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9
Warning: truncated output (original token count: 1945)
Total output lines: 68
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3156, avg length: 789.0, speed: 848.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3519, avg length: 879.8, speed: 939.7 tok/s
vLLM genera…27 tokens truncated…h: 770.5, speed: 833.1 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3807, avg length: 951.8, speed: 1007.6 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1081.6 tok/s
37%|███▋ | 37/100 [05:29<09:24, 8.96s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 260
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "260",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 30.1 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9
Warning: truncated output (original token count: 1741)
Total output lines: 58
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3539, avg length: 884.8, speed: 948.1 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3576, avg length: 894.0, speed: 954.4 tok/s
vLLM genera…27 tokens truncated…th: 770.0, speed: 836.4 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3719, avg length: 929.8, speed: 986.7 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3525, avg length: 881.2, speed: 934.2 tok/s
43%|████▎ | 43/100 [06:23<08:33, 9.01s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 261
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "261",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 32.3 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9
Warning: truncated output (original token count: 1997)
Total output lines: 66
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3625, avg length: 906.2, speed: 971.4 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3996, avg length: 999.0, speed: 1059.3 tok/s
vLLM gener…27 tokens truncated…h: 868.0, speed: 930.8 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3653, avg length: 913.2, speed: 972.7 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3888, avg length: 972.0, speed: 1028.4 tok/s
50%|█████ | 50/100 [07:26<07:29, 8.99s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 262
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "262",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 30.2 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 6
Warning: truncated output (original token count: 2942)
Total output lines: 206
{'loss': -0.0211, 'grad_norm': 0.04648179933428764, 'learning_rate': 2.55e-06, 'on_policy_loss': -0.0211, 'rollout_accuracy': 0.3125, 'epoch': 0.18}
50%|█████ | 50/100 [07:2…27 tokens truncated…56%|█████▌ | 56/100 [08:20<06:34, 8.96s/it]
{'loss': -0.0215, 'grad_norm': 0.055439431220293045, 'learning_rate': 2.25e-06, 'on_policy_loss': -0.0215, 'rollout_accuracy': 0.3125, 'epoch': 0.2}
56%|█████▌ | 56/100 [08:20<06:34, 8.96s/it]At halfway, stronger clipping remains stable but substantially changes the optimization regime (loss around −0.021 versus near zero for v1). This confirms the test is materially different; the full AIME score will decide whether that extra style suppression is beneficial or excessive.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 263
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "263",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 26.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 1867)
Total output lines: 61
vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2541, avg length: 635.2, speed: 702.2 tok/s
vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2588, avg length: 647.0, speed: 714.2 tok/s
vLLM genera…27 tokens truncated…d time: 3.76s, prompts: 4, total tokens: 3972, avg length: 993.0, speed: 1056.8 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3803, avg length: 950.8, speed: 1008.5 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1079.5 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 264
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "264",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 21.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 2000)
Total output lines: 65
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3014, avg length: 753.5, speed: 808.2 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3775, avg length: 943.8, speed: 1011.5 tok/s
vLLM gener…27 tokens truncated…ed time: 3.74s, prompts: 4, total tokens: 3811, avg length: 952.8, speed: 1019.5 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3858, avg length: 964.5, speed: 1023.7 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3870, avg length: 967.5, speed: 1024.2 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 265
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "265",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 20.2 seconds
Output:
Warning: truncated output (original token count: 186)
Total output lines: 7
Warning: truncated output (original token count: 2032)
Total output lines: 69
vLLM generation done - elapsed time: 3.61s, prompts: 4, total tokens: 2864, avg length: 716.0, speed: 792.7 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3137, avg length: 784.2, speed: 846.1 tok/s
vLLM genera…26 tokens truncated…█████▌ | 76/100 [11:20<03:35, 8.99s/it]
{'loss': -0.0245, 'grad_norm': 0.08779964596033096, 'learning_rate': 1.25e-06, 'on_policy_loss': -0.0245, 'rollout_accuracy': 0.1875, 'epoch': 0.28}
76%|███████▌ | 76/100 [11:20<03:35, 8.99s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 266
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "266",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 27.9 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 1882)
Total output lines: 61
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3617, avg length: 904.2, speed: 968.1 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3677, avg length: 919.2, speed: 977.3 tok/s
vLLM genera…27 tokens truncated…psed time: 3.74s, prompts: 4, total tokens: 3547, avg length: 886.8, speed: 949.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3517, avg length: 879.2, speed: 939.5 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3969, avg length: 992.2, speed: 1053.8 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}exec result
Script running with cell ID 267
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "267",
"yield_time_ms": 50000,
"max_tokens": 160
}wait result
Script completed
Wall time 33.3 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8
Warning: truncated output (original token count: 4557)
Total output lines: 393
vLLM generation done - elapsed time: 3.16s, prompts: 4, total tokens: 2789, avg length: 697.2, speed: 883.9 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3248, avg length: 812.0, speed: 876.7 tok/s
vLLM gener…27 tokens truncated… time: 3.70s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1106.6 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3812, avg length: 953.0, speed: 1027.0 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1089.2 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 268
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "268",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 29.1 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 12
Warning: truncated output (original token count: 3102)
Total output lines: 245
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3496, avg length: 874.0, speed: 955.0 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3382, avg length: 845.5, speed: 920.3 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3410, avg length: 852.5, speed: 924.7 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3968, avg length: 992.0, speed: 1064.3 tok/s
90%|█████████ | 90/100 [13:24<01:28, 8…27 tokens truncated…licy_loss': -0.029, 'rollout_accuracy': 0.0625, 'epoch': 0.35}
96%|█████████▌| 96/100 [14:18<00:35, 8.99s/it]vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3804, avg length: 951.0, speed: 1015.7 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4043, avg length: 1010.8, speed: 1073.1 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1080.9 tok/s
vLLM generation done - elapsed time: 3.82s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1071.3 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}exec result
Script completed
Wall time 8.7 seconds
Output:
Warning: truncated output (original token count: 1399)
Total output lines: 37
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1088.5 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3908, avg length: 977.0, speed: 1038.2 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4051, avg length: 1012.8, speed: 1070.8 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1082.0 tok/s
97%|█████████▋| 97/100 [14:27<00:26, 8.98s/it]vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3056, avg length: 764.0, speed: 824.8 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3616, avg length: 904.0, speed: 972.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3856, avg length: 964.0, speed: 1030.1 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1084.4 tok/s
v…899 tokens truncated…s. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 08:21:46.627373896 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 08:21:46.191261281 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 08:21:46.191628873 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/frontier_v4_clip002/capped_seed42/checkpoint-100