Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#451
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 222
Wall time 11.0 seconds
Output:
#452
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "222",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 30.9 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"e939a7","wall_time_seconds":60.00104655,"session_id":11520,"original_token_count":4725,"output":"Warning: truncated output (original token count: 4725)\nTotal output lines: 418\n\nvLLM generation done - elapsed time: 3.17s, prompts: 4, total tokens: 2484, avg length:…57 tokens truncated… 2864, avg length: 716.0, speed: 899.2 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3303, avg length: 825.8, speed: 1011.2 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3394, avg length: 848.5, speed: 1036.8 tok/s\n"}
#453
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 223
Wall time 11.0 seconds
Output:
#454
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "223",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 39.5 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"e84145","wall_time_seconds":60.001387015,"session_id":11520,"original_token_count":3382,"output":"Warning: truncated output (original token count: 3382)\nTotal output lines: 183\n\nvLLM generation done - elapsed time: 3.17s, prompts: 4, total tokens: 2868, avg length…57 tokens truncated… 3555, avg length: 888.8, speed: 1096.3 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3201, avg length: 800.2, speed: 983.4 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3480, avg length: 870.0, speed: 1066.4 tok/s\n"}
#455
narration

The 0.025-clip run is stable so far. Because clipping changes the loss’s effective shape, I’m monitoring its final magnitude as well as accuracy, but only the full AIME score can justify adopting it.

#456
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 224
Wall time 11.0 seconds
Output:
#457
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "224",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 33.0 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"7c7243","wall_time_seconds":60.001517511,"session_id":11520,"original_token_count":2440,"output":"Warning: truncated output (original token count: 2440)\nTotal output lines: 85\n\nvLLM generation done - elapsed time: 3.20s, prompts: 4, total tokens: 2630, avg length:…57 tokens truncated…ion done - elapsed time: 3.33s, prompts: 4, total tokens: 3602, avg length: 900.5, speed: 1081.6 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3490, avg length: 872.5, speed: 1047.1 tok/s\n\r 33%|███▎      | 33/100 [04:19<08:49,  7.90s/it]"}
#458
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 225
Wall time 11.0 seconds
Output:
#459
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "225",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 41.3 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"bade52","wall_time_seconds":60.001749789,"session_id":11520,"original_token_count":2736,"output":"Warning: truncated output (original token count: 2736)\nTotal output lines: 93\n\nvLLM generation done - elapsed time: 3.21s, prompts: 4, total tokens: 2709, avg length:…57 tokens truncated…6, avg length: 1024.0, speed: 1231.5 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.8 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4081, avg length: 1020.2, speed: 1216.7 tok/s\n"}
#460
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 226
Wall time 11.0 seconds
Output:
#461
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "226",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 29.2 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"10915c","wall_time_seconds":60.001295672,"session_id":11520,"original_token_count":3004,"output":"Warning: truncated output (original token count: 3004)\nTotal output lines: 102\n\nvLLM generation done - elapsed time: 3.05s, prompts: 4, total tokens: 2214, avg length…57 tokens truncated…one - elapsed time: 3.32s, prompts: 4, total tokens: 3783, avg length: 945.8, speed: 1139.8 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1222.9 tok/s\n\r 53%|█████▎    | 53/100 [06:57<06:12,  7.93s/it]"}
#462
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 227
Wall time 11.0 seconds
Output:
#463
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "227",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 37.4 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"8b0309","wall_time_seconds":60.000567601,"session_id":11520,"original_token_count":3651,"output":"Warning: truncated output (original token count: 3651)\nTotal output lines: 303\n\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3363, avg length…57 tokens truncated…67, avg length: 916.8, speed: 1107.3 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1233.9 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1233.3 tok/s\n"}
#464
narration

The clip-tuning run has passed halfway with no instability or throughput change. This is the final exploratory candidate; the replicated default-clip method remains the fallback.

#465
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 228
Wall time 11.0 seconds
Output:
#466
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "228",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 31.7 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"f08e25","wall_time_seconds":60.001243554,"session_id":11520,"original_token_count":2777,"output":"Warning: truncated output (original token count: 2777)\nTotal output lines: 94\n\nvLLM generation done - elapsed time: 3.20s, prompts: 4, total tokens: 2922, avg length:…57 tokens truncated…█▏  | 72/100 [09:28<03:43,  7.97s/it]\r                                                \r{'loss': -0.0295, 'grad_norm': 0.06029777601361275, 'learning_rate': 5e-06, 'on_policy_loss': -0.0295, 'epoch': 0.08}\n\r 72%|███████▏  | 72/100 [09:28<03:43,  7.97s/it]"}
#467
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 229
Wall time 11.0 seconds
Output:
#468
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "229",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 43.8 seconds
Output:
Warning: truncated output (original token count: 198)
Total output lines: 1

{"chunk_id":"ad7b4c","wall_time_seconds":60.002866562,"session_id":11520,"original_token_count":2597,"output":"Warning: truncated output (original token count: 2597)\nTotal output lines: 89\n\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3608, avg length:…58 tokens truncated…===\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_clip025/capped_seed42/generations/generations_step_80.json\n================================================================================\n\n\r 81%|████████  | 81/100 [10:39<02:29,  7.89s/it]"}
#469
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 230
Wall time 11.0 seconds
Output:
#470
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "230",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 43.3 seconds
Output:
Warning: truncated output (original token count: 197)
Total output lines: 1

{"chunk_id":"5ddaf6","wall_time_seconds":60.001269076,"session_id":11520,"original_token_count":4432,"output":"Warning: truncated output (original token count: 4432)\nTotal output lines: 378\n\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3293, avg length…57 tokens truncated…█ | 90/100 [11:50<01:19,  7.93s/it]\r                                                \r{'loss': -0.0312, 'grad_norm': 0.050000905990600586, 'learning_rate': 5e-06, 'on_policy_loss': -0.0312, 'epoch': 0.1}\n\r 90%|█████████ | 90/100 [11:50<01:19,  7.93s/it]"}
#471
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 231
Wall time 11.0 seconds
Output:
#472
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "231",
  "yield_time_ms": 60000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 40.3 seconds
Output:
Warning: truncated output (original token count: 561)
Total output lines: 1

{"chunk_id":"e9973d","wall_time_seconds":60.001630738,"session_id":11520,"original_token_count":3422,"output":"Warning: truncated output (original token count: 3422)\nTotal output lines: 218\n\nvLLM generation done - elapsed time: 3.25s, prompts: 4, total tokens: 2801, avg length: 700.2, speed: 862.9 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3366, avg length: 841.5, speed: 1023.1 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3706, avg length: 926.5, speed: 1125.7 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1229.7 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3552, avg length: 888.0, speed: 1073.4 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3483, avg length: 870.8, speed: 1051.2 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4095, avg length: 1023.8, speed: 1226.1 …61 tokens truncated…generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3503, avg length: 875.8, speed: 1069.2 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3354, avg length: 838.5, speed: 1022.0 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1230.5 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1223.4 tok/s\n\r 99%|█████████▉| 99/100 [13:02<00:07,  7.98s/it]vLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3670, avg length: 917.5, speed: 1110.7 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1234.5 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.7 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1226.9 tok/s\n"}
#473
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:30000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{"chunk_id":"f96002","wall_time_seconds":0.000004364,"exit_code":0,"original_token_count":566,"output":"Warning: truncated output (original token count: 566)\nTotal output lines: 12\n\nvLLM generation done - elapsed time: 3.20s, prompts: 4, total tokens: 3159, avg length: 789.8, speed: 985.8 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1241.3 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3836, avg length: 959.0, speed: 1152.3 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4064, avg length: 1016.0, speed: 1214.9 tok/s\n\r100%|██████████| 100/100 [13:10<00:00,  7.99s/it]\r                                                 \r{'loss': -0.033, 'grad_norm': 0.05662483721971512, 'learning_rate': 5e-06, 'on_policy_loss': -0.033, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:10<00:00,  7.99s/it]\r                                                 \r{'train_runtime': 793.7423, 'train_samples_per_second': 4.032, 'train_steps_per_second': 0.126, 'train_loss': -0.022480114698410034, 'epoch': 0.11}\n\r100%|█████████…66 tokens truncated…ources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank3]:[W803 08:27:09.638211187 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank2]:[W803 08:27:09.658859877 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank1]:[W803 08:27:09.659117531 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\nTRAIN_CKPT=/workspace/runs/earlycredit_clip025/capped_seed42/checkpoint-100\n"}
#474
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/earlycredit_clip025/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"4e8cf5","wall_time_seconds":1.002178701,"session_id":70127,"original_token_count":0,"output":""}
#475
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:70127,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 234
Wall time 11.0 seconds
Output: