Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#276
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 138
Wall time 11.0 seconds
Output:
#277
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "138",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 40.1 seconds
Output:
Warning: truncated output (original token count: 278)
Total output lines: 1

{"chunk_id":"0c1ca9","wall_time_seconds":60.001605005,"session_id":94634,"original_token_count":2754,"output":"Warning: truncated output (original token count: 2754)\nTotal output lines: 93\n\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 2960, avg length: 740.0, speed: 916.2 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3408, avg length: 852.0, speed: 1032.0 tok/s\nvLLM generation …58 tokens truncated…psed time: 3.26s, prompts: 4, total tokens: 2492, avg length: 623.0, speed: 763.9 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3353, avg length: 838.2, speed: 1011.9 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1218.9 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3816, avg length: 954.0, speed: 1134.4 tok/s\n"}
#278
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 139
Wall time 11.0 seconds
Output:
#279
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "139",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 40.1 seconds
Output:
Warning: truncated output (original token count: 278)
Total output lines: 1

{"chunk_id":"6ba2f1","wall_time_seconds":60.002777454,"session_id":94634,"original_token_count":2374,"output":"Warning: truncated output (original token count: 2374)\nTotal output lines: 81\n\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3678, avg length: 919.5, speed: 1098.2 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1220.4 tok/s\nvLLM generatio…58 tokens truncated…psed time: 3.35s, prompts: 4, total tokens: 3794, avg length: 948.5, speed: 1131.6 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3935, avg length: 983.8, speed: 1170.2 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1208.2 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 2726, avg length: 681.5, speed: 833.7 tok/s\n"}
#280
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 140
Wall time 11.0 seconds
Output:
#281
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "140",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 43.2 seconds
Output:
Warning: truncated output (original token count: 279)
Total output lines: 1

{"chunk_id":"639f19","wall_time_seconds":60.001903665,"session_id":94634,"original_token_count":2760,"output":"Warning: truncated output (original token count: 2760)\nTotal output lines: 93\n\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3849, avg length: 962.2, speed: 1150.6 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3810, avg length: 952.5, speed: 1138.4 tok/s\nvLLM generation…59 tokens truncated… 1034.9 tok/s\nvLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 3861, avg length: 965.2, speed: 1142.2 tok/s\n\r 82%|████████▏ | 82/100 [10:55<02:23,  7.95s/it]\r                                                \r{'loss': -0.0142, 'grad_norm': 0.08880437165498734, 'learning_rate': 5e-06, 'on_policy_loss': -0.0142, 'epoch': 0.09}\n\r 82%|████████▏ | 82/100 [10:55<02:23,  7.95s/it]"}
#282
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 141
Wall time 11.0 seconds
Output:
#283
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "141",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 42.5 seconds
Output:
Warning: truncated output (original token count: 279)
Total output lines: 1

{"chunk_id":"909b6a","wall_time_seconds":60.001385975,"session_id":94634,"original_token_count":4392,"output":"Warning: truncated output (original token count: 4392)\nTotal output lines: 375\n\nvLLM generation done - elapsed time: 2.95s, prompts: 4, total tokens: 3139, avg length: 784.8, speed: 1063.3 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3138, avg length: 784.5, speed: 942.5 tok/s\nvLLM generation…59 tokens truncated….36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1218.9 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_repeat2/capped_seed42/generations/generations_step_90.json\n================================================================================\n\n\r 91%|█████████ | 91/100 [12:07<01:11,  7.94s/it]"}
#284
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:700});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 142
Wall time 11.0 seconds
Output:
#285
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "142",
  "yield_time_ms": 60000,
  "max_tokens": 700
}
wait result
Script completed
Wall time 43.4 seconds
Output:
Warning: truncated output (original token count: 764)
Total output lines: 1

{"chunk_id":"d9a85c","wall_time_seconds":60.001351246,"session_id":94634,"original_token_count":3446,"output":"Warning: truncated output (original token count: 3446)\nTotal output lines: 223\n\nvLLM generation done - elapsed time: 3.25s, prompts: 4, total tokens: 3635, avg length: 908.8, speed: 1117.5 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3508, avg length: 877.0, speed: 1066.8 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1240.0 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3715, avg length: 928.8, speed: 1120.7 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3204, avg length: 801.0, speed: 973.5 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3261, avg length: 815.2, speed: 986.7 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3090, avg length: 772.5, speed: 934.4 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1215.3 tok/s\n\r 92%|█████████▏| 92/100 [12:14<01:03,  7.92s/it]\r                                                \r{'loss': -0.0134, 'grad_norm': 0.05249234661459923, 'learning_rate': 5e-06, 'on_policy_loss': -0.0134, 'epoch': 0.1}\n\r 92%|█████████▏| …64 tokens truncated…avg length: 951.2, speed: 1156.1 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1238.9 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1235.9 tok/s\nvLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 3182, avg length: 795.5, speed: 996.0 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3462, avg length: 865.5, speed: 1060.6 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1250.3 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1247.8 tok/s\n\r100%|██████████| 100/100 [13:17<00:00,  7.83s/it]\r                                                 \r{'loss': -0.0148, 'grad_norm': 0.0469496063888073, 'learning_rate': 5e-06, 'on_policy_loss': -0.0148, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:17<00:00,  7.83s/it]\r                                                 \r{'train_runtime': 801.1761, 'train_samples_per_second': 3.994, 'train_steps_per_second': 0.125, 'train_loss': -0.005652620196342468, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:21<00:00,  7.83s/it]\r100%|██████████| 100/100 [13:21<00:00,  8.01s/it]\n"}
#286
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:30000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{"chunk_id":"ed52c7","wall_time_seconds":0.000004256,"exit_code":0,"original_token_count":289,"output":"[rank0]:[W803 06:51:01.733706812 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank2]:[W803 06:51:01.406774764 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank1]:[W803 06:51:01.412000192 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank3]:[W803 06:51:01.421899366 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\nTRAIN_CKPT=/workspace/runs/earlycredit_repeat2/capped_seed42/checkpoint-100\n"}
#287
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/earlycredit_repeat2/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":500});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"e97b8b","wall_time_seconds":1.000897785,"session_id":3244,"original_token_count":0,"output":""}
#288
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 145
Wall time 11.0 seconds
Output:
#289
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "145",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.3 seconds
Output:
Warning: truncated output (original token count: 310)
Total output lines: 1

{"chunk_id":"57ad05","wall_time_seconds":60.000659216,"session_id":3244,"original_token_count":5933,"output":"Warning: truncated output (original token count: 5933)\nTotal output lines: 154\n\n`torch_dtype` is deprecated! Use `dtype` instead!\nCheckpoint kind: lora\nMerging LoRA adapter into base weights for evaluation (CPU, once)...\n\rLoading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]\rLoading checkpoint shards: 100%|██████████| 2/2 [00:00<00:00, 53.44it/s]\nG…60 tokens truncated…2 [llm.py:306] Supported_tasks: ['generate']\n\rAdding requests: 100%|██████████| 8/8 [00:00<00:00, 273.79it/s]\n\rProcessed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]\rAdding requests:   0%|          | 0/7 [00:00<?, ?it/s]\rAdding requests: 100%|██████████| 7/7 [00:00<00:00, 267.04it/s]\n\rProcessed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]"}
#290
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 146
Wall time 11.0 seconds
Output:
#291
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "146",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 43.2 seconds
Output:
{"chunk_id":"4c94d0","wall_time_seconds":60.002062226,"session_id":3244,"original_token_count":0,"output":""}
#292
narration

The second independent early-credit checkpoint is now in full evaluation. Its final training loss (−0.0148) is close to the first constant run’s −0.0139, so the optimizer reached a comparable regime despite different sampled rollouts.

#293
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 147
Wall time 11.0 seconds
Output:
#294
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "147",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 41.1 seconds
Output:
{"chunk_id":"568b43","wall_time_seconds":60.001308844,"session_id":3244,"original_token_count":0,"output":""}
#295
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 148
Wall time 11.0 seconds
Output:
#296
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "148",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.1 seconds
Output:
{"chunk_id":"de8980","wall_time_seconds":60.000918106,"session_id":3244,"original_token_count":32,"output":"\rProcessed prompts:  12%|█▎        | 12/96 [03:37<25:20, 18.11s/it, est. speed input: 8.62 toks/s, output: 364.04 toks/s]"}
#297
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 149
Wall time 11.0 seconds
Output:
#298
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "149",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 40.2 seconds
Output:
{"chunk_id":"8638e8","wall_time_seconds":60.000480707,"session_id":3244,"original_token_count":0,"output":""}
#299
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 150
Wall time 11.0 seconds
Output:
#300
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "150",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 39.5 seconds
Output:
{"chunk_id":"74e25a","wall_time_seconds":60.002375204,"session_id":3244,"original_token_count":95,"output":"\rProcessed prompts:  14%|█▍        | 12/84 [05:56<35:36, 29.67s/it, est. speed input: 2.60 toks/s, output: 364.71 toks/s]\rProcessed prompts:  25%|██▌       | 24/96 [06:00<17:23, 14.49s/it, est. speed input: 10.75 toks/s, output: 505.74 toks/s]\rProcessed prompts:  12%|█▎        | 12/96 [06:24<44:51, 32.04s/it, est. speed input: 4.62 toks/s, output: 318.68 toks/s]"}