Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#151
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "79",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.1 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"380469","wall_time_seconds":60.001590314,"session_id":92388,"original_token_count":3525,"output":"Warning: truncated output (original token count: 3525)\nTotal output lines: 251\n\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3607, avg length: 901.8, speed: 1098.8 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1241.5 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 380…58 tokens truncated…sed time: 3.23s, prompts: 4, total tokens: 3102, avg length: 775.5, speed: 961.0 tok/s\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 2918, avg length: 729.5, speed: 902.4 tok/s\nvLLM generation done - elapsed time: 3.25s, prompts: 4, total tokens: 3337, avg length: 834.2, speed: 1027.2 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1248.4 tok/s\n\r 59%|█████▉    | 59/100 [07:43<05:18,  7.76s/it]"}
#152
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92388,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 80
Wall time 11.0 seconds
Output:
#153
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.0 seconds
Output:
Warning: truncated output (original token count: 307)
Total output lines: 1

{"chunk_id":"c65b88","wall_time_seconds":60.001296849,"session_id":92388,"original_token_count":2756,"output":"Warning: truncated output (original token count: 2756)\nTotal output lines: 93\n\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 3412, avg length: 853.0, speed: 1055.6 tok/s\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 3073, avg length: 768.2, speed: 950.0 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3591, …57 tokens truncated… 68/100 [08:54<04:12,  7.88s/it]vLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3528, avg length: 882.0, speed: 1079.5 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3835, avg length: 958.8, speed: 1173.0 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1236.1 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1235.6 tok/s\n"}
#154
narration

Early-credit training is two-thirds complete with no instability or throughput penalty. I’m keeping the evaluation protocol identical and will only retain this change if its full fixed-seed result clears the baseline region; otherwise I’ll pivot to teacher-confidence gating or difficulty filtering.

#155
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92388,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 81
Wall time 11.0 seconds
Output:
#156
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "81",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.6 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"c81ad5","wall_time_seconds":60.001253771,"session_id":92388,"original_token_count":2343,"output":"Warning: truncated output (original token count: 2343)\nTotal output lines: 80\n\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1256.7 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3857, avg length: 964.2, speed: 1179.0 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 4096…58 tokens truncated…6/100 [09:58<03:11,  7.98s/it]vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1230.6 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4038, avg length: 1009.5, speed: 1209.4 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.2 tok/s\nvLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1210.8 tok/s\n"}
#157
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92388,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 82
Wall time 11.0 seconds
Output:
#158
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.6 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"9b099a","wall_time_seconds":60.000955457,"session_id":92388,"original_token_count":2279,"output":"Warning: truncated output (original token count: 2279)\nTotal output lines: 74\n\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1232.6 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1230.3 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 379…58 tokens truncated… | 84/100 [11:02<02:08,  8.04s/it]vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3681, avg length: 920.2, speed: 1108.7 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3845, avg length: 961.2, speed: 1157.3 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1218.7 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3086, avg length: 771.5, speed: 915.2 tok/s\n"}
#159
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92388,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 83
Wall time 11.0 seconds
Output:
#160
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "83",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.0 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"b4ca54","wall_time_seconds":60.002184003,"session_id":92388,"original_token_count":5393,"output":"Warning: truncated output (original token count: 5393)\nTotal output lines: 541\n\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3466, avg length: 866.5, speed: 1052.4 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3801, avg length: 950.2, speed: 1140.7 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4035…58 tokens truncated…| 93/100 [12:14<00:55,  7.96s/it]vLLM generation done - elapsed time: 3.24s, prompts: 4, total tokens: 3412, avg length: 853.0, speed: 1051.8 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3414, avg length: 853.5, speed: 1032.1 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3891, avg length: 972.8, speed: 1167.6 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1223.8 tok/s\n"}
#161
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92388,chars:\"\",yield_time_ms:60000,max_output_tokens:1000});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 84
Wall time 11.0 seconds
Output:
#162
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "84",
  "yield_time_ms": 60000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 42.3 seconds
Output:
Warning: truncated output (original token count: 1064)
Total output lines: 1

{"chunk_id":"036051","wall_time_seconds":57.571885329,"exit_code":0,"original_token_count":2300,"output":"Warning: truncated output (original token count: 2300)\nTotal output lines: 69\n\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3803, avg length: 950.8, speed: 1143.1 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.6 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1217.6 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1216.8 tok/s\n\r 94%|█████████▍| 94/100 [12:22<00:47,  7.96s/it]\r                                                \r{'loss': -0.0131, 'grad_norm': 0.05200755223631859, 'learning_rate': 5e-06, 'on_policy_loss': -0.0131, 'epoch': 0.1}\n\r 94%|█████████▍| 94/100 [12:22<00:47,  7.96s/it]vLLM generation done - elapsed time: 3.17s, prompts: 4, total tokens: 2544, avg length: 636.0, speed: 801.3 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3721, avg length: 930.2, speed: 1124.0 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1219.8 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1217.3 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3268, avg length: 817.0, speed: 990.4 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3521, avg length: 880.2, speed: 1065.9 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3535, avg length: 883.8, speed: 1065.3 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.1 tok/s\n\r 95%|█████████▌| 95/100 [12:30<00:39,  7.98s/it]vLLM generation done - elapsed time:…64 tokens truncated…s: 4, total tokens: 4066, avg length: 1016.5, speed: 1212.1 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1218.6 tok/s\n\r100%|██████████| 100/100 [13:10<00:00,  7.99s/it]\r                                                 \r{'loss': -0.0139, 'grad_norm': 0.055231135338544846, 'learning_rate': 5e-06, 'on_policy_loss': -0.0139, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:10<00:00,  7.99s/it]\r                                                 \r{'train_runtime': 793.8716, 'train_samples_per_second': 4.031, 'train_steps_per_second': 0.126, 'train_loss': -0.005379514992237091, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:13<00:00,  7.99s/it]\r100%|██████████| 100/100 [13:13<00:00,  7.94s/it]\n[rank0]:[W803 05:47:53.323354887 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank3]:[W803 05:47:54.983563400 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank2]:[W803 05:47:54.446412276 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank1]:[W803 05:47:55.496787237 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\nTRAIN_CKPT=/workspace/runs/earlycredit/capped_seed42/checkpoint-100\n"}
#163
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/earlycredit/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":1000});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"9ee7fa","wall_time_seconds":1.001236605,"session_id":17991,"original_token_count":0,"output":""}
#164
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 86
Wall time 11.0 seconds
Output:
#165
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "86",
  "yield_time_ms": 60000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 43.9 seconds
Output:
Warning: truncated output (original token count: 563)
Total output lines: 1

{"chunk_id":"517380","wall_time_seconds":60.000781309,"session_id":17991,"original_token_count":5933,"output":"Warning: truncated output (original token count: 5933)\nTotal output lines: 154\n\n`torch_dtype` is deprecated! Use `dtype` instead!\nCheckpoint kind: lora\nMerging LoRA adapter into base weights for evaluation (CPU, once)...\n\rLoading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]\rLoading checkpoint shards: 100%|██████████| 2/2 [00:00<00:00, 49.19it/s]\nGenerating with data_parallel_size=4 (TP=1 per engine) ...\nINFO 08-03 05:48:37 [__init__.py:216] Automatically detected platform cuda.\nINFO 08-03 05:48:38 [__init__.py:216] Automatically detected platform cuda.\nINFO 08-03 05:48:38 [__init__.py:216] Automatically detected platform cuda.\nINFO 08-03 05:48:38 [__init__.py:216] Automatically detected platform cuda.\nINFO 08-03 05:48:40 [arg_utils.py:504] HF_HUB_OFFLINE is True, replace model_id [/tmp/opsd_merged_omh4tgib] to model_path [/tmp/opsd…63 tokens truncated…81] Cudagraph is disabled under eager mode\nINFO 08-03 05:48:56 [llm.py:306] Supported_tasks: ['generate']\n\rAdding requests:   0%|          | 0/7 [00:00<?, ?it/s]\rAdding requests: 100%|██████████| 7/7 [00:00<00:00, 306.29it/s]\n\rProcessed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]INFO 08-03 05:48:56 [llm.py:306] Supported_tasks: ['generate']\n\rAdding requests:   0%|          | 0/8 [00:00<?, ?it/s]\rAdding requests: 100%|██████████| 8/8 [00:00<00:00, 260.26it/s]\n\rProcessed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]INFO 08-03 05:48:56 [llm.py:306] Supported_tasks: ['generate']\n\rAdding requests:   0%|          | 0/8 [00:00<?, ?it/s]\rAdding requests: 100%|██████████| 8/8 [00:00<00:00, 263.49it/s]\n\rProcessed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]"}
#166
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 87
Wall time 11.0 seconds
Output:
#167
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "87",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 45.5 seconds
Output:
{"chunk_id":"dad766","wall_time_seconds":60.001094627,"session_id":17991,"original_token_count":0,"output":""}
#168
narration

While evaluation runs, I found an important reproducibility clue in the supplied artifacts: the baseline log uses a linear decay from 5e−6 to zero, whereas the capped mirror’s nominal default explicitly passes a constant schedule (our logs confirm constant). Since lr_scheduler_type is a permitted method knob, the next candidate will restore linear decay and test the method under the supplied baseline’s actual optimization dynamics.

#169
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 88
Wall time 11.0 seconds
Output:
#170
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "88",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 30.4 seconds
Output:
{"chunk_id":"e9be6f","wall_time_seconds":60.001647347,"session_id":17991,"original_token_count":32,"output":"\rProcessed prompts:  12%|█▎        | 12/96 [02:44<19:13, 13.73s/it, est. speed input: 11.36 toks/s, output: 494.52 toks/s]"}
#171
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 89
Wall time 11.0 seconds
Output:
#172
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "89",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 45.3 seconds
Output:
{"chunk_id":"12b168","wall_time_seconds":60.002058458,"session_id":17991,"original_token_count":32,"output":"\rProcessed prompts:  25%|██▌       | 24/96 [04:05<11:31,  9.60s/it, est. speed input: 15.80 toks/s, output: 746.08 toks/s]"}
#173
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 90
Wall time 11.0 seconds
Output:
#174
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "90",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 44.2 seconds
Output:
{"chunk_id":"d0a184","wall_time_seconds":60.00093737,"session_id":17991,"original_token_count":32,"output":"\rProcessed prompts:  14%|█▍        | 12/84 [04:40<28:02, 23.37s/it, est. speed input: 3.29 toks/s, output: 434.59 toks/s]"}
#175
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 91
Wall time 11.0 seconds
Output: