Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#251
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5527,chars:\"\",yield_time_ms:60000,max_output_tokens:300});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 127
Wall time 11.0 seconds
Output:
#252
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "127",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 40.3 seconds
Output:
{"chunk_id":"82b336","wall_time_seconds":60.001679754,"session_id":5527,"original_token_count":139,"output":"\rProcessed prompts:  62%|██████▎   | 60/96 [12:36<05:03,  8.43s/it, est. speed input: 10.82 toks/s, output: 1360.60 toks/s]\rProcessed prompts:  71%|███████▏  | 60/84 [12:51<03:20,  8.35s/it, est. speed input: 9.80 toks/s, output: 1510.49 toks/s]\rProcessed prompts:  86%|████████▌ | 72/84 [12:59<01:08,  5.75s/it, est. speed input: 11.24 toks/s, output: 1874.57 toks/s]\rProcessed prompts:  75%|███████▌  | 72/96 [13:05<02:33,  6.39s/it, est. speed input: 12.32 toks/s, output: 1603.89 toks/s]"}
#253
narration

Two evaluation shards are complete and the remaining two are above 70%. This run is slightly faster than the constant early-credit evaluation, suggesting it did not induce the same lengthening of test-time reasoning.

#254
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5527,chars:\"\",yield_time_ms:60000,max_output_tokens:300});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 128
Wall time 11.0 seconds
Output:
#255
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "128",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 38.9 seconds
Output:
{"chunk_id":"3bc2c2","wall_time_seconds":60.002059452,"session_id":5527,"original_token_count":108,"output":"\rProcessed prompts: 100%|██████████| 84/84 [14:22<00:00,  6.13s/it, est. speed input: 11.53 toks/s, output: 2037.81 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [14:22<00:00,  6.13s/it, est. speed input: 11.53 toks/s, output: 2037.81 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [14:22<00:00, 10.27s/it, est. speed input: 11.53 toks/s, output: 2037.81 toks/s]\n"}
#256
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5527,chars:\"\",yield_time_ms:60000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 129
Wall time 11.0 seconds
Output:
#257
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 32.8 seconds
Output:
{"chunk_id":"1410d1","wall_time_seconds":48.891979715,"exit_code":0,"original_token_count":179,"output":"\rProcessed prompts:  88%|████████▊ | 84/96 [15:21<01:36,  8.02s/it, est. speed input: 11.91 toks/s, output: 1726.01 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [15:39<00:00,  5.93s/it, est. speed input: 13.92 toks/s, output: 2136.62 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [15:39<00:00,  5.93s/it, est. speed input: 13.92 toks/s, output: 2136.62 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [15:39<00:00,  9.78s/it, est. speed input: 13.92 toks/s, output: 2136.62 toks/s]\n\nFINAL avg@12 accuracy: 54.72% (197/360 solutions correct on 30 problems)\nResults written to /workspace/eval_results/dev_eval_1785738016.json\n"}
#258
narration

Linear early-credit scored 54.72% (197/360), below the constant version’s 56.11. I’m reverting to constant and beginning an independent re-train of the winning method. This repeat is more valuable than adding another mechanism now: it tests whether the +1.89 result survives the stated ~1.5-point per-run variation.

#259
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n-lr_scheduler_type=linear\\n+# lr_scheduler_type=constant\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/earlycredit_repeat2 PORT=12953 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":1000,\"max_output_tokens\":800});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{}{"chunk_id":"378cdd","wall_time_seconds":1.00248083,"session_id":94634,"original_token_count":0,"output":""}
#260
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:300});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 131
Wall time 11.0 seconds
Output:
#261
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "131",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 44.8 seconds
Output:
Warning: truncated output (original token count: 360)
Total output lines: 1

{"chunk_id":"358f82","wall_time_seconds":60.002052877,"session_id":94634,"original_token_count":3883,"output":"Warning: truncated output (original token count: 3883)\nTotal output lines: 190\n\n<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.\n\n================================================================================\nRUN CONFIGURATION\n================================================================================\nWandB Run Name: capped_seed42_lr5e-6_bs32\nOutput Directory: /workspace/runs/earlycredit_repeat2/…60 tokens truncated….16804006695747375, 'learning_rate': 5e-06, 'on_policy_loss': 0.0127, 'epoch': 0.0}\n\r  2%|▏         | 2/100 [00:16<13:13,  8.10s/it]vLLM generation done - elapsed time: 3.12s, prompts: 4, total tokens: 2192, avg length: 548.0, speed: 702.5 tok/s\nvLLM generation done - elapsed time: 3.18s, prompts: 4, total tokens: 2259, avg length: 564.8, speed: 710.3 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3645, avg length: 911.2, speed: 1091.3 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3992, avg length: 998.0, speed: 1189.9 tok/s\n"}
#262
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 132
Wall time 11.0 seconds
Output:
#263
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "132",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 42.5 seconds
Output:
Warning: truncated output (original token count: 280)
Total output lines: 1

{"chunk_id":"5d9c5e","wall_time_seconds":60.002025035,"session_id":94634,"original_token_count":3977,"output":"Warning: truncated output (original token count: 3977)\nTotal output lines: 349\n\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 2907, avg length: 726.8, speed: 886.5 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3000, avg length: 750.0, speed: 911.8 tok/s\nvLLM generation …60 tokens truncated…|im_end|>\n================================================================================\n\n\n================================================================================\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_repeat2/capped_seed42/generations/generations_step_10.json\n================================================================================\n\n\r 11%|█         | 11/100 [01:27<11:48,  7.96s/it]"}
#264
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 133
Wall time 11.0 seconds
Output:
#265
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "133",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 42.0 seconds
Output:
Warning: truncated output (original token count: 278)
Total output lines: 1

{"chunk_id":"efb68d","wall_time_seconds":60.000301794,"session_id":94634,"original_token_count":2344,"output":"Warning: truncated output (original token count: 2344)\nTotal output lines: 78\n\nvLLM generation done - elapsed time: 3.14s, prompts: 4, total tokens: 2199, avg length: 549.8, speed: 700.3 tok/s\nvLLM generation done - elapsed time: 3.21s, prompts: 4, total tokens: 2540, avg length: 635.0, speed: 790.7 tok/s\nvLLM generation d…58 tokens truncated…apsed time: 2.87s, prompts: 4, total tokens: 2616, avg length: 654.0, speed: 913.1 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3493, avg length: 873.2, speed: 1048.9 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3461, avg length: 865.2, speed: 1036.1 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3400, avg length: 850.0, speed: 1014.6 tok/s\n"}
#266
narration

The independent constant-schedule repeat is underway and showing the expected rollout nondeterminism despite the same seed. That is useful: the final comparison will genuinely sample the method’s training variance rather than accidentally duplicate the first checkpoint.

#267
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 134
Wall time 11.0 seconds
Output:
#268
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "134",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 36.4 seconds
Output:
Warning: truncated output (original token count: 278)
Total output lines: 1

{"chunk_id":"ae3cb1","wall_time_seconds":60.000148037,"session_id":94634,"original_token_count":3418,"output":"Warning: truncated output (original token count: 3418)\nTotal output lines: 219\n\nvLLM generation done - elapsed time: 3.22s, prompts: 4, total tokens: 2579, avg length: 644.8, speed: 801.6 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3260, avg length: 815.0, speed: 999.6 tok/s\nvLLM generation …58 tokens truncated…lapsed time: 3.18s, prompts: 4, total tokens: 2629, avg length: 657.2, speed: 826.4 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 2913, avg length: 728.2, speed: 882.2 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3336, avg length: 834.0, speed: 1005.2 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3703, avg length: 925.8, speed: 1108.2 tok/s\n"}
#269
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 135
Wall time 11.0 seconds
Output:
#270
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "135",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 38.9 seconds
Output:
Warning: truncated output (original token count: 278)
Total output lines: 1

{"chunk_id":"af4dbf","wall_time_seconds":60.00095257,"session_id":94634,"original_token_count":2334,"output":"Warning: truncated output (original token count: 2334)\nTotal output lines: 81\n\n\r 29%|██▉       | 29/100 [03:51<09:25,  7.97s/it]vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3629, avg length: 907.2, speed: 1085.3 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4043, …58 tokens truncated…h: 889.8, speed: 1090.9 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3684, avg length: 921.0, speed: 1099.1 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3935, avg length: 983.8, speed: 1171.3 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1209.4 tok/s\n\r 37%|███▋      | 37/100 [04:55<08:25,  8.03s/it]"}
#271
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 136
Wall time 11.0 seconds
Output:
#272
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "136",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 44.6 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 1

{"chunk_id":"98b398","wall_time_seconds":60.00243665,"session_id":94634,"original_token_count":2485,"output":"Warning: truncated output (original token count: 2485)\nTotal output lines: 82\n\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3266, avg length: 816.5, speed: 985.1 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1220.5 tok/s\nvLLM generation …57 tokens truncated…lapsed time: 3.29s, prompts: 4, total tokens: 2953, avg length: 738.2, speed: 897.7 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3014, avg length: 753.5, speed: 912.5 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3572, avg length: 893.0, speed: 1074.7 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3433, avg length: 858.2, speed: 1023.3 tok/s\n"}
#273
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:94634,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 137
Wall time 11.0 seconds
Output:
#274
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "137",
  "yield_time_ms": 60000,
  "max_tokens": 220
}
wait result
Script completed
Wall time 45.0 seconds
Output:
Warning: truncated output (original token count: 280)
Total output lines: 1

{"chunk_id":"716cf2","wall_time_seconds":60.000924666,"session_id":94634,"original_token_count":3534,"output":"Warning: truncated output (original token count: 3534)\nTotal output lines: 239\n\n\n================================================================================\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_repeat2/capped_seed42/generations/generations_step_45.json\n=======================================…60 tokens truncated… 786.8, speed: 958.6 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3269, avg length: 817.2, speed: 988.1 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3351, avg length: 837.8, speed: 1010.6 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1217.8 tok/s\n\r 55%|█████▌    | 55/100 [07:20<05:59,  8.00s/it]"}
#275
narration

The repeat has passed halfway with the same stable ~8 seconds/step profile. I’ll evaluate it fully, then use the two-run mean to decide whether to tune the early-weight slope or lock the current 1.5→0.5 schedule.