Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#426
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "210",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 40.9 seconds
Output:
{"chunk_id":"f212aa","wall_time_seconds":60.001755625,"session_id":99902,"original_token_count":0,"output":""}
#427
narration

Evaluation is progressing normally: one shard is at 25%, another at 14%, with the remaining engines still in long initial traces.

#428
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 211
Wall time 11.0 seconds
Output:
#429
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "211",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 44.3 seconds
Output:
{"chunk_id":"9160d0","wall_time_seconds":60.000838765,"session_id":99902,"original_token_count":32,"output":"\rProcessed prompts:  12%|█▎        | 12/96 [07:02<49:18, 35.22s/it, est. speed input: 4.20 toks/s, output: 266.10 toks/s]"}
#430
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 212
Wall time 11.0 seconds
Output:
#431
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "212",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 39.6 seconds
Output:
{"chunk_id":"d806fd","wall_time_seconds":60.001797845,"session_id":99902,"original_token_count":0,"output":""}
#432
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 213
Wall time 11.0 seconds
Output:
#433
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "213",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 37.2 seconds
Output:
{"chunk_id":"3c92e1","wall_time_seconds":60.002366138,"session_id":99902,"original_token_count":64,"output":"\rProcessed prompts:  14%|█▍        | 12/84 [09:06<54:41, 45.58s/it, est. speed input: 1.84 toks/s, output: 273.89 toks/s]\rProcessed prompts:  38%|███▊      | 36/96 [09:29<17:24, 17.41s/it, est. speed input: 8.30 toks/s, output: 571.71 toks/s] "}
#434
narration

Two shards are in unusually long generations, while the leading shard is at 38%. This is within the benchmark’s normal variability and the run remains healthy.

#435
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 214
Wall time 11.0 seconds
Output:
#436
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "214",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 42.4 seconds
Output:
{"chunk_id":"68d67e","wall_time_seconds":60.002340552,"session_id":99902,"original_token_count":98,"output":"\rProcessed prompts:  50%|█████     | 48/96 [09:51<08:59, 11.23s/it, est. speed input: 9.79 toks/s, output: 845.94 toks/s]\rProcessed prompts:  29%|██▊       | 24/84 [10:08<25:51, 25.87s/it, est. speed input: 4.18 toks/s, output: 600.73 toks/s]\rProcessed prompts:  57%|█████▋    | 48/84 [10:50<06:26, 10.73s/it, est. speed input: 8.82 toks/s, output: 1344.73 toks/s]"}
#437
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 215
Wall time 11.0 seconds
Output:
#438
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "215",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 43.6 seconds
Output:
Warning: truncated output (original token count: 206)
Total output lines: 1

{"chunk_id":"62df51","wall_time_seconds":60.001190718,"session_id":99902,"original_token_count":244,"output":"Warning: truncated output (original token count: 244)\nTotal output lines: 1\n\n\rProcessed prompts:  71%|███████▏  | 60/84 [11:11<03:11,  7.97s/it, est. speed input: 9.84 to…56 tokens truncated… toks/s]\rProcessed prompts: 100%|██████████| 84/84 [12:05<00:00,  5.07s/it, est. speed input: 12.04 toks/s, output: 2129.77 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [12:05<00:00,  8.64s/it, est. speed input: 12.04 toks/s, output: 2129.77 toks/s]\n"}
#439
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 216
Wall time 11.0 seconds
Output:
#440
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "216",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 43.3 seconds
Output:
Warning: truncated output (original token count: 206)
Total output lines: 1

{"chunk_id":"9cadda","wall_time_seconds":60.001386239,"session_id":99902,"original_token_count":245,"output":"Warning: truncated output (original token count: 245)\nTotal output lines: 2\n\n\rProcessed prompts:  75%|███████▌  | 72/96 [12:28<03:22,  8.44s/it, est. speed input: 16.34 t…56 tokens truncated…utput: 2097.50 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:10<00:00,  8.24s/it, est. speed input: 19.20 toks/s, output: 2097.50 toks/s]\n\rProcessed prompts:  38%|███▊      | 36/96 [13:26<19:06, 19.10s/it, est. speed input: 6.95 toks/s, output: 750.20 toks/s]"}
#441
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 217
Wall time 11.0 seconds
Output:
#442
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "217",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 40.0 seconds
Output:
{"chunk_id":"208d07","wall_time_seconds":60.001079274,"session_id":99902,"original_token_count":98,"output":"\rProcessed prompts:  50%|█████     | 48/96 [13:36<09:30, 11.88s/it, est. speed input: 8.11 toks/s, output: 1045.12 toks/s]\rProcessed prompts:  29%|██▊       | 24/84 [14:08<33:34, 33.57s/it, est. speed input: 2.70 toks/s, output: 450.74 toks/s]\rProcessed prompts:  43%|████▎     | 36/84 [14:17<14:51, 18.57s/it, est. speed input: 4.07 toks/s, output: 779.20 toks/s]"}
#443
narration

Two shards are complete; the remaining pair are at roughly 50% and 43%. I’ll report the replicated mean immediately after they finish.

#444
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 218
Wall time 11.0 seconds
Output:
#445
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "218",
  "yield_time_ms": 60000,
  "max_tokens": 150
}
wait result
Script completed
Wall time 43.3 seconds
Output:
Warning: truncated output (original token count: 206)
Total output lines: 1

{"chunk_id":"e70957","wall_time_seconds":60.000879185,"session_id":99902,"original_token_count":245,"output":"Warning: truncated output (original token count: 245)\nTotal output lines: 1\n\n\rProcessed prompts:  57%|█████▋    | 48/84 [14:37<07:08, 11.90s/it, est. speed input: 5.62 toks/s…56 tokens truncated… toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:36<00:00,  4.77s/it, est. speed input: 10.63 toks/s, output: 2017.62 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:36<00:00, 11.14s/it, est. speed input: 10.63 toks/s, output: 2017.62 toks/s]\n"}
#446
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99902,chars:\"\",yield_time_ms:60000,max_output_tokens:150});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{"chunk_id":"bfc336","wall_time_seconds":0.000004105,"exit_code":0,"original_token_count":142,"output":"\rProcessed prompts:  75%|███████▌  | 72/96 [15:59<03:21,  8.38s/it, est. speed input: 10.08 toks/s, output: 1439.18 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [15:59<00:00,  8.38s/it, est. speed input: 13.64 toks/s, output: 2135.04 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [15:59<00:00,  9.99s/it, est. speed input: 13.64 toks/s, output: 2135.04 toks/s]\n\nFINAL avg@12 accuracy: 55.28% (199/360 solutions correct on 30 problems)\nResults written to /workspace/eval_results/dev_eval_1785743764.json\n"}
#447
narration

The strong repeat scored 55.28%, giving a two-run mean of 56.11—the same mean as the milder schedule, but with much higher spread (56.94/55.28 versus 56.11/56.11). I’m selecting the stable 1.5→0.5 credit schedule. One final orthogonal test will halve the pointwise clip to 0.025, as the paper explicitly identifies that knob as untuned; it must clearly beat 56.11 to be retained.

#448
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/repo/opsd_trainer.py\\n@@\\n-        # scale.  The linear 1.75 -> 0.25 schedule has mean one on every non-empty\\n+        # scale.  The linear 1.5 -> 0.5 schedule has mean one on every non-empty\\n@@\\n-            credit = (1.75 - 1.5 * progress).clamp_min(0.25) * mask\\n+            credit = (1.5 - progress).clamp_min(0.5) * mask\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n-# jsd_token_clip=0.05\\n+jsd_token_clip=0.025\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/earlycredit_clip025 PORT=12956 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{}{"chunk_id":"2a8254","wall_time_seconds":1.002017246,"session_id":11520,"original_token_count":0,"output":""}
#449
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:11520,chars:\"\",yield_time_ms:60000,max_output_tokens:140});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 221
Wall time 11.0 seconds
Output:
#450
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "221",
  "yield_time_ms": 60000,
  "max_tokens": 140
}
wait result
Script completed
Wall time 37.7 seconds
Output:
Warning: truncated output (original token count: 198)
Total output lines: 1

{"chunk_id":"c585d8","wall_time_seconds":60.001900917,"session_id":11520,"original_token_count":4208,"output":"Warning: truncated output (original token count: 4208)\nTotal output lines: 200\n\n<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version…58 tokens truncated…811, avg length: 952.8, speed: 1139.1 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1222.2 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3939, avg length: 984.8, speed: 1172.0 tok/s\n"}