Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#176
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "91",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 44.5 seconds
Output:
{"chunk_id":"1b6848","wall_time_seconds":60.001307553,"session_id":17991,"original_token_count":32,"output":"\rProcessed prompts:  12%|█▎        | 12/96 [05:54<41:19, 29.52s/it, est. speed input: 5.01 toks/s, output: 324.65 toks/s]"}
#177
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 92
Wall time 11.0 seconds
Output:
#178
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "92",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 44.8 seconds
Output:
{"chunk_id":"3b0cb3","wall_time_seconds":60.001173578,"session_id":17991,"original_token_count":33,"output":"\rProcessed prompts:  38%|███▊      | 36/96 [07:08<12:11, 12.19s/it, est. speed input: 11.04 toks/s, output: 760.19 toks/s]"}
#179
narration

The early-credit checkpoint is producing somewhat longer evaluation traces than the first candidate, but all engines are progressing normally. The accuracy result—not trace length—will determine whether temporal weighting survives.

#180
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 93
Wall time 11.0 seconds
Output:
#181
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "93",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 43.7 seconds
Output:
{"chunk_id":"1920c4","wall_time_seconds":60.000987859,"session_id":17991,"original_token_count":32,"output":"\rProcessed prompts:  14%|█▍        | 12/84 [07:53<47:19, 39.44s/it, est. speed input: 2.13 toks/s, output: 285.27 toks/s]"}
#182
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 94
Wall time 11.0 seconds
Output:
#183
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "94",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 44.0 seconds
Output:
{"chunk_id":"2acd2a","wall_time_seconds":60.002191209,"session_id":17991,"original_token_count":99,"output":"\rProcessed prompts:  29%|██▊       | 24/84 [09:13<23:00, 23.02s/it, est. speed input: 3.47 toks/s, output: 510.25 toks/s]\rProcessed prompts:  50%|█████     | 48/96 [09:14<09:12, 11.51s/it, est. speed input: 11.15 toks/s, output: 915.56 toks/s]\rProcessed prompts:  62%|██████▎   | 60/96 [09:43<04:56,  8.24s/it, est. speed input: 12.40 toks/s, output: 1204.51 toks/s]"}
#184
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 95
Wall time 11.0 seconds
Output:
#185
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "95",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 45.5 seconds
Output:
{"chunk_id":"4535ca","wall_time_seconds":60.001133118,"session_id":17991,"original_token_count":33,"output":"\rProcessed prompts:  43%|████▎     | 36/84 [09:57<11:19, 14.16s/it, est. speed input: 4.68 toks/s, output: 794.96 toks/s]"}
#186
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 96
Wall time 11.0 seconds
Output:
#187
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "96",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 45.2 seconds
Output:
{"chunk_id":"58d24c","wall_time_seconds":60.000654341,"session_id":17991,"original_token_count":100,"output":"\rProcessed prompts:  57%|█████▋    | 48/84 [11:48<07:19, 12.22s/it, est. speed input: 5.71 toks/s, output: 1108.65 toks/s]\rProcessed prompts:  25%|██▌       | 24/96 [11:52<35:40, 29.72s/it, est. speed input: 4.46 toks/s, output: 465.22 toks/s]\rProcessed prompts:  75%|███████▌  | 72/96 [11:58<03:42,  9.26s/it, est. speed input: 11.34 toks/s, output: 1325.42 toks/s]"}
#188
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 97
Wall time 11.0 seconds
Output:
#189
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "97",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 41.8 seconds
Output:
{"chunk_id":"dd4885","wall_time_seconds":60.000956547,"session_id":17991,"original_token_count":177,"output":"\rProcessed prompts:  88%|████████▊ | 84/96 [12:33<01:26,  7.19s/it, est. speed input: 17.44 toks/s, output: 1735.74 toks/s]\rProcessed prompts:  71%|███████▏  | 60/84 [12:41<03:45,  9.41s/it, est. speed input: 6.80 toks/s, output: 1449.02 toks/s]\rProcessed prompts:  86%|████████▌ | 72/84 [12:54<01:18,  6.58s/it, est. speed input: 8.77 toks/s, output: 1743.98 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [12:54<00:00,  6.58s/it, est. speed input: 11.28 toks/s, output: 2051.33 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [12:54<00:00,  9.22s/it, est. speed input: 11.28 toks/s, output: 2051.33 toks/s]\n"}
#190
narration

One evaluation shard is complete and another is nearly done; two slower shards remain. After this result, I’ll use the remaining budget for a linear-decay run and at least one repeated full evaluation of the strongest method, since a single run carries roughly 1.5 points of noise.

#191
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 98
Wall time 11.0 seconds
Output:
#192
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "98",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 43.5 seconds
Output:
{"chunk_id":"41ca7a","wall_time_seconds":60.001039586,"session_id":17991,"original_token_count":140,"output":"\rProcessed prompts: 100%|██████████| 96/96 [13:20<00:00,  6.16s/it, est. speed input: 18.95 toks/s, output: 2107.83 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:20<00:00,  6.16s/it, est. speed input: 18.95 toks/s, output: 2107.83 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:20<00:00,  8.34s/it, est. speed input: 18.95 toks/s, output: 2107.83 toks/s]\n\rProcessed prompts:  29%|██▊       | 24/84 [14:01<34:18, 34.30s/it, est. speed input: 2.72 toks/s, output: 436.00 toks/s]"}
#193
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 99
Wall time 11.0 seconds
Output:
#194
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "99",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 42.9 seconds
Output:
{"chunk_id":"15a944","wall_time_seconds":60.00212509,"session_id":17991,"original_token_count":133,"output":"\rProcessed prompts:  43%|████▎     | 36/84 [14:36<15:59, 19.98s/it, est. speed input: 3.98 toks/s, output: 758.44 toks/s]\rProcessed prompts:  38%|███▊      | 36/96 [15:03<23:25, 23.42s/it, est. speed input: 4.65 toks/s, output: 673.12 toks/s]\rProcessed prompts:  57%|█████▋    | 48/84 [15:11<07:55, 13.22s/it, est. speed input: 5.41 toks/s, output: 1072.52 toks/s]\rProcessed prompts:  71%|███████▏  | 60/84 [15:15<03:26,  8.59s/it, est. speed input: 6.68 toks/s, output: 1464.55 toks/s]"}
#195
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:17991,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 100
Wall time 11.0 seconds
Output:
#196
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 35.9 seconds
Output:
Warning: truncated output (original token count: 457)
Total output lines: 1

{"chunk_id":"0eb5ad","wall_time_seconds":55.559779809,"exit_code":0,"original_token_count":422,"output":"Warning: truncated output (original token count: 422)\nTotal output lines: 5\n\n\rProcessed prompts:  50%|█████     | 48/96 [15:25<11:54, 14.88s/it, est. speed input: 6.15 toks/s, output: 926.05 toks/s]\rProcessed prompts:  86%|████████▌ | 72/84 [15:30<01:13,  6.09s/it, est. speed input: 7.85 toks/s, output: 1727.21 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:35<00:00,  4.23s/it, est. speed input: 10.63 toks/s, output: 2053.87 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:35<00:00,  4.23s/it, est. speed input: 10.63 toks/s, output: 2053.87 toks/s]\rProcessed prompts: 100%|█████████…57 tokens truncated…6:27<03:07,  7.83s/it, est. speed input: 9.80 toks/s, output: 1392.89 toks/s]\rProcessed prompts:  88%|████████▊ | 84/96 [16:28<01:03,  5.31s/it, est. speed input: 11.10 toks/s, output: 1741.75 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:33<00:00,  3.73s/it, est. speed input: 13.17 toks/s, output: 2111.73 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:33<00:00,  3.73s/it, est. speed input: 13.17 toks/s, output: 2111.73 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:33<00:00, 10.34s/it, est. speed input: 13.17 toks/s, output: 2111.73 toks/s]\n\nFINAL avg@12 accuracy: 56.11% (202/360 solutions correct on 30 problems)\nResults written to /workspace/eval_results/dev_eval_1785736095.json\n"}
#197
narration

Early-token credit scored 56.11% (202/360), a +1.89-point single-run gain over the 54.22 five-run baseline and +4.17 over the rejected candidate. This is large enough to pursue. I’m now rerunning the same method with the supplied baseline log’s linear decay; this isolates scheduler dynamics while preserving the demonstrated method improvement.

#198
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n # max_completion_length=1024        # <= 4096\\n+lr_scheduler_type=linear\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/earlycredit_linear PORT=12952 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":1000,\"max_output_tokens\":1000});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{}{"chunk_id":"ed890b","wall_time_seconds":1.00082904,"session_id":89905,"original_token_count":0,"output":""}
#199
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 102
Wall time 11.0 seconds
Output:
#200
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "102",
  "yield_time_ms": 60000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 45.2 seconds
Output:
Warning: truncated output (original token count: 563)
Total output lines: 1

{"chunk_id":"d90839","wall_time_seconds":60.001025025,"session_id":89905,"original_token_count":4148,"output":"Warning: truncated output (original token count: 4148)\nTotal output lines: 197\n\n<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.\n\n================================================================================\nRUN CONFIGURATION\n================================================================================\nWandB Run Name: capped_seed42_lr5e-6_bs32\nOutput Directory: /workspace/runs/earlycredit_linear/capped_seed42\n================================================================================\n\n\n================================================================================\nLoading model with dtype: torch.bfloat16\nUsing attention implementation: flash_attention_2\n================================================================================\n\n<string>:24: FutureWarning: `torch_dtype`…63 tokens truncated…00:16<12:55,  7.91s/it]vLLM generation done - elapsed time: 3.05s, prompts: 4, total tokens: 2165, avg length: 541.2, speed: 710.6 tok/s\nvLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 2489, avg length: 622.2, speed: 779.5 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3480, avg length: 870.0, speed: 1067.4 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3824, avg length: 956.0, speed: 1164.5 tok/s\nvLLM generation done - elapsed time: 3.21s, prompts: 4, total tokens: 3036, avg length: 759.0, speed: 945.8 tok/s\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 3323, avg length: 830.8, speed: 1029.0 tok/s\nvLLM generation done - elapsed time: 3.24s, prompts: 4, total tokens: 2940, avg length: 735.0, speed: 907.5 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3540, avg length: 885.0, speed: 1082.2 tok/s\n\r  3%|▎         | 3/100 [00:23<12:40,  7.84s/it]"}