Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#226
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "123",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 32.0 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 2402)
Total output lines: 164

vLLM generation done - elapsed time: 3.61s, prompts: 4, total tokens: 3228, avg length: 807.0, speed: 893.6 tok/s
vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2917, avg length: 729.2, speed: 804.9 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2999, avg length: 74…27 tokens truncated…vLLM generation done - elapsed time: 3.59s, prompts: 4, total tokens: 2584, avg length: 646.0, speed: 720.5 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3506, avg length: 876.5, speed: 946.0 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4056, avg length: 1014.0, speed: 1078.7 tok/s

 13%|█▎        | 13/100 [01:58<13:41,  9.44s/it]
#227
narration

The v2 refinement is now training: dense OPSD still prioritizes early decision points, while the verified term now anchors every token of successful traces uniformly, including the late final answer. This is a narrow one-line methodological change, so its full score will be directly interpretable against the two v1 runs.

#228
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 124
Wall time 11.0 seconds
Output:
#229
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "124",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.2 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 1838)
Total output lines: 61

vLLM generation done - elapsed time: 3.63s, prompts: 4, total tokens: 2740, avg length: 685.0, speed: 754.0 tok/s
vLLM generation done - elapsed time: 3.65s, prompts: 4, total tokens: 2996, avg length: 749.0, speed: 820.7 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3332, avg length: 833…27 tokens truncated…l tokens: 2828, avg length: 707.0, speed: 778.2 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3127, avg length: 781.8, speed: 841.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3430, avg length: 857.5, speed: 917.9 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3707, avg length: 926.8, speed: 983.6 tok/s
#230
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 125
Wall time 11.0 seconds
Output:
#231
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "125",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 31.4 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 3026)
Total output lines: 188

vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3165, avg length: 791.2, speed: 848.9 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3297, avg length: 824.2, speed: 882.2 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3400, avg length: 85…27 tokens truncated…==========================================================================


 26%|██▌       | 26/100 [03:55<11:05,  9.00s/it]
                                                
{'loss': 0.0023, 'grad_norm': 0.07633288949728012, 'learning_rate': 3.7500000000000005e-06, 'on_policy_loss': 0.0023, 'rollout_accuracy': 0.4375, 'epoch': 0.09}

 26%|██▌       | 26/100 [03:55<11:05,  9.00s/it]
#232
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 126
Wall time 11.0 seconds
Output:
#233
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "126",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.9 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 1979)
Total output lines: 66

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3279, avg length: 819.8, speed: 892.3 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3176, avg length: 794.0, speed: 857.0 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3361, avg length: 840…27 tokens truncated…eneration done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1101.0 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3568, avg length: 892.0, speed: 958.5 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4054, avg length: 1013.5, speed: 1087.8 tok/s

 33%|███▎      | 33/100 [04:57<09:55,  8.88s/it]
#234
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 127
Wall time 11.0 seconds
Output:
#235
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "127",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 33.6 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 9

Warning: truncated output (original token count: 2057)
Total output lines: 67

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3563, avg length: 890.8, speed: 971.2 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3403, avg length: 850.8, speed: 925.7 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3772, avg length: 943…27 tokens truncated…ed time: 3.75s, prompts: 4, total tokens: 3923, avg length: 980.8, speed: 1045.7 tok/s

 40%|████      | 40/100 [06:01<09:06,  9.11s/it]
                                                
{'loss': 0.0001, 'grad_norm': 0.05959784612059593, 'learning_rate': 3.05e-06, 'on_policy_loss': 0.0001, 'rollout_accuracy': 0.5625, 'epoch': 0.14}

 40%|████      | 40/100 [06:01<09:06,  9.11s/it]
#236
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 128
Wall time 11.0 seconds
Output:
#237
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "128",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 25.4 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 2062)
Total output lines: 72

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2736, avg length: 684.0, speed: 746.0 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3037, avg length: 759.2, speed: 824.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3778, avg length: 944…27 tokens truncated…eneration done - elapsed time: 3.73s, prompts: 4, total tokens: 3466, avg length: 866.5, speed: 929.9 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3628, avg length: 907.0, speed: 965.4 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1084.5 tok/s

 47%|████▋     | 47/100 [07:03<07:54,  8.95s/it]
#238
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 129
Wall time 11.0 seconds
Output:
#239
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.3 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 9

Warning: truncated output (original token count: 3244)
Total output lines: 190

vLLM generation done - elapsed time: 3.65s, prompts: 4, total tokens: 2934, avg length: 733.5, speed: 803.1 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3747, avg length: 936.8, speed: 1003.3 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3709, avg length: 9…27 tokens truncated… time: 3.76s, prompts: 4, total tokens: 3864, avg length: 966.0, speed: 1028.0 tok/s

 54%|█████▍    | 54/100 [08:06<06:53,  8.99s/it]
                                                
{'loss': 0.0, 'grad_norm': 0.055839866399765015, 'learning_rate': 2.35e-06, 'on_policy_loss': 0.0, 'rollout_accuracy': 0.3125, 'epoch': 0.2}

 54%|█████▍    | 54/100 [08:06<06:53,  8.99s/it]
#240
narration

v2 is past halfway and remains stable; its logged rollout accuracy is not collapsing late as sharply as v1 did in the corresponding windows, though that statistic is too noisy to count as evidence. The full fixed evaluation remains the selection criterion.

#241
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 130
Wall time 11.0 seconds
Output:
#242
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "130",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 31.8 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 9

Warning: truncated output (original token count: 1751)
Total output lines: 58

vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3743, avg length: 935.8, speed: 1014.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1095.8 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: …27 tokens truncated….73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.1 tok/s

 60%|██████    | 60/100 [08:59<05:53,  8.83s/it]
                                                
{'loss': -0.0038, 'grad_norm': 0.0669080913066864, 'learning_rate': 2.05e-06, 'on_policy_loss': -0.0038, 'rollout_accuracy': 0.0625, 'epoch': 0.22}

 60%|██████    | 60/100 [08:59<05:53,  8.83s/it]
#243
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 131
Wall time 11.0 seconds
Output:
#244
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "131",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 30.3 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 2072)
Total output lines: 72

vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3389, avg length: 847.2, speed: 931.1 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3635, avg length: 908.8, speed: 981.9 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3940, avg length: 985…27 tokens truncated…eration done - elapsed time: 3.66s, prompts: 4, total tokens: 3272, avg length: 818.0, speed: 893.4 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3533, avg length: 883.2, speed: 962.9 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3559, avg length: 889.8, speed: 956.5 tok/s

 67%|██████▋   | 67/100 [10:01<04:51,  8.83s/it]
#245
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 132
Wall time 11.0 seconds
Output:
#246
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "132",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 27.5 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 1758)
Total output lines: 58

vLLM generation done - elapsed time: 3.65s, prompts: 4, total tokens: 2924, avg length: 731.0, speed: 800.9 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3518, avg length: 879.5, speed: 959.9 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3406, avg length: 851…27 tokens truncated…done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1099.9 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1096.5 tok/s

 73%|███████▎  | 73/100 [10:53<03:56,  8.77s/it]
#247
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 133
Wall time 11.0 seconds
Output:
#248
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "133",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 35.1 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 1766)
Total output lines: 58

vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3497, avg length: 874.2, speed: 950.3 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3897, avg length: 974.2, speed: 1052.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 4096, avg length: 10…27 tokens truncated…ion done - elapsed time: 3.68s, prompts: 4, total tokens: 3069, avg length: 767.2, speed: 833.1 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3764, avg length: 941.0, speed: 1021.3 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3796, avg length: 949.0, speed: 1026.3 tok/s

 79%|███████▉  | 79/100 [11:46<03:04,  8.79s/it]
#249
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 134
Wall time 11.0 seconds
Output:
#250
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "134",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.5 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10

Warning: truncated output (original token count: 3273)
Total output lines: 196

vLLM generation done - elapsed time: 3.56s, prompts: 4, total tokens: 2756, avg length: 689.0, speed: 773.1 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3637, avg length: 909.2, speed: 979.3 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3917, avg length: 97…27 tokens truncated…okens: 3277, avg length: 819.2, speed: 895.9 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3533, avg length: 883.2, speed: 961.1 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3788, avg length: 947.0, speed: 1028.1 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.8 tok/s