Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#326
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 179
Wall time 11.0 seconds
Output:
#327
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "179",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 862)
Total output lines: 31

vLLM generation done - elapsed time: 4.54s, prompts: 4, total tokens: 4717, avg le…26 tokens truncated…5826, avg length: 1456.5, speed: 875.0 tok/s
vLLM generation done - elapsed time: 6.75s, prompts: 4, total tokens: 6333, avg length: 1583.2, speed: 938.2 tok/s
#328
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 180
Wall time 11.0 seconds
Output:
#329
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "180",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 922)
Total output lines: 32

vLLM generation done - elapsed time: 4.09s, prompts: 4, total tokens: 4069, avg le…26 tokens truncated…6035, avg length: 1508.8, speed: 911.1 tok/s
vLLM generation done - elapsed time: 6.76s, prompts: 4, total tokens: 6496, avg length: 1624.0, speed: 961.5 tok/s
#330
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 181
Wall time 11.0 seconds
Output:
#331
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "181",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 796)
Total output lines: 25

vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 2632, avg le…26 tokens truncated…413, avg length: 1603.2, speed: 951.0 tok/s
vLLM generation done - elapsed time: 6.80s, prompts: 4, total tokens: 7074, avg length: 1768.5, speed: 1040.7 tok/s
#332
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 182
Wall time 11.0 seconds
Output:
#333
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "182",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 810)
Total output lines: 29

vLLM generation done - elapsed time: 5.25s, prompts: 4, total tokens: 4495, avg le…26 tokens truncated…744299337267876, 'learning_rate': 2.4500000000000003e-06, 'on_policy_loss': 0.0113, 'epoch': 0.45}

 52%|█████▏    | 52/100 [13:19<12:20, 15.43s/it]
#334
narration

At the midpoint, robust KL is still positive (0.0113) with ordinary gradient norms. The model is not trivially collapsing the teacher/student gap to zero, so the capped objective continues supplying signal across all 100 steps.

#335
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 183
Wall time 11.0 seconds
Output:
#336
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "183",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.8 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2214)
Total output lines: 116

vLLM generation done - elapsed time: 4.13s, prompts: 4, total tokens: 3739, avg …27 tokens truncated…5371, avg length: 1342.8, speed: 809.7 tok/s
vLLM generation done - elapsed time: 6.65s, prompts: 4, total tokens: 5901, avg length: 1475.2, speed: 887.2 tok/s
#337
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 184
Wall time 11.0 seconds
Output:
#338
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "184",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 5.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 864)
Total output lines: 31

vLLM generation done - elapsed time: 6.34s, prompts: 4, total tokens: 3946, avg le…26 tokens truncated…5073, avg length: 1268.2, speed: 794.1 tok/s
vLLM generation done - elapsed time: 6.67s, prompts: 4, total tokens: 5828, avg length: 1457.0, speed: 874.1 tok/s
#339
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 185
Wall time 11.0 seconds
Output:
#340
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "185",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 736)
Total output lines: 23

vLLM generation done - elapsed time: 6.63s, prompts: 4, total tokens: 5800, avg le…26 tokens truncated…d_norm': 0.049084246158599854, 'learning_rate': 2.05e-06, 'on_policy_loss': 0.0121, 'epoch': 0.52}

 60%|██████    | 60/100 [15:22<10:12, 15.32s/it]
#341
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 186
Wall time 11.0 seconds
Output:
#342
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "186",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.3 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 854)
Total output lines: 31

vLLM generation done - elapsed time: 6.14s, prompts: 4, total tokens: 4494, avg le…26 tokens truncated…99, avg length: 1774.8, speed: 1038.7 tok/s
vLLM generation done - elapsed time: 6.92s, prompts: 4, total tokens: 8028, avg length: 2007.0, speed: 1159.6 tok/s
#343
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 187
Wall time 11.0 seconds
Output:
#344
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "187",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 680)
Total output lines: 21


 63%|██████▎   | 63/100 [16:07<09:26, 15.32s/it]vLLM generation don…26 tokens truncated…6492, avg length: 1623.0, speed: 973.6 tok/s
vLLM generation done - elapsed time: 6.68s, prompts: 4, total tokens: 6236, avg length: 1559.0, speed: 933.1 tok/s
#345
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 188
Wall time 11.0 seconds
Output:
#346
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "188",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 813)
Total output lines: 29

vLLM generation done - elapsed time: 5.76s, prompts: 4, total tokens: 4863, avg le…26 tokens truncated…7385134100914, 'learning_rate': 1.6500000000000003e-06, 'on_policy_loss': 0.0103, 'epoch': 0.59}

 68%|██████▊   | 68/100 [17:24<08:11, 15.37s/it]
#347
narration

The robust run is at step 68 with loss 0.0103 and no instability. Its longer rollouts often reach 1,500–2,000 tokens, so this test combines corrected robust distillation with materially better coverage of late reasoning and final answers.

#348
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 189
Wall time 11.0 seconds
Output:
#349
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "189",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 8

Warning: truncated output (original token count: 870)
Total output lines: 32

vLLM generation done - elapsed time: 5.50s, prompts: 4, total tokens: 5709, avg le…26 tokens truncated…s_step_70.json
================================================================================


 71%|███████   | 71/100 [18:10<07:25, 15.37s/it]
#350
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 190
Wall time 11.0 seconds
Output: