Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#301
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "166",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 848)
Total output lines: 26

vLLM generation done - elapsed time: 6.17s, prompts: 4, total tokens: 3356, avg le…26 tokens truncated…5541, avg length: 1385.2, speed: 849.5 tok/s
vLLM generation done - elapsed time: 6.73s, prompts: 4, total tokens: 6526, avg length: 1631.5, speed: 969.0 tok/s
#302
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"tail -c 1800 /workspace/runs/robust_kl_2048/train_seed42.log | rg -o \\\"[0-9]+%\\\\|[^\\\\r]*|\\\\{'loss'.*\\\" | tail -3\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":500});\ntext(r.output);\n"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
{'loss': 0.0121, 'grad_norm': 0.06966597586870193, 'learning_rate': 4.350000000000001e-06, 'on_policy_loss': 0.0121, 'epoch': 0.12}
14%|█▍        | 14/100 [03:35<22:01, 15.37s/it]vLLM generation done - elapsed time: 4.33s, prompts: 4, total tokens: 3878, avg length: 969.5, speed: 896.2 tok/s
15%|█▌        | 15/100 [03:50<21:40, 15.30s/it]
#303
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 168
Wall time 11.0 seconds
Output:
#304
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "168",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 15.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 856)
Total output lines: 31

vLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 3746, avg le…26 tokens truncated…5696, avg length: 1424.0, speed: 861.5 tok/s
vLLM generation done - elapsed time: 6.65s, prompts: 4, total tokens: 5974, avg length: 1493.5, speed: 898.6 tok/s
#305
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 169
Wall time 11.0 seconds
Output:
#306
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "169",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 543)
Total output lines: 17

vLLM generation done - elapsed time: 4.34s, prompts: 4, total tokens: 3630, avg le…26 tokens truncated…5613, avg length: 1403.2, speed: 842.9 tok/s
vLLM generation done - elapsed time: 6.69s, prompts: 4, total tokens: 6534, avg length: 1633.5, speed: 977.1 tok/s
#307
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 170
Wall time 11.0 seconds
Output:
#308
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "170",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 3001)
Total output lines: 221

vLLM generation done - elapsed time: 6.38s, prompts: 4, total tokens: 4773, avg …27 tokens truncated…123, 'grad_norm': 0.060684684664011, 'learning_rate': 3.95e-06, 'on_policy_loss': 0.0123, 'epoch': 0.19}

 22%|██▏       | 22/100 [05:38<20:02, 15.41s/it]
#309
narration

By step 22 the robust KL remains positive (0.0123) and the gradient norm has settled to 0.061, so the student is matching the privileged teacher instead of entering the baseline estimator’s negative-loss regime. The run still has to prove that this cleaner objective translates to AIME accuracy.

#310
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 171
Wall time 11.0 seconds
Output:
#311
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "171",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 660)
Total output lines: 21

vLLM generation done - elapsed time: 4.71s, prompts: 4, total tokens: 4139, avg le…26 tokens truncated…5251, avg length: 1312.8, speed: 799.0 tok/s
vLLM generation done - elapsed time: 6.79s, prompts: 4, total tokens: 6726, avg length: 1681.5, speed: 990.4 tok/s
#312
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 172
Wall time 11.0 seconds
Output:
#313
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "172",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 747)
Total output lines: 28

vLLM generation done - elapsed time: 6.24s, prompts: 4, total tokens: 6168, avg le…26 tokens truncated…tion done - elapsed time: 6.79s, prompts: 4, total tokens: 7197, avg length: 1799.2, speed: 1059.2 tok/s

 27%|██▋       | 27/100 [06:55<18:43, 15.39s/it]
#314
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"python - <<'PY'\\nimport json\\nb=json.load(open('/workspace/baseline/logs/official_eval_result.json'))['results']\\nr=json.load(open('/workspace/eval_results/refined_long_full.json'))['results']\\nbm={x['problem_id']:x['num_correct'] for x in b}\\nprint('id base refined delta')\\nfor x in r:\\n i=x['problem_id']; d=x['num_correct']-bm[i]\\n if d: print(i,bm[i],x['num_correct'],f'{d:+}')\\nPY\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);\n"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
id base refined delta
65 6 7 +1
66 11 9 -2
67 12 11 -1
70 5 9 +4
74 10 8 -2
78 5 7 +2
79 12 11 -1
80 2 8 +6
82 7 9 +2
84 11 12 +1
#315
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 174
Wall time 11.0 seconds
Output:
#316
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "174",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 659)
Total output lines: 21

vLLM generation done - elapsed time: 5.23s, prompts: 4, total tokens: 3891, avg le…26 tokens truncated…6735, avg length: 1683.8, speed: 993.0 tok/s
vLLM generation done - elapsed time: 6.81s, prompts: 4, total tokens: 6784, avg length: 1696.0, speed: 995.6 tok/s
#317
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 175
Wall time 11.0 seconds
Output:
#318
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "175",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 920)
Total output lines: 32

vLLM generation done - elapsed time: 6.53s, prompts: 4, total tokens: 5262, avg le…26 tokens truncated…656, avg length: 1664.0, speed: 990.6 tok/s
vLLM generation done - elapsed time: 6.81s, prompts: 4, total tokens: 7248, avg length: 1812.0, speed: 1065.0 tok/s
#319
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 176
Wall time 11.0 seconds
Output:
#320
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "176",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 663)
Total output lines: 21

vLLM generation done - elapsed time: 6.41s, prompts: 4, total tokens: 4794, avg le…26 tokens truncated…6092, avg length: 1523.0, speed: 916.2 tok/s
vLLM generation done - elapsed time: 6.69s, prompts: 4, total tokens: 6271, avg length: 1567.8, speed: 936.8 tok/s
#321
narration

The validated 57.22% model’s gains are concentrated on harder AIME items (notably +4 and +6 correct samples on two low-baseline problems), which supports the difficulty/format-matched replay hypothesis rather than a mere formatting-rate effect.

#322
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 177
Wall time 11.0 seconds
Output:
#323
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "177",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 749)
Total output lines: 27


 35%|███▌      | 35/100 [08:58<16:36, 15.34s/it]vLLM generation done - el…26 tokens truncated…5905, avg length: 1476.2, speed: 898.4 tok/s
vLLM generation done - elapsed time: 6.74s, prompts: 4, total tokens: 6413, avg length: 1603.2, speed: 951.7 tok/s
#324
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 178
Wall time 11.0 seconds
Output:
#325
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "178",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 5.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 850)
Total output lines: 26

vLLM generation done - elapsed time: 6.01s, prompts: 4, total tokens: 4484, avg le…26 tokens truncated…5672, avg length: 1418.0, speed: 851.5 tok/s
vLLM generation done - elapsed time: 6.73s, prompts: 4, total tokens: 6136, avg length: 1534.0, speed: 911.8 tok/s