Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#501
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "257",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 29.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 2982)
Total output lines: 192

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2918, avg length: 729.5, speed: 795.2 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3052, avg length: 763.0, speed: 824.1 tok/s
vLLM gener…27 tokens truncated…9.7 tok/s

 24%|██▍       | 24/100 [03:33<11:15,  8.89s/it]
                                                
{'loss': -0.017, 'grad_norm': 0.053049515932798386, 'learning_rate': 3.85e-06, 'on_policy_loss': -0.017, 'rollout_accuracy': 0.1875, 'epoch': 0.09}

 24%|██▍       | 24/100 [03:33<11:15,  8.89s/it]
#502
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 258
Wall time 11.0 seconds
Output:
#503
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "258",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 29.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 1854)
Total output lines: 61

vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2876, avg length: 719.0, speed: 793.7 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3075, avg length: 768.8, speed: 830.0 tok/s
vLLM genera…27 tokens truncated…d time: 3.73s, prompts: 4, total tokens: 3153, avg length: 788.2, speed: 844.8 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4084, avg length: 1021.0, speed: 1080.2 tok/s
vLLM generation done - elapsed time: 3.80s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1078.5 tok/s
#504
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 259
Wall time 11.0 seconds
Output:
#505
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "259",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 20.1 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9

Warning: truncated output (original token count: 1945)
Total output lines: 68

vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3156, avg length: 789.0, speed: 848.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3519, avg length: 879.8, speed: 939.7 tok/s
vLLM genera…27 tokens truncated…h: 770.5, speed: 833.1 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3807, avg length: 951.8, speed: 1007.6 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1081.6 tok/s

 37%|███▋      | 37/100 [05:29<09:24,  8.96s/it]
#506
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 260
Wall time 11.0 seconds
Output:
#507
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "260",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 30.1 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9

Warning: truncated output (original token count: 1741)
Total output lines: 58

vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3539, avg length: 884.8, speed: 948.1 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3576, avg length: 894.0, speed: 954.4 tok/s
vLLM genera…27 tokens truncated…th: 770.0, speed: 836.4 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3719, avg length: 929.8, speed: 986.7 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3525, avg length: 881.2, speed: 934.2 tok/s

 43%|████▎     | 43/100 [06:23<08:33,  9.01s/it]
#508
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 261
Wall time 11.0 seconds
Output:
#509
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "261",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 32.3 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 9

Warning: truncated output (original token count: 1997)
Total output lines: 66

vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3625, avg length: 906.2, speed: 971.4 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3996, avg length: 999.0, speed: 1059.3 tok/s
vLLM gener…27 tokens truncated…h: 868.0, speed: 930.8 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3653, avg length: 913.2, speed: 972.7 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3888, avg length: 972.0, speed: 1028.4 tok/s

 50%|█████     | 50/100 [07:26<07:29,  8.99s/it]
#510
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 262
Wall time 11.0 seconds
Output:
#511
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "262",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 30.2 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 6

Warning: truncated output (original token count: 2942)
Total output lines: 206


                                                
{'loss': -0.0211, 'grad_norm': 0.04648179933428764, 'learning_rate': 2.55e-06, 'on_policy_loss': -0.0211, 'rollout_accuracy': 0.3125, 'epoch': 0.18}

 50%|█████     | 50/100 [07:2…27 tokens truncated…56%|█████▌    | 56/100 [08:20<06:34,  8.96s/it]
                                                
{'loss': -0.0215, 'grad_norm': 0.055439431220293045, 'learning_rate': 2.25e-06, 'on_policy_loss': -0.0215, 'rollout_accuracy': 0.3125, 'epoch': 0.2}

 56%|█████▌    | 56/100 [08:20<06:34,  8.96s/it]
#512
narration

At halfway, stronger clipping remains stable but substantially changes the optimization regime (loss around −0.021 versus near zero for v1). This confirms the test is materially different; the full AIME score will decide whether that extra style suppression is beneficial or excessive.

#513
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 263
Wall time 11.0 seconds
Output:
#514
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "263",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 26.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 1867)
Total output lines: 61

vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2541, avg length: 635.2, speed: 702.2 tok/s
vLLM generation done - elapsed time: 3.62s, prompts: 4, total tokens: 2588, avg length: 647.0, speed: 714.2 tok/s
vLLM genera…27 tokens truncated…d time: 3.76s, prompts: 4, total tokens: 3972, avg length: 993.0, speed: 1056.8 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3803, avg length: 950.8, speed: 1008.5 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1079.5 tok/s
#515
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 264
Wall time 11.0 seconds
Output:
#516
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "264",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 21.8 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 2000)
Total output lines: 65

vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3014, avg length: 753.5, speed: 808.2 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3775, avg length: 943.8, speed: 1011.5 tok/s
vLLM gener…27 tokens truncated…ed time: 3.74s, prompts: 4, total tokens: 3811, avg length: 952.8, speed: 1019.5 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3858, avg length: 964.5, speed: 1023.7 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3870, avg length: 967.5, speed: 1024.2 tok/s
#517
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 265
Wall time 11.0 seconds
Output:
#518
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "265",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 20.2 seconds
Output:
Warning: truncated output (original token count: 186)
Total output lines: 7

Warning: truncated output (original token count: 2032)
Total output lines: 69

vLLM generation done - elapsed time: 3.61s, prompts: 4, total tokens: 2864, avg length: 716.0, speed: 792.7 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3137, avg length: 784.2, speed: 846.1 tok/s
vLLM genera…26 tokens truncated…█████▌  | 76/100 [11:20<03:35,  8.99s/it]
                                                
{'loss': -0.0245, 'grad_norm': 0.08779964596033096, 'learning_rate': 1.25e-06, 'on_policy_loss': -0.0245, 'rollout_accuracy': 0.1875, 'epoch': 0.28}

 76%|███████▌  | 76/100 [11:20<03:35,  8.99s/it]
#519
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 266
Wall time 11.0 seconds
Output:
#520
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "266",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 27.9 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 1882)
Total output lines: 61

vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3617, avg length: 904.2, speed: 968.1 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3677, avg length: 919.2, speed: 977.3 tok/s
vLLM genera…27 tokens truncated…psed time: 3.74s, prompts: 4, total tokens: 3547, avg length: 886.8, speed: 949.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3517, avg length: 879.2, speed: 939.5 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3969, avg length: 992.2, speed: 1053.8 tok/s
#521
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:160});\ntext(r.output);"
}
exec result
Script running with cell ID 267
Wall time 11.0 seconds
Output:
#522
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "267",
  "yield_time_ms": 50000,
  "max_tokens": 160
}
wait result
Script completed
Wall time 33.3 seconds
Output:
Warning: truncated output (original token count: 187)
Total output lines: 8

Warning: truncated output (original token count: 4557)
Total output lines: 393

vLLM generation done - elapsed time: 3.16s, prompts: 4, total tokens: 2789, avg length: 697.2, speed: 883.9 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3248, avg length: 812.0, speed: 876.7 tok/s
vLLM gener…27 tokens truncated… time: 3.70s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1106.6 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3812, avg length: 953.0, speed: 1027.0 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1089.2 tok/s
#523
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 268
Wall time 11.0 seconds
Output:
#524
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "268",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 29.1 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 12

Warning: truncated output (original token count: 3102)
Total output lines: 245

vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3496, avg length: 874.0, speed: 955.0 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3382, avg length: 845.5, speed: 920.3 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3410, avg length: 852.5, speed: 924.7 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3968, avg length: 992.0, speed: 1064.3 tok/s

 90%|█████████ | 90/100 [13:24<01:28,  8…27 tokens truncated…licy_loss': -0.029, 'rollout_accuracy': 0.0625, 'epoch': 0.35}

 96%|█████████▌| 96/100 [14:18<00:35,  8.99s/it]vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3804, avg length: 951.0, speed: 1015.7 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4043, avg length: 1010.8, speed: 1073.1 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1080.9 tok/s
vLLM generation done - elapsed time: 3.82s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1071.3 tok/s
#525
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89779,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script completed
Wall time 8.7 seconds
Output:
Warning: truncated output (original token count: 1399)
Total output lines: 37

vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1088.5 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3908, avg length: 977.0, speed: 1038.2 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4051, avg length: 1012.8, speed: 1070.8 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1082.0 tok/s

 97%|█████████▋| 97/100 [14:27<00:26,  8.98s/it]vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3056, avg length: 764.0, speed: 824.8 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3616, avg length: 904.0, speed: 972.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3856, avg length: 964.0, speed: 1030.1 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1084.4 tok/s
v…899 tokens truncated…s. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 08:21:46.627373896 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 08:21:46.191261281 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 08:21:46.191628873 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/frontier_v4_clip002/capped_seed42/checkpoint-100