Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#151
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "85",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 34.2 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11

Warning: truncated output (original token count: 1939)
Total output lines: 67

vLLM generation done - elapsed time: 3.60s, prompts: 4, total tokens: 3200, avg length: 800.0, speed: 888.7 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3555, avg length: 888.8, speed: 968.4 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3162, avg length: 790.5, speed: 862.9 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3772, …27 tokens truncated…    | 51/100 [07:31<07:12,  8.83s/it]vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3444, avg length: 861.0, speed: 945.5 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3753, avg length: 938.2, speed: 1015.0 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3479, avg length: 869.8, speed: 941.0 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.2 tok/s
#152
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 86
Wall time 11.0 seconds
Output:
#153
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "86",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.4 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11

Warning: truncated output (original token count: 3111)
Total output lines: 193

vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3593, avg length: 898.2, speed: 981.5 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1113.2 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3493, avg length: 873.2, speed: 949.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 339…27 tokens truncated…otal tokens: 4096, avg length: 1024.0, speed: 1107.0 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1104.5 tok/s

 58%|█████▊    | 58/100 [08:33<06:09,  8.79s/it]
                                                
{'loss': -0.0013, 'grad_norm': 0.055766116827726364, 'learning_rate': 2.15e-06, 'on_policy_loss': -0.0013, 'rollout_accuracy': 0.1875, 'epoch': 0.21}

 58%|█████▊    | 58/100 [08:33<06:09,  8.79s/it]
#154
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 87
Wall time 11.0 seconds
Output:
#155
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "87",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.4 seconds
Output:
Warning: truncated output (original token count: 276)
Total output lines: 11

Warning: truncated output (original token count: 1874)
Total output lines: 61

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1115.2 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3693, avg length: 923.2, speed: 999.7 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3811, avg length: 952.8, speed: 1027.6 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 409…26 tokens truncated…   | 64/100 [09:26<05:16,  8.79s/it]vLLM generation done - elapsed time: 3.63s, prompts: 4, total tokens: 3136, avg length: 784.0, speed: 862.9 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3794, avg length: 948.5, speed: 1024.9 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3826, avg length: 956.5, speed: 1031.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3711, avg length: 927.8, speed: 992.2 tok/s
#156
narration

The repeat has passed step 60 with the same stable loss/gradient profile as run one. Its on-policy correctness windows are somewhat lower, reinforcing that the decisive evidence must be the frozen full evaluation rather than training-rollout accuracy.

#157
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 88
Wall time 11.0 seconds
Output:
#158
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "88",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.3 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11

Warning: truncated output (original token count: 1760)
Total output lines: 57

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1117.3 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3464, avg length: 866.0, speed: 943.4 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3454, avg length: 863.5, speed: 933.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4096…27 tokens truncated…   | 70/100 [10:19<04:23,  8.77s/it]vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3817, avg length: 954.2, speed: 1038.8 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3555, avg length: 888.8, speed: 960.7 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3968, avg length: 992.0, speed: 1073.5 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4041, avg length: 1010.2, speed: 1083.2 tok/s
#159
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 89
Wall time 11.0 seconds
Output:
#160
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "89",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.2 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 13

Warning: truncated output (original token count: 1951)
Total output lines: 67

vLLM generation done - elapsed time: 3.61s, prompts: 4, total tokens: 3097, avg length: 774.2, speed: 857.3 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2993, avg length: 748.2, speed: 815.9 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1105.9 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4096…27 tokens truncated… length: 1001.2, speed: 1077.7 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3337, avg length: 834.2, speed: 912.3 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3479, avg length: 869.8, speed: 947.5 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3494, avg length: 873.5, speed: 945.4 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4068, avg length: 1017.0, speed: 1091.0 tok/s
#161
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 90
Wall time 11.0 seconds
Output:
#162
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "90",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.7 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11

Warning: truncated output (original token count: 1784)
Total output lines: 58


 77%|███████▋  | 77/100 [11:21<03:22,  8.81s/it]vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3544, avg length: 886.0, speed: 968.2 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3645, avg length: 911.2, speed: 995.2 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3598, avg length: 899.5, speed: 975.5 tok/s
vLLM genera…27 tokens truncated…ime: 3.63s, prompts: 4, total tokens: 3157, avg length: 789.2, speed: 870.0 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1104.0 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3714, avg length: 928.5, speed: 1000.4 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4056, avg length: 1014.0, speed: 1091.8 tok/s

 83%|████████▎ | 83/100 [12:14<02:30,  8.83s/it]
#163
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 91
Wall time 11.0 seconds
Output:
#164
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "91",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.5 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11

Warning: truncated output (original token count: 4217)
Total output lines: 323

vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3281, avg length: 820.2, speed: 900.2 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1106.3 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.9 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 3…27 tokens truncated…ns: 3852, avg length: 963.0, speed: 1032.6 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.6 tok/s

 90%|█████████ | 90/100 [13:15<01:28,  8.84s/it]
                                                
{'loss': -0.0038, 'grad_norm': 0.10706433653831482, 'learning_rate': 5.5e-07, 'on_policy_loss': -0.0038, 'rollout_accuracy': 0.0625, 'epoch': 0.33}

 90%|█████████ | 90/100 [13:16<01:28,  8.84s/it]
#165
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 92
Wall time 11.0 seconds
Output:
#166
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "92",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 35.6 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 12

Warning: truncated output (original token count: 2934)
Total output lines: 214

vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3552, avg length: 888.0, speed: 963.2 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3733, avg length: 933.2, speed: 1003.7 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1098.3 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 40…27 tokens truncated…e: 3.68s, prompts: 4, total tokens: 3785, avg length: 946.2, speed: 1028.9 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3694, avg length: 923.5, speed: 1003.3 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3846, avg length: 961.5, speed: 1027.5 tok/s

 97%|█████████▋| 97/100 [14:17<00:26,  8.82s/it]
#167
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:5669,chars:\"\",yield_time_ms:50000,max_output_tokens:900});\ntext(r.output);"
}
exec result
Script running with cell ID 93
Wall time 11.0 seconds
Output:
#168
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "93",
  "yield_time_ms": 50000,
  "max_tokens": 900
}
wait result
Script completed
Wall time 15.0 seconds
Output:
Warning: truncated output (original token count: 927)
Total output lines: 24

Warning: truncated output (original token count: 1265)
Total output lines: 33

vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3991, avg length: 997.8, speed: 1084.3 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3380, avg length: 845.0, speed: 918.8 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3896, avg length: 974.0, speed: 1053.0 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1090.9 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3259, avg length: 814.8, speed: 888.8 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3764, avg length: 941.0, speed: 1019.5 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3770, avg length: 942.5, speed: 1017.5 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 4053, avg length: 1013.2, speed: 1089.0 tok/s

 98%|█████████▊| 98/100 [14:26<00:17,  8.81s/it]
                                                
{'loss': -0.0024, 'grad_norm': 0.08249122649431229, 'learning_rate': 1.5000000000000002e-07, 'on_policy_loss': -0.0024, 'rollout_accuracy': 0.3125, 'epoch': 0.36}

 98%|█████████▊| 98/100 [14:26<00:17,  8.81s/it]vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3140, avg length: 785.0, speed: 863.1 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 4006, avg length: 1001.5, speed: 1080.7 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3633, avg length: 908.2, speed: 980.2 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0…27 tokens truncated…00:00,  8.83s/it]
                                                 
{'loss': -0.0032, 'grad_norm': 0.05919903144240379, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': -0.0032, 'rollout_accuracy': 0.125, 'epoch': 0.36}

100%|██████████| 100/100 [14:44<00:00,  8.83s/it]
                                                 
{'train_runtime': 887.4829, 'train_samples_per_second': 3.606, 'train_steps_per_second': 0.113, 'train_loss': 0.001642743945121765, 'epoch': 0.36}

100%|██████████| 100/100 [14:47<00:00,  8.83s/it]
100%|██████████| 100/100 [14:47<00:00,  8.87s/it]
[rank0]:[W803 05:40:02.412210858 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 05:40:04.885965380 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 05:40:04.885972563 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 05:40:04.926445576 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/frontier_v1_rep2/capped_seed42/checkpoint-100
#169
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/frontier_v1_rep2/capped_seed42/checkpoint-100 /workspace/eval_results/frontier_v1_rep2_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":800,\"tty\":true});\ntext(JSON.stringify(r));"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"228f9c","wall_time_seconds":1.002491768,"session_id":35310,"original_token_count":0,"output":""}
#170
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:400});\ntext(r.output);"
}
exec result
Script running with cell ID 95
Wall time 11.0 seconds
Output:
#171
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "95",
  "yield_time_ms": 50000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 35.5 seconds
Output:
Warning: truncated output (original token count: 427)
Total output lines: 18

Warning: truncated output (original token count: 5926)
Total output lines: 146

Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...
`torch_dtype` is deprecated! Use `dtype` instead!

Loading checkpoint shards:   0%|                         | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards: 100%|█████████████████| 2/2 [00:00<00:00, 46.50it/s]
Generating with data_parallel_size=4 (TP=1 per engine) ...
INFO 08-03 05:40:47 [__init__.py:216] Automatically detected platform cuda.
INFO 08-03 05:40:47 [__init__.py:216] Automatically detected platform cuda.
INFO 08-03 05:40:47 [__init__.py:216] Automatically detected platform cuda.
INFO 08-03 05:40:47 [__init__.py:216] Automatically detected platform cuda.
INFO 08-03 0…27 tokens truncated…s]

Processed prompts:   0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,INFO 08-03 05:41:07 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|                                   | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 8/8 [00:00<00:00, 239.67it/s]

Processed prompts:   0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,INFO 08-03 05:41:07 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|                                   | 0/7 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 7/7 [00:00<00:00, 231.51it/s]

Processed prompts:   0%| | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,
#172
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 96
Wall time 11.0 seconds
Output:
#173
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "96",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 35.3 seconds
Output:
#174
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 97
Wall time 11.0 seconds
Output:
#175
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "97",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 34.3 seconds
Output: