Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#301
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:400});\ntext(r.output);"
}
exec result
Script running with cell ID 159
Wall time 11.0 seconds
Output:
#302
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "159",
  "yield_time_ms": 50000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 29.8 seconds
Output:

Processed prompts: 100%|█| 96/96 [19:34<00:00,  6.76s/it, est. speed input: 11.
Processed prompts: 100%|█| 96/96 [19:34<00:00,  6.76s/it, est. speed input: 11.
Processed prompts: 100%|█| 96/96 [19:34<00:00, 12.23s/it, est. speed input: 11.
#303
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:400});\ntext(r.output);"
}
exec result
Script running with cell ID 160
Wall time 11.0 seconds
Output:
#304
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "160",
  "yield_time_ms": 50000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 28.0 seconds
Output:

Processed prompts:  86%|▊| 72/84 [21:30<02:53, 14.49s/it, est. speed input: 6.7
#305
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 161
Wall time 11.0 seconds
Output:
#306
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "161",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 10.5 seconds
Output:

Processed prompts: 100%|█| 84/84 [22:15<00:00, 10.97s/it, est. speed input: 7.4
Processed prompts: 100%|█| 84/84 [22:15<00:00, 10.97s/it, est. speed input: 7.4
Processed prompts: 100%|█| 84/84 [22:15<00:00, 15.89s/it, est. speed input: 7.4

FINAL avg@12 accuracy: 55.56% (200/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/frontier_v2_full.json
#307
narration

v2 scored 55.56% (200/360), beating the v1 two-run mean by 0.42 and baseline mean by 1.34. This is promising but within one-run noise, so I’m repeating v2 unchanged. The final choice will use the replicated mean, not the best single checkpoint.

#308
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"cat /workspace/.timer/remaining_secs && SEED=42 OUT_DIR=/workspace/runs/frontier_v2_rep2 PORT=12953 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":500,\"tty\":true});\ntext(JSON.stringify(r));"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"6fba6c","wall_time_seconds":1.001686439,"session_id":92058,"original_token_count":2,"output":"14185\r\n"}
#309
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 163
Wall time 11.0 seconds
Output:
#310
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "163",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 31.0 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 18

Warning: truncated output (original token count: 2569)
Total output lines: 153

<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.

================================================================================
RUN CONFIGURATION
================================================================================
WandB Run Name: capped_seed42_lr5e-6_bs32
Output Directory: /workspace/runs/frontier_v2_rep2/capped_seed42
==…27 tokens truncated… peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0

Loading safetensors checkpoint shards:   0% Completed | 0/2 [00:00<?, ?it/s]

Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00,  2.78it/s]

Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00,  2.78it/s]

#311
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 164
Wall time 11.0 seconds
Output:
#312
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "164",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 33.5 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 7

Warning: truncated output (original token count: 2356)
Total output lines: 68


Capturing CUDA graphs (mixed prefill-decode, PIECEWISE):   0%|          | 0/5 [00:00<?, ?it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE):  40%|████      | 2/5 [00:00<00:00, 17.38it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|████…27 tokens truncated…: 979.7 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 2791, avg length: 697.8, speed: 841.7 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3136, avg length: 784.0, speed: 853.9 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3506, avg length: 876.5, speed: 952.9 tok/s
#313
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 165
Wall time 11.0 seconds
Output:
#314
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "165",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 31.3 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 3979)
Total output lines: 293

vLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 3003, avg length: 750.8, speed: 888.4 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3607, avg length: 901.8, speed: 968.7 tok/s
vLLM generation done - elapsed time: 3.74s, prompt…27 tokens truncated… 823.3 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3546, avg length: 886.5, speed: 967.3 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3054, avg length: 763.5, speed: 827.5 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3930, avg length: 982.5, speed: 1039.0 tok/s
#315
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 166
Wall time 11.0 seconds
Output:
#316
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "166",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 34.3 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 9

Warning: truncated output (original token count: 2158)
Total output lines: 71

vLLM generation done - elapsed time: 3.57s, prompts: 4, total tokens: 2547, avg length: 636.8, speed: 713.3 tok/s
vLLM generation done - elapsed time: 3.60s, prompts: 4, total tokens: 2521, avg length: 630.2, speed: 700.1 tok/s
vLLM generation done - elapsed time: 3.68s, prompts…27 tokens truncated…tal tokens: 3491, avg length: 872.8, speed: 927.7 tok/s

 20%|██        | 20/100 [02:59<11:58,  8.98s/it]
                                                
{'loss': 0.0058, 'grad_norm': 0.06886154413223267, 'learning_rate': 4.05e-06, 'on_policy_loss': 0.0058, 'rollout_accuracy': 0.375, 'epoch': 0.07}

 20%|██        | 20/100 [03:00<11:58,  8.98s/it]
#317
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 167
Wall time 11.0 seconds
Output:
#318
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "167",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 27.4 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 3065)
Total output lines: 182

vLLM generation done - elapsed time: 3.63s, prompts: 4, total tokens: 2771, avg length: 692.8, speed: 762.4 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3093, avg length: 773.2, speed: 829.4 tok/s
vLLM generation done - elapsed time: 3.82s, prompt…27 tokens truncated…8s, prompts: 4, total tokens: 3472, avg length: 868.0, speed: 942.4 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3087, avg length: 771.8, speed: 833.9 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3411, avg length: 852.8, speed: 917.5 tok/s

 27%|██▋       | 27/100 [04:02<10:53,  8.95s/it]
#319
narration

The v2 repeat is past step 25 with the same stable behavior as the first v2 run. I’ll evaluate it fully and use the v2 two-run mean versus v1’s 55.14; no further method changes will be made unless the replicated evidence clearly supports them.

#320
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 168
Wall time 11.0 seconds
Output:
#321
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "168",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 25.0 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 9

Warning: truncated output (original token count: 2051)
Total output lines: 67

vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3246, avg length: 811.5, speed: 877.6 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3745, avg length: 936.2, speed: 1001.3 tok/s
vLLM generation done - elapsed time: 3.78s, prompt…27 tokens truncated…ngth: 948.8, speed: 1009.5 tok/s

 34%|███▍      | 34/100 [05:05<09:51,  8.96s/it]
                                                
{'loss': 0.0025, 'grad_norm': 0.07858163863420486, 'learning_rate': 3.3500000000000005e-06, 'on_policy_loss': 0.0025, 'rollout_accuracy': 0.4375, 'epoch': 0.12}

 34%|███▍      | 34/100 [05:05<09:51,  8.96s/it]
#322
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 169
Wall time 11.0 seconds
Output:
#323
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "169",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 27.4 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 1859)
Total output lines: 61

vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3221, avg length: 805.2, speed: 866.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3564, avg length: 891.0, speed: 953.1 tok/s
vLLM generation done - elapsed time: 3.76s, prompts…27 tokens truncated… 776.6 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 2884, avg length: 721.0, speed: 788.0 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3730, avg length: 932.5, speed: 1002.0 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3658, avg length: 914.5, speed: 981.9 tok/s
#324
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 170
Wall time 11.0 seconds
Output:
#325
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "170",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 34.2 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 1951)
Total output lines: 68

vLLM generation done - elapsed time: 2.80s, prompts: 4, total tokens: 2473, avg length: 618.2, speed: 882.4 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.6 tok/s
vLLM generation done - elapsed time: 3.76s, promp…27 tokens truncated…mpts: 4, total tokens: 3361, avg length: 840.2, speed: 912.7 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3743, avg length: 935.8, speed: 1000.5 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1080.7 tok/s

 47%|████▋     | 47/100 [07:01<07:55,  8.97s/it]