Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#201
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 117
Wall time 11.0 seconds
Output:
#202
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 932)
Total output lines: 32

vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5081, avg le…26 tokens truncated…59, avg length: 1439.8, speed: 1131.3 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1196.6 tok/s
#203
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 118
Wall time 11.0 seconds
Output:
#204
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "118",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.2 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1064)
Total output lines: 37

vLLM generation done - elapsed time: 4.99s, prompts: 4, total tokens: 4831, avg l…27 tokens truncated…rm': 0.04970723018050194, 'learning_rate': 1.45e-06, 'on_policy_loss': -0.0085, 'epoch': 0.63}

 72%|███████▏  | 72/100 [14:15<05:32, 11.87s/it]
#205
narration

The refined run has passed step 72 without memory pressure, NaNs, or gradient spikes. I’m letting it finish and will immediately run the exact full evaluation; no conclusions will be drawn from its training loss alone because the released clipped estimator is not a proper nonnegative KL.

#206
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 119
Wall time 11.0 seconds
Output:
#207
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "119",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 806)
Total output lines: 26

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4578, avg le…26 tokens truncated…- elapsed time: 5.13s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1198.3 tok/s

 75%|███████▌  | 75/100 [14:51<04:57, 11.90s/it]
#208
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 120
Wall time 11.0 seconds
Output:
#209
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "120",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 15.4 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1166)
Total output lines: 40

vLLM generation done - elapsed time: 5.00s, prompts: 4, total tokens: 4857, avg l…27 tokens truncated…995, avg length: 1248.8, speed: 988.0 tok/s
vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1199.6 tok/s
#210
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 121
Wall time 11.0 seconds
Output:
#211
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "121",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 955)
Total output lines: 33


 79%|███████▉  | 79/100 [15:38<04:09, 11.87s/it]vLLM generation d…26 tokens truncated…332868576, 'learning_rate': 9.500000000000001e-07, 'on_policy_loss': -0.0099, 'epoch': 0.72}

 82%|████████▏ | 82/100 [16:14<03:34, 11.90s/it]
#212
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 122
Wall time 11.0 seconds
Output:
#213
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "122",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 808)
Total output lines: 26

vLLM generation done - elapsed time: 5.10s, prompts: 4, total tokens: 5411, avg le…26 tokens truncated…elapsed time: 5.16s, prompts: 4, total tokens: 6140, avg length: 1535.0, speed: 1190.3 tok/s

 85%|████████▌ | 85/100 [16:50<02:58, 11.91s/it]
#214
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 123
Wall time 11.0 seconds
Output:
#215
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "123",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.9 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 5093)
Total output lines: 437

vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5257, avg …27 tokens truncated…m': 0.05567526817321777, 'learning_rate': 6.5e-07, 'on_policy_loss': -0.0106, 'epoch': 0.77}

 88%|████████▊ | 88/100 [17:26<02:23, 11.94s/it]
#216
narration

The run is at step 88 and remains comparable to baseline optimization. Once its full score is available, I’ll use the remaining budget for one loss-level candidate: a properly nonnegative, influence-capped token KL that preserves the paper’s robustness motivation without the current asymmetric vocabulary-term clipping.

#217
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 124
Wall time 11.0 seconds
Output:
#218
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "124",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 990)
Total output lines: 35

vLLM generation done - elapsed time: 5.03s, prompts: 4, total tokens: 4985, avg le…26 tokens truncated…44, avg length: 1536.0, speed: 1193.8 tok/s
vLLM generation done - elapsed time: 5.16s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1190.6 tok/s
#219
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 125
Wall time 11.0 seconds
Output:
#220
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "125",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 9.9 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2290)
Total output lines: 273

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5283, avg …27 tokens truncated…19, avg length: 1354.8, speed: 1064.8 tok/s
vLLM generation done - elapsed time: 5.09s, prompts: 4, total tokens: 5962, avg length: 1490.5, speed: 1171.0 tok/s
#221
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 126
Wall time 11.0 seconds
Output:
#222
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "126",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 3.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 960)
Total output lines: 33


 95%|█████████▌| 95/100 [18:49<00:59, 11.92s/it]vLLM generati…26 tokens truncated…235336, 'learning_rate': 1.5000000000000002e-07, 'on_policy_loss': -0.0105, 'epoch': 0.86}

 98%|█████████▊| 98/100 [19:25<00:23, 11.91s/it]
#223
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}
exec result
Script running with cell ID 127
Wall time 11.0 seconds
Output:
#224
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "127",
  "yield_time_ms": 30000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 6.0 seconds
Output:
vLLM generation done - elapsed time: 4.90s, prompts: 4, total tokens: 4621, avg length: 1155.2, speed: 942.6 tok/s
vLLM generation done - elapsed time: 5.04s, prompts: 4, total tokens: 5201, avg length: 1300.2, speed: 1031.9 tok/s
vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1201.0 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6091, avg length: 1522.8, speed: 1188.3 tok/s
vLLM generation done - elapsed time: 5.00s, prompts: 4, total tokens: 4575, avg length: 1143.8, speed: 915.3 tok/s
vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5083, avg length: 1270.8, speed: 1006.0 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 5615, avg length: 1403.8, speed: 1099.9 tok/s
vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1200.4 tok/s

 99%|█████████▉| 99/100 [19:37<00:11, 11.94s/it]vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5506, avg length: 1376.5, speed: 1084.5 tok/s
vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5504, avg length: 1376.0, speed: 1083.5 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1202.0 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1194.8 tok/s
vLLM generation done - elapsed time: 4.94s, prompts: 4, total tokens: 4635, avg length: 1158.8, speed: 937.4 tok/s
vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5261, avg length: 1315.2, speed: 1040.9 tok/s
vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5100, avg length: 1275.0, speed: 1005.1 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6077, avg length: 1519.2, speed: 1183.9 tok/s

100%|██████████| 100/100 [19:49<00:00, 11.96s/it]
                                                 
{'loss': -0.0097, 'grad_norm': 0.05189653858542442, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': -0.0097, 'epoch': 0.87}

100%|██████████| 100/100 [19:49<00:00, 11.96s/it]
                                                 
{'train_runtime': 1192.4141, 'train_samples_per_second': 2.684, 'train_steps_per_second': 0.084, 'train_loss': -0.005053745154291391, 'epoch': 0.87}

100%|██████████| 100/100 [19:52<00:00, 11.96s/it]
100%|██████████| 100/100 [19:52<00:00, 11.92s/it]
[rank0]:[W802 19:25:55.693550690 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W802 19:25:55.955422838 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W802 19:25:55.965909605 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W802 19:25:55.448476056 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/refined_long/capped_seed42/checkpoint-100
#225
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/refined_long/capped_seed42/checkpoint-100 /workspace/eval_results/refined_long_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":1000});\ntext(r.output); if(r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=63518