Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#326
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 171
Wall time 11.0 seconds
Output:
#327
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "171",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 32.2 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 3516)
Total output lines: 139

vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 2943, avg length: 735.8, speed: 802.3 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3681, avg length: 920.2, speed: 987.7 tok/s
vLLM generation done - elapsed time: 3.74s, prompt…27 tokens truncated…7 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.2 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1093.3 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1090.1 tok/s
#328
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 172
Wall time 11.0 seconds
Output:
#329
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "172",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 28.6 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 2066)
Total output lines: 71

vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3364, avg length: 841.0, speed: 907.9 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3346, avg length: 836.5, speed: 902.8 tok/s
vLLM generation done - elapsed time: 3.73s, prompts…27 tokens truncated…819.5 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3348, avg length: 837.0, speed: 902.7 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3916, avg length: 979.0, speed: 1045.9 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3944, avg length: 986.0, speed: 1044.8 tok/s
#330
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 173
Wall time 11.0 seconds
Output:
#331
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "173",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 27.4 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 1756)
Total output lines: 57

vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3730, avg length: 932.5, speed: 1002.6 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3549, avg length: 887.2, speed: 950.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompt…27 tokens truncated…9.1 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3335, avg length: 833.8, speed: 905.4 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1088.2 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1080.0 tok/s
#332
narration

The v2 repeat is beyond the midpoint with finite, modest gradients and no failures. At this point the only remaining uncertainty is benchmark variance; the implementation itself has now completed hundreds of optimizer steps across repeated capped runs without instability.

#333
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 174
Wall time 11.0 seconds
Output:
#334
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "174",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 28.3 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 2079)
Total output lines: 66

vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1087.9 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1086.7 tok/s
vLLM generation done - elapsed time: 3.78s, pro…27 tokens truncated…944.0 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3490, avg length: 872.5, speed: 938.1 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 3915, avg length: 978.8, speed: 1037.7 tok/s
vLLM generation done - elapsed time: 3.78s, prompts: 4, total tokens: 3955, avg length: 988.8, speed: 1046.6 tok/s
#335
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 175
Wall time 11.0 seconds
Output:
#336
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "175",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 32.6 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 2084)
Total output lines: 71

vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3550, avg length: 887.5, speed: 952.0 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1091.3 tok/s
vLLM generation done - elapsed time: 3.77s, promp…27 tokens truncated…17.3 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 3801, avg length: 950.2, speed: 1026.6 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3440, avg length: 860.0, speed: 927.0 tok/s
vLLM generation done - elapsed time: 3.77s, prompts: 4, total tokens: 4020, avg length: 1005.0, speed: 1067.3 tok/s
#337
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}
exec result
Script running with cell ID 176
Wall time 11.0 seconds
Output:
#338
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "176",
  "yield_time_ms": 50000,
  "max_tokens": 180
}
wait result
Script completed
Wall time 33.7 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10

Warning: truncated output (original token count: 4442)
Total output lines: 327

vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3479, avg length: 869.8, speed: 936.8 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 3776, avg length: 944.0, speed: 1014.4 tok/s
vLLM generation done - elapsed time: 3.74s, promp…27 tokens truncated…=========================


 88%|████████▊ | 88/100 [13:08<01:47,  8.95s/it]
                                                
{'loss': 0.0006, 'grad_norm': 0.05855390429496765, 'learning_rate': 6.5e-07, 'on_policy_loss': 0.0006, 'rollout_accuracy': 0.3125, 'epoch': 0.32}

 88%|████████▊ | 88/100 [13:08<01:47,  8.95s/it]
#339
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 177
Wall time 11.0 seconds
Output:
#340
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "177",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 27.0 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 12

Warning: truncated output (original token count: 3034)
Total output lines: 237

vLLM generation done - elapsed time: 3.60s, prompts: 4, total tokens: 2931, avg length: 732.8, speed: 813.1 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3693, avg length: 923.2, speed: 989.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3440, avg length: 860.0, speed: 919.5 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096,…27 tokens truncated…: 3.74s, prompts: 4, total tokens: 3490, avg length: 872.5, speed: 933.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1093.8 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.1 tok/s
vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1088.3 tok/s

 95%|█████████▌| 95/100 [14:10<00:44,  8.96s/it]
#341
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:92058,chars:\"\",yield_time_ms:50000,max_output_tokens:700});\ntext(r.output);"
}
exec result
Script running with cell ID 178
Wall time 11.0 seconds
Output:
#342
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "178",
  "yield_time_ms": 50000,
  "max_tokens": 700
}
wait result
Script completed
Wall time 15.4 seconds
Output:
Warning: truncated output (original token count: 727)
Total output lines: 24

Warning: truncated output (original token count: 1907)
Total output lines: 56

vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3345, avg length: 836.2, speed: 907.6 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 3890, avg length: 972.5, speed: 1042.5 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3769, avg length: 942.2, speed: 1008.3 tok/s
vLLM generation done - elapsed time: 3.75s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1092.3 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1100.0 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3810, avg length: 952.5, speed: 1017.8 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3828, avg length: 957.0, speed: 1022.6 tok/s
vLLM generation done - elapsed time: 3.80s, prompts: 4, total tokens: 4089, avg length: 1022.2, speed: 1075.3 tok/s

================================================================================
Saved 40 generation outputs to:
  /workspace/runs/frontier_v2_rep2/capped_seed42/generations/generations_step_95.json
================================================================================


 96%|█████████▌| 96/100 [14:19<00:35,  8.96s/it]
                                 …27 tokens truncated…cond': 3.561, 'train_steps_per_second': 0.111, 'train_loss': 0.0012850296497344971, 'epoch': 0.36}

100%|██████████| 100/100 [14:58<00:00,  8.89s/it]
100%|██████████| 100/100 [14:58<00:00,  8.99s/it]
[rank0]:[W803 06:57:38.530152832 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 06:57:39.398656085 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 06:57:39.459447751 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 06:57:39.459596798 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/frontier_v2_rep2/capped_seed42/checkpoint-100
#343
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/frontier_v2_rep2/capped_seed42/checkpoint-100 /workspace/eval_results/frontier_v2_rep2_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300,\"tty\":true});\ntext(JSON.stringify(r));"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"19c11d","wall_time_seconds":1.001775453,"session_id":51685,"original_token_count":0,"output":""}
#344
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}
exec result
Script running with cell ID 180
Wall time 11.0 seconds
Output:
#345
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "180",
  "yield_time_ms": 50000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 34.6 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 13

Warning: truncated output (original token count: 5926)
Total output lines: 146

Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...
`torch_dtype` is deprecated! Use `dtype` instead!

Loading checkpoint shards:   0%|                         | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards: 100%|█████████████████| 2/2 [00:00<00:00, 45.99it/s]
Generating with data_parallel_size=4 (TP=1 per engine) ...
INFO 08-03 06:58:28 …27 tokens truncated…█████████| 7/7 [00:00<00:00, 193.44it/s]

Processed prompts:   0%| | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,INFO 08-03 06:58:48 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|                                   | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 8/8 [00:00<00:00, 218.89it/s]

Processed prompts:   0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,
#346
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 181
Wall time 11.0 seconds
Output:
#347
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "181",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 31.8 seconds
Output:
#348
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 182
Wall time 11.0 seconds
Output:
#349
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "182",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 25.8 seconds
Output:
#350
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 183
Wall time 11.0 seconds
Output: