Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#601
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 308
Wall time 11.0 seconds
Output:
#602
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "308",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.3 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1121)
Total output lines: 39

vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 6144, avg l…27 tokens truncated…e - elapsed time: 5.13s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1197.5 tok/s

 63%|██████▎   | 63/100 [12:26<07:20, 11.91s/it]
#603
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 309
Wall time 11.0 seconds
Output:
#604
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "309",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 9.5 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1044)
Total output lines: 36

vLLM generation done - elapsed time: 4.92s, prompts: 4, total tokens: 4547, avg l…27 tokens truncated…37, avg length: 1434.2, speed: 1118.1 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 5786, avg length: 1446.5, speed: 1126.6 tok/s
#605
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 310
Wall time 11.0 seconds
Output:
#606
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "310",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 4.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1114)
Total output lines: 34

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 5121, avg l…27 tokens truncated…77, avg length: 1319.2, speed: 1038.9 tok/s
vLLM generation done - elapsed time: 5.09s, prompts: 4, total tokens: 5866, avg length: 1466.5, speed: 1152.6 tok/s
#607
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 311
Wall time 11.0 seconds
Output:
#608
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "311",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.3 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1066)
Total output lines: 37

vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5612, avg l…27 tokens truncated…3542637825, 'learning_rate': 1.3500000000000002e-06, 'on_policy_loss': -0.0091, 'epoch': 0.65}

 74%|███████▍  | 74/100 [14:37<05:09, 11.91s/it]
#609
narration

The final replicate is at step 74 with expected loss and gradient statistics. It should finish shortly, after which I’ll run the last full evaluation and compute the three-run mean.

#610
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 312
Wall time 11.0 seconds
Output:
#611
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "312",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.2 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 989)
Total output lines: 35

vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5821, avg le…26 tokens truncated…61, avg length: 1490.2, speed: 1164.9 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1198.4 tok/s
#612
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 313
Wall time 11.0 seconds
Output:
#613
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "313",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 1.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 869)
Total output lines: 26

vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 5084, avg le…26 tokens truncated…04, avg length: 1501.0, speed: 1175.1 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 6096, avg length: 1524.0, speed: 1193.6 tok/s
#614
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 314
Wall time 11.0 seconds
Output:
#615
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "314",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 993)
Total output lines: 35

vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 4879, avg le…26 tokens truncated…44, avg length: 1536.0, speed: 1199.0 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 5948, avg length: 1487.0, speed: 1156.9 tok/s
#616
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 315
Wall time 11.0 seconds
Output:
#617
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "315",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.8 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 5

Warning: truncated output (original token count: 5339)
Total output lines: 493


 84%|████████▍ | 84/100 [16:37<03:11, 11.99s/it]
            …27 tokens truncated…44, avg length: 1536.0, speed: 1205.2 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1203.4 tok/s
#618
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 316
Wall time 11.0 seconds
Output:
#619
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "316",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.6 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1072)
Total output lines: 37

vLLM generation done - elapsed time: 5.04s, prompts: 4, total tokens: 5198, avg l…27 tokens truncated…6519699, 'learning_rate': 4.5000000000000003e-07, 'on_policy_loss': -0.0099, 'epoch': 0.8}

 92%|█████████▏| 92/100 [18:13<01:36, 12.01s/it]
#620
narration

The last training run is at step 92 and remains clean. I’ll proceed directly to its full evaluation, then verify the final diff and recipe against the hard constraints.

#621
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 317
Wall time 11.0 seconds
Output:
#622
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "317",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 9.0 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2374)
Total output lines: 238

vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 4496, avg …27 tokens truncated…98, avg length: 1299.5, speed: 1023.9 tok/s
vLLM generation done - elapsed time: 5.10s, prompts: 4, total tokens: 5554, avg length: 1388.5, speed: 1089.3 tok/s
#623
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 318
Wall time 11.0 seconds
Output:
#624
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "318",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 8

Warning: truncated output (original token count: 962)
Total output lines: 33


================================================================================
…26 tokens truncated…apsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1194.0 tok/s

 99%|█████████▉| 99/100 [19:36<00:11, 11.92s/it]
#625
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:800});\ntext(r.output);\n"
}
exec result
Script completed
Wall time 2.0 seconds
Output:
vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5339, avg length: 1334.8, speed: 1057.5 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 5824, avg length: 1456.0, speed: 1139.8 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 5943, avg length: 1485.8, speed: 1162.6 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1201.4 tok/s
vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4557, avg length: 1139.2, speed: 909.6 tok/s
vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4674, avg length: 1168.5, speed: 932.7 tok/s
vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 5388, avg length: 1347.0, speed: 1067.4 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 5953, avg length: 1488.2, speed: 1161.2 tok/s

100%|██████████| 100/100 [19:48<00:00, 11.92s/it]
                                                 
{'loss': -0.0095, 'grad_norm': 0.050380807369947433, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': -0.0095, 'epoch': 0.87}

100%|██████████| 100/100 [19:48<00:00, 11.92s/it]
                                                 
{'train_runtime': 1191.5662, 'train_samples_per_second': 2.686, 'train_steps_per_second': 0.084, 'train_loss': -0.004995019435882569, 'epoch': 0.87}

100%|██████████| 100/100 [19:51<00:00, 11.92s/it]
100%|██████████| 100/100 [19:51<00:00, 11.92s/it]
[rank0]:[W802 21:25:43.722890895 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W802 21:25:44.033183994 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W802 21:25:44.033680501 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W802 21:25:44.074828517 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/refined_repeat3/capped_seed42/checkpoint-100