Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#576
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 297
Wall time 11.0 seconds
Output:
#577
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "297",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 976)
Total output lines: 31

vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 5224, avg le…26 tokens truncated…tion done - elapsed time: 5.06s, prompts: 4, total tokens: 5260, avg length: 1315.0, speed: 1040.4 tok/s

 25%|██▌       | 25/100 [04:54<14:45, 11.81s/it]
#578
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 298
Wall time 11.0 seconds
Output:
#579
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "298",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1165)
Total output lines: 41

vLLM generation done - elapsed time: 3.44s, prompts: 4, total tokens: 3148, avg l…27 tokens truncated…tion done - elapsed time: 5.14s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1194.5 tok/s

 29%|██▉       | 29/100 [05:42<13:58, 11.81s/it]
#580
narration

The third run is near step 30 and continues to mirror the first two. Once it finishes, one final full evaluation will give a three-run local mean to compare with the official five-run baseline.

#581
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 299
Wall time 11.0 seconds
Output:
#582
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "299",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 3.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1169)
Total output lines: 41

vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4360, avg l…27 tokens truncated…on done - elapsed time: 5.11s, prompts: 4, total tokens: 5754, avg length: 1438.5, speed: 1126.1 tok/s

 33%|███▎      | 33/100 [06:29<13:11, 11.81s/it]
#583
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 300
Wall time 11.0 seconds
Output:
#584
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "300",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 956)
Total output lines: 33

vLLM generation done - elapsed time: 4.70s, prompts: 4, total tokens: 4093, avg le…26 tokens truncated…█▌      | 36/100 [07:05<12:36, 11.82s/it]vLLM generation done - elapsed time: 4.82s, prompts: 4, total tokens: 3796, avg length: 949.0, speed: 788.1 tok/s
#585
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 301
Wall time 11.0 seconds
Output:
#586
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "301",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 884)
Total output lines: 28

vLLM generation done - elapsed time: 4.98s, prompts: 4, total tokens: 4688, avg le…26 tokens truncated…63, avg length: 1265.8, speed: 1006.6 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1193.5 tok/s
#587
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 302
Wall time 11.0 seconds
Output:
#588
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "302",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 9.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1036)
Total output lines: 36

vLLM generation done - elapsed time: 4.98s, prompts: 4, total tokens: 4609, avg l…27 tokens truncated…56, avg length: 1289.0, speed: 1015.6 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1191.9 tok/s
#589
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 303
Wall time 11.0 seconds
Output:
#590
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "303",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.7 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 942)
Total output lines: 33


 43%|████▎     | 43/100 [08:28<11:17, 11.88s/it]vLLM generation done - …26 tokens truncated…05659085512161255, 'learning_rate': 2.7500000000000004e-06, 'on_policy_loss': -0.0054, 'epoch': 0.4}

 46%|████▌     | 46/100 [09:03<10:41, 11.88s/it]
#591
narration

At step 46, the third run’s clipped objective (−0.0054) and gradient norm (0.057) remain in the same range as the validated runs. Training is on schedule.

#592
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 304
Wall time 11.0 seconds
Output:
#593
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "304",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 799)
Total output lines: 26

vLLM generation done - elapsed time: 5.00s, prompts: 4, total tokens: 4367, avg le…26 tokens truncated… done - elapsed time: 5.10s, prompts: 4, total tokens: 6031, avg length: 1507.8, speed: 1182.6 tok/s

 49%|████▉     | 49/100 [09:39<10:06, 11.89s/it]
#594
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 305
Wall time 11.0 seconds
Output:
#595
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "305",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.5 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1074)
Total output lines: 37

vLLM generation done - elapsed time: 5.02s, prompts: 4, total tokens: 4712, avg l…27 tokens truncated…144, avg length: 1536.0, speed: 1197.2 tok/s
vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4662, avg length: 1165.5, speed: 938.9 tok/s
#596
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 306
Wall time 11.0 seconds
Output:
#597
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "306",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.2 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2516)
Total output lines: 170

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5572, avg …27 tokens truncated…d_norm': 0.056033145636320114, 'learning_rate': 2.25e-06, 'on_policy_loss': -0.008, 'epoch': 0.49}

 56%|█████▌    | 56/100 [11:03<08:46, 11.97s/it]
#598
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 307
Wall time 11.0 seconds
Output:
#599
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "307",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 857)
Total output lines: 27

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 5143, avg le…26 tokens truncated…4169, avg length: 1042.2, speed: 843.7 tok/s
vLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4395, avg length: 1098.8, speed: 884.2 tok/s
#600
narration

The third replicate is past halfway and remains stable. No further method changes are planned; the remaining work is replication, final integrity checks, and concise handoff.