Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#626
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "310",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.4 seconds
Output:
 tokens: 3659, avg length: 914.8, speed: 1105.6 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1235.5 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3798, avg length: 949.5, speed: 1132.2 tok/s
vLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1212.1 tok/s

 99%|█████████▉| 99/100 [13:09<00:08,  8.00s/it]
#627
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:30000,max_output_tokens:800}); text(r.output.slice(-1500));\n"
}
exec result
Script running with cell ID 311
Wall time 11.0 seconds
Output:
#628
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "311",
  "yield_time_ms": 30000,
  "max_tokens": 800
}
wait result
Script completed
Wall time 3.1 seconds
Output:
0%|██████████| 100/100 [13:17<00:00,  8.00s/it]
                                                 
{'train_runtime': 801.2703, 'train_samples_per_second': 3.994, 'train_steps_per_second': 0.125, 'train_loss': -0.008110508918762208, 'epoch': 0.11}

100%|██████████| 100/100 [13:21<00:00,  8.00s/it]
100%|██████████| 100/100 [13:21<00:00,  8.01s/it]
[rank0]:[W803 09:34:17.386161596 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 09:34:18.546885125 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 09:34:18.557258449 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 09:34:18.557411248 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/earlycredit_lr75/capped_seed42/checkpoint-100
#629
narration

The 7.5e-6 candidate finished all 100 steps cleanly. I’m running its full evaluation now; a weak or marginal result will cause the recipe override to be removed.

#630
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/earlycredit_lr75/capped_seed42/checkpoint-100\",workdir:\"/workspace\",yield_time_ms:1000,max_output_tokens:300}); text(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"f8eba7","wall_time_seconds":1.000999851,"session_id":67176,"original_token_count":0,"output":""}
#631
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 313
Wall time 11.0 seconds
Output:
#632
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "313",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.8 seconds
Output:
100%|██████████| 7/7 [00:00<00:00, 231.68it/s]

Processed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]INFO 08-03 09:35:10 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|          | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████| 8/8 [00:00<00:00, 265.24it/s]

Processed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]
#633
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 314
Wall time 11.0 seconds
Output:
#634
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "314",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.5 seconds
Output:
#635
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 315
Wall time 11.0 seconds
Output:
#636
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "315",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.0 seconds
Output:
#637
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 316
Wall time 11.0 seconds
Output:
#638
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "316",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.3 seconds
Output:

Processed prompts:  12%|█▎        | 12/96 [03:28<24:19, 17.37s/it, est. speed input: 9.61 toks/s, output: 479.46 toks/s]
#639
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 317
Wall time 11.0 seconds
Output:
#640
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "317",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 46.3 seconds
Output:

Processed prompts:  14%|█▍        | 12/84 [03:51<23:08, 19.29s/it, est. speed input: 3.99 toks/s, output: 469.94 toks/s]
Processed prompts:  25%|██▌       | 24/96 [04:19<11:35,  9.66s/it, est. speed input: 14.93 toks/s, output: 688.96 toks/s]
#641
narration

Full generation is progressing normally. Because this comparison differs only in learning rate, it gives a clean test of whether the front-loaded objective was under-updated at the baseline rate.

#642
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 318
Wall time 11.0 seconds
Output:
#643
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "318",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 46.1 seconds
Output:
#644
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 319
Wall time 11.0 seconds
Output:
#645
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "319",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.8 seconds
Output:
#646
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}
exec result
Script running with cell ID 320
Wall time 11.0 seconds
Output:
#647
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "320",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 46.4 seconds
Output:
#648
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:67176,chars:\"\",yield_time_ms:60000,max_output_tokens:300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 321
Wall time 11.0 seconds
Output:
#649
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "321",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.1 seconds
Output:

Processed prompts:  12%|█▎        | 12/96 [08:21<58:31, 41.80s/it, est. speed input: 3.54 toks/s, output: 231.72 toks/s]
Processed prompts:  38%|███▊      | 36/96 [08:32<14:51, 14.86s/it, est. speed input: 9.23 toks/s, output: 602.71 toks/s] 
Processed prompts:  29%|██▊       | 24/84 [08:37<21:58, 21.97s/it, est. speed input: 3.48 toks/s, output: 470.89 toks/s]
Processed prompts:  43%|████▎     | 36/84 [08:56<10:08, 12.67s/it, est. speed input: 5.68 toks/s, output: 932.69 toks/s]
#650
narration

The slowest shard is dominated by unusually long reasoning traces, but there are no engine errors. If 7.5e-6 wins only narrowly, I will still prefer the replicated 5e-6 recipe because single-run training noise is about 1.5 points.