Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#526
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 257
Wall time 11.0 seconds
Output:
#527
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "257",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 47.2 seconds
Output:
29s, prompts: 4, total tokens: 3361, avg length: 840.2, speed: 1021.8 tok/s
vLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3735, avg length: 933.8, speed: 1133.8 tok/s
vLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3289, avg length: 822.2, speed: 996.0 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3947, avg length: 986.8, speed: 1182.2 tok/s
#528
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 258
Wall time 11.0 seconds
Output:
#529
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "258",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 47.0 seconds
Output:
21s, prompts: 4, total tokens: 2601, avg length: 650.2, speed: 810.0 tok/s
vLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3201, avg length: 800.2, speed: 980.2 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3401, avg length: 850.2, speed: 1023.5 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1222.0 tok/s
#530
narration

The candidate has passed the midpoint with stable loss and gradients. Equalizing trajectory contributions changes only the reduction—not sampling, batch size, completion budget, or optimizer-step budget.

#531
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 259
Wall time 11.0 seconds
Output:
#532
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "259",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 47.2 seconds
Output:
vg length: 988.8, speed: 1181.6 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1217.3 tok/s

 68%|██████▊   | 68/100 [09:02<04:16,  8.03s/it]
                                                
{'loss': -0.0102, 'grad_norm': 0.05043625459074974, 'learning_rate': 5e-06, 'on_policy_loss': -0.0102, 'epoch': 0.07}

 68%|██████▊   | 68/100 [09:02<04:16,  8.03s/it]
#533
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 260
Wall time 11.0 seconds
Output:
#534
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "260",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 46.7 seconds
Output:
rlycredit_trajbalanced/capped_seed42/generations/generations_step_75.json
================================================================================


 76%|███████▌  | 76/100 [10:06<03:11,  7.98s/it]
                                                
{'loss': 0.0026, 'grad_norm': 0.20465494692325592, 'learning_rate': 5e-06, 'on_policy_loss': 0.0026, 'epoch': 0.08}

 76%|███████▌  | 76/100 [10:06<03:11,  7.98s/it]
#535
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 261
Wall time 11.0 seconds
Output:
#536
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "261",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 41.7 seconds
Output:
732.2, speed: 896.8 tok/s
vLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3560, avg length: 890.0, speed: 1086.3 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3694, avg length: 923.5, speed: 1111.6 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3597, avg length: 899.2, speed: 1081.5 tok/s

 84%|████████▍ | 84/100 [11:10<02:07,  7.98s/it]
#537
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 250}); text(r.output.slice(-420));\n"
}
exec result
Script running with cell ID 262
Wall time 11.0 seconds
Output:
#538
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "262",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 46.8 seconds
Output:
3s, prompts: 4, total tokens: 2833, avg length: 708.2, speed: 877.1 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3732, avg length: 933.0, speed: 1123.4 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3348, avg length: 837.0, speed: 1007.4 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1229.8 tok/s
#539
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 60000, max_output_tokens: 350}); text(r.output.slice(-700));\n"
}
exec result
Script running with cell ID 263
Wall time 11.0 seconds
Output:
#540
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "263",
  "yield_time_ms": 60000,
  "max_tokens": 350
}
wait result
Script completed
Wall time 46.7 seconds
Output:
03 tokens truncated…tokens: 3870, avg length: 967.5, speed: 1165.2 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3774, avg length: 943.5, speed: 1133.8 tok/s

 99%|█████████▉| 99/100 [13:10<00:08,  8.01s/it]vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1220.9 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1228.8 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3855, avg length: 963.8, speed: 1148.7 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1216.3 tok/s
#541
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 69361, chars: \"\", yield_time_ms: 30000, max_output_tokens: 1000}); text(r.output.slice(-1800));\n"
}
exec result
Script completed
Wall time 7.6 seconds
Output:
rompts: 4, total tokens: 3757, avg length: 939.2, speed: 1123.2 tok/s

100%|██████████| 100/100 [13:18<00:00,  8.01s/it]
                                                 
{'loss': -0.0135, 'grad_norm': 0.04120411351323128, 'learning_rate': 5e-06, 'on_policy_loss': -0.0135, 'epoch': 0.11}

100%|██████████| 100/100 [13:18<00:00,  8.01s/it]
                                                 
{'train_runtime': 802.1186, 'train_samples_per_second': 3.989, 'train_steps_per_second': 0.125, 'train_loss': -0.003942320942878723, 'epoch': 0.11}

100%|██████████| 100/100 [13:22<00:00,  8.01s/it]
100%|██████████| 100/100 [13:22<00:00,  8.02s/it]
[rank0]:[W803 08:59:55.428232128 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W803 08:59:56.561157708 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W803 08:59:56.571407167 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W803 08:59:56.063813142 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/earlycredit_trajbalanced/capped_seed42/checkpoint-100
#542
narration

Training completed at exactly 100 optimizer steps. I’m starting the full 30-problem, 12-sample evaluation now; this is the same meaningful comparison used for all retained/rejected variants.

#543
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const t = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/earlycredit_trajbalanced/capped_seed42/checkpoint-100\",workdir:\"/workspace\",yield_time_ms:1000,max_output_tokens:500}); text(JSON.stringify(t));\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"4fcd28","wall_time_seconds":1.001368478,"session_id":58459,"original_token_count":0,"output":""}
#544
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 266
Wall time 11.0 seconds
Output:
#545
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "266",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 44.6 seconds
Output:
da.
INFO 08-03 09…5633 tokens truncated…y:306] Supported_tasks: ['generate']

Adding requests:   0%|          | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████| 8/8 [00:00<00:00, 272.15it/s]

Processed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]INFO 08-03 09:00:50 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|          | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████| 8/8 [00:00<00:00, 253.85it/s]

Processed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]
#546
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 267
Wall time 11.0 seconds
Output:
#547
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "267",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 38.8 seconds
Output:
#548
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 268
Wall time 11.0 seconds
Output:
#549
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "268",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.7 seconds
Output:

Processed prompts:  12%|█▎        | 12/96 [02:21<16:32, 11.82s/it, est. speed input: 13.20 toks/s, output: 501.12 toks/s]
#550
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 269
Wall time 11.0 seconds
Output: