gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 596
- duration
- 7h 59m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__kuc95M7
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 299
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "299",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 44.6 seconds
Output:
eneration done - elapsed time: 2.33s, prompts: 4, total tokens: 2252, avg length: 563.0, speed: 966.1 tok/s
vLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 2987, avg length: 746.8, speed: 936.5 tok/s
vLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 2966, avg length: 741.5, speed: 907.8 tok/s
vLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3271, avg length: 817.8, speed: 993.6 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 300
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "300",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 45.9 seconds
Output:
ation done - elapsed time: 3.17s, prompts: 4, total tokens: 2710, avg length: 677.5, speed: 853.9 tok/s
vLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3380, avg length: 845.0, speed: 1022.0 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3651, avg length: 912.8, speed: 1100.1 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1231.4 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:400}); text(r.output.slice(-750));\n"
}exec result
Script running with cell ID 301
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "301",
"yield_time_ms": 60000,
"max_tokens": 400
}wait result
Script completed
Wall time 46.3 seconds
Output:
time: 3.30s, prompts: 4, total tokens: 3539, avg length: 884.8, speed: 1073.5 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3795, avg length: 948.8, speed: 1139.6 tok/s
================================================================================
Saved 40 generation outputs to:
/workspace/runs/earlycredit_lr75/capped_seed42/generations/generations_step_25.json
================================================================================
26%|██▌ | 26/100 [03:26<09:47, 7.94s/it]
{'loss': -0.0036, 'grad_norm': 0.056837618350982666, 'learning_rate': 7.5e-06, 'on_policy_loss': -0.0036, 'epoch': 0.03}
26%|██▌ | 26/100 [03:26<09:47, 7.94s/it]The 7.5e-6 run has reached step 26 with finite, unsaturated gradients and the intended constant rate. This changes optimization intensity only; the 100-step/global-batch budget and all rollout settings remain fixed.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 302
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "302",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 42.2 seconds
Output:
: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.4 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1219.1 tok/s
34%|███▍ | 34/100 [04:29<08:43, 7.94s/it]
{'loss': -0.0043, 'grad_norm': 0.057652849704027176, 'learning_rate': 7.5e-06, 'on_policy_loss': -0.0043, 'epoch': 0.04}
34%|███▍ | 34/100 [04:29<08:43, 7.94s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 303
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "303",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 45.6 seconds
Output:
ion done - elapsed time: 3.28s, prompts: 4, total tokens: 3169, avg length: 792.2, speed: 964.8 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.6 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1222.1 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1221.3 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 304
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "304",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 43.4 seconds
Output:
tion done - elapsed time: 3.28s, prompts: 4, total tokens: 3090, avg length: 772.5, speed: 941.6 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3781, avg length: 945.2, speed: 1136.6 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.5 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4074, avg length: 1018.5, speed: 1218.5 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 305
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "305",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 42.1 seconds
Output:
tion done - elapsed time: 3.30s, prompts: 4, total tokens: 3299, avg length: 824.8, speed: 998.3 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1234.0 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3897, avg length: 974.2, speed: 1171.9 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.8 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 306
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "306",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 46.2 seconds
Output:
tion done - elapsed time: 3.27s, prompts: 4, total tokens: 3092, avg length: 773.0, speed: 945.7 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.9 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1225.4 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3790, avg length: 947.5, speed: 1136.3 tok/s
Training remains stable past the midpoint; the larger rate has not triggered clipping saturation or runaway gradients. I’ll still select it only by full AIME accuracy, not training-loss appearance.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:300}); text(r.output.slice(-600));\n"
}exec result
Script running with cell ID 307
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "307",
"yield_time_ms": 60000,
"max_tokens": 300
}wait result
Script completed
Wall time 45.6 seconds
Output:
okens truncated…906, 'learning_rate': 7.5e-06, 'on_policy_loss': -0.0128, 'epoch': 0.08}
74%|███████▍ | 74/100 [09:49<03:27, 7.97s/it]vLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3384, avg length: 846.0, speed: 1023.9 tok/s
vLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3766, avg length: 941.5, speed: 1136.9 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1227.8 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1216.1 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 308
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "308",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 35.9 seconds
Output:
tion done - elapsed time: 3.29s, prompts: 4, total tokens: 3080, avg length: 770.0, speed: 936.7 tok/s
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3485, avg length: 871.2, speed: 1049.6 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1230.6 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1215.9 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 309
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "309",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 46.3 seconds
Output:
ion done - elapsed time: 3.30s, prompts: 4, total tokens: 3323, avg length: 830.8, speed: 1007.9 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3808, avg length: 952.0, speed: 1141.1 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1224.9 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1222.7 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:28348,chars:\"\",yield_time_ms:60000,max_output_tokens:250}); text(r.output.slice(-450));\n"
}exec result
Script running with cell ID 310
Wall time 11.0 seconds
Output: