gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 428
- duration
- 9h 21m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__eNDoWrD
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "96",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 96
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "96",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script completed
Wall time 14.9 seconds
Output:
Warning: truncated output (original token count: 184)
Total output lines: 15
| 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]
[03:14<08:31, 7.11s/it, est. speed input: 19.93 toks/s, output: 903.35 toks…104 tokens truncated…:25<00:00, 8.87s/it, est. speed input: 11.73 toks/s, output: 2117.92 toks/s]
[13:24<05:56, 9.90s/it, est. speed input: 6.13 toks/s, output: 1188.72 toks/s]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let out=\"\";for(let i=0;i<7;i++){const r=await tools.write_stdin({session_id:88389,chars:\"\",yield_time_ms:60000,max_output_tokens:60});out+=r.output.slice(-120)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}exec result
Script running with cell ID 97
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "97",
"yield_time_ms": 60000,
"max_tokens": 100
}wait result
Script running with cell ID 97
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "97",
"yield_time_ms": 60000,
"max_tokens": 100
}wait result
Script completed
Wall time 22.4 seconds
Output:
okens truncated… 75%|███████▌ | 72/96 [14:17<03:12, 8.02s/it, est. speed input: 11.28 toks/s, output: 1526.87 toks/s]
: 53.33% (192/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/dev_eval_1785705452.json
EXIT0Teacher-support KL underperforms at 53.33%, so I’m dropping it. The two safest full-run results remain tied at 54.44%. With the remaining budget I’m testing one capacity-controlled variant: keep the released OPSD objective and schedule, but give its low-rank update enough subspace to represent both algebraic and geometric corrections (rank 128 with the same LoRA scaling). This stays within the designated model and fixed optimizer budget.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n-# Teacher-support OPSD: distill the privileged teacher's likely alternatives\\n-# and preserve all remaining probability as one exact tail-mass bucket.\\n+# Higher-capacity OPSD adapter with the released update scale and schedule.\\n lmbda=0\\n-top_k_loss=64\\n+top_k_loss=0\\n+lora_r=128\\n+lora_alpha=256\\n max_completion_length=1024\\n lr_scheduler_type=linear\\n*** End Patch\";\ntext(await tools.apply_patch(patch));"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/rank128 PORT=12958 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=78856encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let out=\"\";for(let i=0;i<12;i++){const r=await tools.write_stdin({session_id:78856,chars:\"\",yield_time_ms:60000,max_output_tokens:50});out+=r.output.slice(-100)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}exec result
Script running with cell ID 100
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 60000,
"max_tokens": 80
}wait result
Script completed
Wall time 33.9 seconds
Output:
Warning: truncated output (original token count: 319)
Total output lines: 24
done - elapsed time: 3.38s, prompts: 4, total tokens: 3947, avg length: 986.8, speed: 1167.5 tok/s
d…05e-06, 'on_policy_loss': 0.0006, 'epoch': 0.01}
10%|…239 tokens truncated…██████▌ | 76/100 [10:19<03:16, 8.19s/it]
done - elapsed time: 3.40s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1206.3 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let out=\"\";for(let i=0;i<5;i++){const r=await tools.write_stdin({session_id:78856,chars:\"\",yield_time_ms:60000,max_output_tokens:70});out+=r.output.slice(-140)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}exec result
Script running with cell ID 101
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "101",
"yield_time_ms": 60000,
"max_tokens": 100
}wait result
Script running with cell ID 101
Wall time 61.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "101",
"yield_time_ms": 60000,
"max_tokens": 100
}wait result
Script completed
Wall time 51.1 seconds
Output:
Warning: truncated output (original token count: 113)
Total output lines: 9
13 tokens truncated…ate': 4.5000000000000003e-07, 'on_policy_loss': -0.0151, 'epoch': 0.1}
92%|█████████▏| 92/100 [12:30<01:05, 8.16s/it]
.5, speed: 1198.6 tok/s
vLLM generati…13 tokens truncated…okens: 4032, avg length: 1008.0, speed: 1192.3 tok/s
s://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/rank128/capped_seed42/checkpoint-100
EXIT0encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/rank128/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=82081encrypted chain-of-thought — recorded, not exposed
args
{
"input": "let out=\"\";for(let i=0;i<16;i++){const r=await tools.write_stdin({session_id:82081,chars:\"\",yield_time_ms:60000,max_output_tokens:40});out+=r.output.slice(-80)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}exec result
Script running with cell ID 103
Wall time 11.0 seconds
Output: