Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
428
duration
9h 21m
recorded
2026-08-02

run idopsd-aime24-4gpus__eNDoWrD

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol444 turns16 narration428 reasoning428 action428 observation1 final
#276
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "96",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 96
Wall time 61.0 seconds
Output:
#277
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "96",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.9 seconds
Output:
Warning: truncated output (original token count: 184)
Total output lines: 15

     | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]


 [03:14<08:31,  7.11s/it, est. speed input: 19.93 toks/s, output: 903.35 toks…104 tokens truncated…:25<00:00,  8.87s/it, est. speed input: 11.73 toks/s, output: 2117.92 toks/s]

 [13:24<05:56,  9.90s/it, est. speed input: 6.13 toks/s, output: 1188.72 toks/s]
#278
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<7;i++){const r=await tools.write_stdin({session_id:88389,chars:\"\",yield_time_ms:60000,max_output_tokens:60});out+=r.output.slice(-120)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 97
Wall time 11.0 seconds
Output:
#279
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "97",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 97
Wall time 61.0 seconds
Output:
#280
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "97",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 22.4 seconds
Output:
okens truncated…  75%|███████▌  | 72/96 [14:17<03:12,  8.02s/it, est. speed input: 11.28 toks/s, output: 1526.87 toks/s]
: 53.33% (192/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/dev_eval_1785705452.json

EXIT0
#281
narration

Teacher-support KL underperforms at 53.33%, so I’m dropping it. The two safest full-run results remain tied at 54.44%. With the remaining budget I’m testing one capacity-controlled variant: keep the released OPSD objective and schedule, but give its low-rank update enough subspace to represent both algebraic and geometric corrections (rank 128 with the same LoRA scaling). This stays within the designated model and fixed optimizer budget.

#282
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n-# Teacher-support OPSD: distill the privileged teacher's likely alternatives\\n-# and preserve all remaining probability as one exact tail-mass bucket.\\n+# Higher-capacity OPSD adapter with the released update scale and schedule.\\n lmbda=0\\n-top_k_loss=64\\n+top_k_loss=0\\n+lora_r=128\\n+lora_alpha=256\\n max_completion_length=1024\\n lr_scheduler_type=linear\\n*** End Patch\";\ntext(await tools.apply_patch(patch));"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{}
#283
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/rank128 PORT=12958 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=78856
#284
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<12;i++){const r=await tools.write_stdin({session_id:78856,chars:\"\",yield_time_ms:60000,max_output_tokens:50});out+=r.output.slice(-100)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 100
Wall time 11.0 seconds
Output:
#285
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#286
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#287
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#288
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#289
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#290
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#291
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#292
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#293
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#294
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 100
Wall time 61.0 seconds
Output:
#295
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "100",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 33.9 seconds
Output:
Warning: truncated output (original token count: 319)
Total output lines: 24

 done - elapsed time: 3.38s, prompts: 4, total tokens: 3947, avg length: 986.8, speed: 1167.5 tok/s

d…05e-06, 'on_policy_loss': 0.0006, 'epoch': 0.01}

 10%|…239 tokens truncated…██████▌  | 76/100 [10:19<03:16,  8.19s/it]
done - elapsed time: 3.40s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1206.3 tok/s

#296
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<5;i++){const r=await tools.write_stdin({session_id:78856,chars:\"\",yield_time_ms:60000,max_output_tokens:70});out+=r.output.slice(-140)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 101
Wall time 11.0 seconds
Output:
#297
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "101",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 101
Wall time 61.0 seconds
Output:
#298
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "101",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 51.1 seconds
Output:
Warning: truncated output (original token count: 113)
Total output lines: 9

13 tokens truncated…ate': 4.5000000000000003e-07, 'on_policy_loss': -0.0151, 'epoch': 0.1}

 92%|█████████▏| 92/100 [12:30<01:05,  8.16s/it]
.5, speed: 1198.6 tok/s
vLLM generati…13 tokens truncated…okens: 4032, avg length: 1008.0, speed: 1192.3 tok/s

s://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/rank128/capped_seed42/checkpoint-100

EXIT0
#299
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/rank128/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=82081
#300
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<16;i++){const r=await tools.write_stdin({session_id:82081,chars:\"\",yield_time_ms:60000,max_output_tokens:40});out+=r.output.slice(-80)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 103
Wall time 11.0 seconds
Output: