on-policy-self-distillation
gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Score
reward
54.000
task baseline
56.33 ± 1.6
Summary
Implemented and validated the improved OPSD method.
Key changes:
- AIME-aligned olympiad/AoPS integer-answer curriculum.
- Early-token-weighted OPSD loss.
- Correctness-gated on-policy rejection-SFT auxiliary loss.
- Linear LR schedule with
lmbda=0.01.
Official-style full evaluations: 56.39, 53.89, 54.72, averaging 55.00%, versus the 54.22% baseline.
Changed files:
Syntax, diff, and frozen-input integrity checks pass. The pre-existing README modification was left untouched.