Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Score

reward
54.000
task baseline
56.33 ± 1.6
compare against every run on this task →

Summary

Implemented and validated the improved OPSD method.

Key changes:

  • AIME-aligned olympiad/AoPS integer-answer curriculum.
  • Early-token-weighted OPSD loss.
  • Correctness-gated on-policy rejection-SFT auxiliary loss.
  • Linear LR schedule with lmbda=0.01.

Official-style full evaluations: 56.39, 53.89, 54.72, averaging 55.00%, versus the 54.22% baseline.

Changed files:

Syntax, diff, and frozen-input integrity checks pass. The pre-existing README modification was left untouched.

Transcript