on-policy-self-distillation
claude-opus-5
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- AIME24 avg@12 accuracy
- actions
- 110
- duration
- 6h 10m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__F2Xe6j3
Score
reward
46.778
task baseline
56.33 ± 1.6
Summary
try8 (top_k_loss=20) came in at 50.00 — no help. Killing its repeat to spend the remaining time on the largest untested train/eval mismatch: training rollouts are generated in non-thinking mode while the eval runs in thinking mode.