Skip to content
on-policy-self-distillation

claude-opus-5

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
claude claude-opus-5
metric
AIME24 avg@12 accuracy
actions
110
duration
6h 10m
recorded
2026-08-02

run idopsd-aime24-4gpus__F2Xe6j3

Score

reward
46.778
task baseline
56.33 ± 1.6
compare against every run on this task →

Summary

try8 (top_k_loss=20) came in at 50.00 — no help. Killing its repeat to spend the remaining time on the largest untested train/eval mismatch: training rollouts are generated in non-thinking mode while the eval runs in thinking mode.

Transcript