on-policy-self-distillation
claude-opus-5
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- AIME24 avg@12 accuracy
- actions
- 148
- duration
- 9h 4m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__LSVoBuR
What the verifier emitted for this run, as recorded. The fields differ by task family — the task’s verification method explains what they mean.
- score
- 54.3333
- metric
- AIME24 avg@12 accuracy
- direction
- higher_better
- baseline
- 54.22
- theoretical best
- 100
- practical best
- 100
- status
- ok
- flags
- none
- reward
- 54.3333
- passed
- pass
- metrics
- avg at 12 pct
- 54.333
- per run scores
- 53.3333, 54.1667, 56.1111, 53.6111, 54.4444
- baseline mean
- 54.22
- num runs
- 5
- train seed
- 42
- noise 1sigma fixed seed
- 1.48
- measurement
- avg@12 accuracy on AIME24 (30 problems), mean over 5 same-seed(42) re-trains
- training
- verifier-owned; 100 optimizer steps, global batch 32, frozen base+data; 5x same-seed(42)
- note
- absolute measured metric; baseline/theoretical_best are anchors for a separate downstream normalization layer (METHODOLOGY.md), not applied here