on-policy-self-distillation
gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 428
- duration
- 9h 21m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__eNDoWrD
What the verifier emitted for this run, as recorded. The fields differ by task family — the task’s verification method explains what they mean.
- score
- 56.5556
- metric
- AIME24 avg@12 accuracy
- direction
- higher_better
- baseline
- 54.22
- theoretical best
- 100
- practical best
- 100
- status
- ok
- flags
- none
- reward
- 56.5556
- passed
- pass
- metrics
- avg at 12 pct
- 56.556
- per run scores
- 56.6667, 56.9444, 55.8333, 56.3889, 56.9444
- baseline mean
- 54.22
- num runs
- 5
- train seed
- 42
- noise 1sigma fixed seed
- 1.48
- measurement
- avg@12 accuracy on AIME24 (30 problems), mean over 5 same-seed(42) re-trains
- training
- verifier-owned; 100 optimizer steps, global batch 32, frozen base+data; 5x same-seed(42)
- note
- absolute measured metric; baseline/theoretical_best are anchors for a separate downstream normalization layer (METHODOLOGY.md), not applied here