Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
428
duration
9h 21m
recorded
2026-08-02

run idopsd-aime24-4gpus__eNDoWrD

What the verifier emitted for this run, as recorded. The fields differ by task family — the task’s verification method explains what they mean.

score
56.5556
metric
AIME24 avg@12 accuracy
direction
higher_better
baseline
54.22
theoretical best
100
practical best
100
status
ok
flags
none
reward
56.5556
passed
pass
metrics
avg at 12 pct
56.556
per run scores
56.6667, 56.9444, 55.8333, 56.3889, 56.9444
baseline mean
54.22
num runs
5
train seed
42
noise 1sigma fixed seed
1.48
measurement
avg@12 accuracy on AIME24 (30 problems), mean over 5 same-seed(42) re-trains
training
verifier-owned; 100 optimizer steps, global batch 32, frozen base+data; 5x same-seed(42)
note
absolute measured metric; baseline/theoretical_best are anchors for a separate downstream normalization layer (METHODOLOGY.md), not applied here