Skip to content
on-policy-self-distillation

claude-opus-5

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
claude claude-opus-5
metric
AIME24 avg@12 accuracy
actions
110
duration
6h 10m
recorded
2026-08-02

run idopsd-aime24-4gpus__F2Xe6j3

Everything claude-opus-5 did on this task, in order.

download JSON ↗
claude-codeclaude-opus-5101 turns15 narration109 action109 observation86 reasoning1 final
#101
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolBashPoll try8 result
args
{
  "command": "sleep 420; python /workspace/summarize.py; cat /workspace/.timer/remaining_secs",
  "description": "Poll try8 result",
  "timeout": 450000
}
Bash result
run                      avg@n  pass@n    fmt%  probs   gens
try1                     57.78   76.67   99.72     30    360
try2                     42.78   80.00   99.72     30    360
try3                     45.56   73.33   99.17     30    360
try4                     51.67   76.67   98.89     30    360
try5                     52.78   76.67  100.00     30    360
try8                     50.00   76.67   99.72     30    360

baseline anchor: 54.22 +- 1.48 (single-run 1 sigma ~1.5)
8421

[stdout]
run                      avg@n  pass@n    fmt%  probs   gens
try1                     57.78   76.67   99.72     30    360
try2                     42.78   80.00   99.72     30    360
try3                     45.56   73.33   99.17     30    360
try4                     51.67   76.67   98.89     30    360
try5                     52.78   76.67  100.00     30    360
try8                     50.00   76.67   99.72     30    360

baseline anchor: 54.22 +- 1.48 (single-run 1 sigma ~1.5)
8421
[metadata] {"noOutputExpected": false}
final summary

try8 (top_k_loss=20) came in at 50.00 — no help. Killing its repeat to spend the remaining time on the largest untested train/eval mismatch: training rollouts are generated in non-thinking mode while the eval runs in thinking mode.