on-policy-self-distillation
claude-opus-5
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- AIME24 avg@12 accuracy
- actions
- 110
- duration
- 6h 10m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__F2Xe6j3
Everything claude-opus-5 did on this task, in order.
download JSON ↗claude-codeclaude-opus-5101 turns15 narration109 action109 observation86 reasoning1 final
#101
reasoning (redacted)
encrypted chain-of-thought — recorded, not exposed
toolBashPoll try8 result
args
{
"command": "sleep 420; python /workspace/summarize.py; cat /workspace/.timer/remaining_secs",
"description": "Poll try8 result",
"timeout": 450000
}Bash result
run avg@n pass@n fmt% probs gens
try1 57.78 76.67 99.72 30 360
try2 42.78 80.00 99.72 30 360
try3 45.56 73.33 99.17 30 360
try4 51.67 76.67 98.89 30 360
try5 52.78 76.67 100.00 30 360
try8 50.00 76.67 99.72 30 360
baseline anchor: 54.22 +- 1.48 (single-run 1 sigma ~1.5)
8421
[stdout]
run avg@n pass@n fmt% probs gens
try1 57.78 76.67 99.72 30 360
try2 42.78 80.00 99.72 30 360
try3 45.56 73.33 99.17 30 360
try4 51.67 76.67 98.89 30 360
try5 52.78 76.67 100.00 30 360
try8 50.00 76.67 99.72 30 360
baseline anchor: 54.22 +- 1.48 (single-run 1 sigma ~1.5)
8421
[metadata] {"noOutputExpected": false}final summary
try8 (top_k_loss=20) came in at 50.00 — no help. Killing its repeat to spend the remaining time on the largest untested train/eval mismatch: training rollouts are generated in non-thinking mode while the eval runs in thinking mode.