Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#101
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "61",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 36.1 seconds
Output:

Processed prompts:  25%|▎| 24/96 [08:45<24:49, 20.68s/it, est. speed input: 6.0
#102
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 62
Wall time 11.0 seconds
Output:
#103
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "62",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 36.1 seconds
Output:

Processed prompts:  57%|▌| 48/84 [10:21<07:09, 11.93s/it, est. speed input: 6.5
#104
narration

The slowest shards are about one-quarter complete and the fastest over halfway; this is normal for AIME because trace lengths vary greatly by problem. Once the fixed-seed score lands, I’ll compare per-problem deltas against the shipped baseline to distinguish a genuine reasoning gain from a formatting or length shift.

#105
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 63
Wall time 11.0 seconds
Output:
#106
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "63",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 35.9 seconds
Output:

Processed prompts:  38%|▍| 36/96 [10:29<19:26, 19.44s/it, est. speed input: 8.4
Processed prompts:  50%|▌| 48/96 [10:44<09:48, 12.27s/it, est. speed input: 9.9
Processed prompts:  62%|▋| 60/96 [10:51<04:50,  8.06s/it, est. speed input: 11.
#107
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 64
Wall time 11.0 seconds
Output:
#108
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "64",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 34.4 seconds
Output:

Processed prompts:  14%|▏| 12/84 [11:44<1:10:27, 58.71s/it, est. speed input: 2
#109
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 65
Wall time 11.0 seconds
Output:
#110
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "65",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 36.0 seconds
Output:

Processed prompts:  71%|▋| 60/84 [12:23<04:30, 11.27s/it, est. speed input: 7.6
Processed prompts:  75%|▊| 72/96 [12:53<03:30,  8.77s/it, est. speed input: 10.
Processed prompts:  29%|▎| 24/84 [12:56<27:42, 27.70s/it, est. speed input: 3.4
Processed prompts:  43%|▍| 36/84 [13:11<12:30, 15.63s/it, est. speed input: 4.9
#111
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 66
Wall time 11.0 seconds
Output:
#112
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "66",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 35.4 seconds
Output:

Processed prompts:  86%|▊| 72/84 [13:45<01:57,  9.76s/it, est. speed input: 9.2
#113
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 67
Wall time 11.0 seconds
Output:
#114
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "67",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 36.1 seconds
Output:

Processed prompts:  57%|▌| 48/84 [14:34<07:18, 12.18s/it, est. speed input: 5.8
Processed prompts:  88%|▉| 84/96 [14:38<01:45,  8.76s/it, est. speed input: 11.
#115
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 68
Wall time 11.0 seconds
Output:
#116
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "68",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 35.4 seconds
Output:

Processed prompts: 100%|█| 84/84 [15:20<00:00,  9.15s/it, est. speed input: 9.5
Processed prompts: 100%|█| 84/84 [15:20<00:00,  9.15s/it, est. speed input: 9.5
Processed prompts: 100%|█| 84/84 [15:20<00:00, 10.95s/it, est. speed input: 9.5

Processed prompts: 100%|█| 96/96 [15:24<00:00,  7.20s/it, est. speed input: 16.
Processed prompts: 100%|█| 96/96 [15:24<00:00,  7.20s/it, est. speed input: 16.
Processed prompts: 100%|█| 96/96 [15:24<00:00,  9.63s/it, est. speed input: 16.
#117
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 69
Wall time 11.0 seconds
Output:
#118
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "69",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 34.5 seconds
Output:

Processed prompts:  71%|▋| 60/84 [16:19<04:22, 10.94s/it, est. speed input: 7.9
#119
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 70
Wall time 11.0 seconds
Output:
#120
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "70",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 33.7 seconds
Output:

Processed prompts:  38%|▍| 36/96 [17:50<31:58, 31.98s/it, est. speed input: 4.9
Processed prompts:  50%|▌| 48/96 [17:54<15:36, 19.51s/it, est. speed input: 6.1
#121
narration

Two shards have finished; the remaining two are dominated by a few very long traces, with one at about 70% and the other 50%. This confirms why the official full evaluation is much more trustworthy—and slower—than the tempting partial screen.

#122
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 71
Wall time 11.0 seconds
Output:
#123
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "71",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 34.4 seconds
Output:

Processed prompts:  86%|▊| 72/84 [18:15<02:06, 10.52s/it, est. speed input: 8.1
Processed prompts:  62%|▋| 60/96 [18:25<08:01, 13.39s/it, est. speed input: 7.3
#124
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:2739,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 72
Wall time 11.0 seconds
Output:
#125
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "72",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 34.6 seconds
Output:

Processed prompts: 100%|█| 84/84 [19:45<00:00,  9.52s/it, est. speed input: 8.3
Processed prompts: 100%|█| 84/84 [19:45<00:00,  9.52s/it, est. speed input: 8.3
Processed prompts: 100%|█| 84/84 [19:45<00:00, 14.11s/it, est. speed input: 8.3