Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#251
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "142",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 12.9 seconds
Output:

Processed prompts:  29%|██▊       | 24/84 [06:27<14:37, 14.63s/it, est. speed input: 4.96 toks/s, output: 778.78 toks/s]
#252
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 143
Wall time 11.0 seconds
Output:
#253
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "143",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 8.1 seconds
Output:

Processed prompts:  12%|█▎        | 12/96 [07:17<51:00, 36.44s/it, est. speed input: 4.06 toks/s, output: 291.88 toks/s]
#254
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 144
Wall time 11.0 seconds
Output:
#255
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "144",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 10.2 seconds
Output:
#256
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 145
Wall time 11.0 seconds
Output:
#257
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "145",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 7.9 seconds
Output:
#258
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 146
Wall time 11.0 seconds
Output:
#259
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "146",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 11.5 seconds
Output:

Processed prompts:  43%|████▎     | 36/84 [09:25<11:47, 14.74s/it, est. speed input: 4.94 toks/s, output: 766.54 toks/s]
#260
narration

The full evaluation is still running; some shards are slow because their sampled solutions hit the 38,912-token cap, which is expected under the frozen settings. No partial score is being used for selection.

#261
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 147
Wall time 11.0 seconds
Output:
#262
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "147",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 12.5 seconds
Output:

Processed prompts:  38%|███▊      | 36/96 [09:43<18:43, 18.73s/it, est. speed input: 8.11 toks/s, output: 544.59 toks/s] 
Processed prompts:  14%|█▍        | 12/84 [09:50<59:03, 49.22s/it, est. speed input: 1.71 toks/s, output: 246.08 toks/s]
Processed prompts:  25%|██▌       | 24/96 [10:14<28:27, 23.71s/it, est. speed input: 5.17 toks/s, output: 547.92 toks/s]
#263
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 148
Wall time 11.0 seconds
Output:
#264
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "148",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 6.5 seconds
Output:

Processed prompts:  50%|█████     | 48/96 [10:57<11:02, 13.79s/it, est. speed input: 9.40 toks/s, output: 816.85 toks/s]
#265
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 149
Wall time 11.0 seconds
Output:
#266
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "149",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 15.2 seconds
Output:

Processed prompts:  62%|██████▎   | 60/96 [11:29<05:51,  9.77s/it, est. speed input: 10.50 toks/s, output: 1081.60 toks/s]
#267
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 150
Wall time 11.0 seconds
Output:
#268
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "150",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 15.3 seconds
Output:

Processed prompts:  57%|█████▋    | 48/84 [11:48<08:10, 13.63s/it, est. speed input: 6.23 toks/s, output: 1010.37 toks/s]
Processed prompts:  71%|███████▏  | 60/84 [12:01<03:38,  9.10s/it, est. speed input: 7.85 toks/s, output: 1385.27 toks/s]
Processed prompts:  75%|███████▌  | 72/96 [12:03<02:58,  7.42s/it, est. speed input: 11.26 toks/s, output: 1329.90 toks/s]
#269
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 151
Wall time 11.0 seconds
Output:
#270
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "151",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 14.5 seconds
Output:

Processed prompts:  86%|████████▌ | 72/84 [12:21<01:18,  6.55s/it, est. speed input: 10.27 toks/s, output: 1674.04 toks/s]
Processed prompts:  88%|████████▊ | 84/96 [12:25<01:07,  5.59s/it, est. speed input: 17.63 toks/s, output: 1752.87 toks/s]
Processed prompts: 100%|██████████| 84/84 [12:52<00:00,  5.28s/it, est. speed input: 11.30 toks/s, output: 2066.10 toks/s]
Processed prompts: 100%|██████████| 84/84 [12:52<00:00,  5.28s/it, est. speed input: 11.30 toks/s, output: 2066.10 toks/s]
Processed prompts: 100%|██████████| 84/84 [12:52<00:00,  9.20s/it, est. speed input: 11.30 toks/s, output: 2066.10 toks/s]
#271
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 152
Wall time 11.0 seconds
Output:
#272
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "152",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 8.8 seconds
Output:

Processed prompts: 100%|██████████| 96/96 [13:25<00:00,  5.40s/it, est. speed input: 18.85 toks/s, output: 2067.87 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:25<00:00,  5.40s/it, est. speed input: 18.85 toks/s, output: 2067.87 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:25<00:00,  8.39s/it, est. speed input: 18.85 toks/s, output: 2067.87 toks/s]
#273
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63518,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 153
Wall time 11.0 seconds
Output:
#274
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "153",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 7.3 seconds
Output:

Processed prompts:  38%|███▊      | 36/96 [13:40<20:41, 20.70s/it, est. speed input: 6.83 toks/s, output: 751.97 toks/s]
#275
narration

Two of four evaluation shards have completed; the remaining shards contain longer generations but are healthy. The scorer uses the merged 360-result file, so I’m waiting for all shards before launching the new loss run.