Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#551
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "269",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.4 seconds
Output:
#552
narration

The full evaluation is generating normally across all four single-GPU engines. Long Qwen reasoning traces make this the slowest stage; I’m waiting for the exact aggregate before deciding whether to retain trajectory balancing.

#553
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 270
Wall time 11.0 seconds
Output:
#554
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "270",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.7 seconds
Output:

Processed prompts:  25%|██▌       | 24/96 [03:51<11:05,  9.25s/it, est. speed input: 11.78 toks/s, output: 820.97 toks/s]
Processed prompts:  38%|███▊      | 36/96 [03:52<05:03,  5.06s/it, est. speed input: 20.37 toks/s, output: 1265.47 toks/s]
Processed prompts:  38%|███▊      | 36/96 [04:09<05:03,  5.06s/it, est. speed input: 20.37 toks/s, output: 1265.47 toks/s]
#555
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 271
Wall time 11.0 seconds
Output:
#556
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "271",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 45.7 seconds
Output:

Processed prompts:  14%|█▍        | 12/84 [05:29<32:54, 27.42s/it, est. speed input: 2.81 toks/s, output: 377.56 toks/s]
Processed prompts:  12%|█▎        | 12/96 [05:43<40:03, 28.62s/it, est. speed input: 5.17 toks/s, output: 318.36 toks/s]
#557
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 272
Wall time 11.0 seconds
Output:
#558
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "272",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.3 seconds
Output:

Processed prompts:  50%|█████     | 48/96 [06:15<06:14,  7.79s/it, est. speed input: 16.44 toks/s, output: 1222.86 toks/s]
#559
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 273
Wall time 11.0 seconds
Output:
#560
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "273",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.9 seconds
Output:

Processed prompts:  29%|██▊       | 24/84 [07:04<15:57, 15.96s/it, est. speed input: 4.24 toks/s, output: 627.26 toks/s]
#561
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 274
Wall time 11.0 seconds
Output:
#562
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "274",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.3 seconds
Output:

Processed prompts:  43%|████▎     | 36/84 [08:11<08:59, 11.25s/it, est. speed input: 5.68 toks/s, output: 846.56 toks/s]
#563
narration

Generation is about halfway through the slowest shard. I’m using a strict selection rule here: a merely similar one-run score will not displace the twice-replicated 56.11 method.

#564
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 275
Wall time 11.0 seconds
Output:
#565
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "275",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 47.1 seconds
Output:

Processed prompts:  62%|██████▎   | 60/96 [09:20<06:19, 10.53s/it, est. speed input: 12.91 toks/s, output: 1171.66 toks/s]
Processed prompts:  14%|█▍        | 12/84 [09:36<57:39, 48.05s/it, est. speed input: 1.75 toks/s, output: 242.87 toks/s]
Processed prompts:  75%|███████▌  | 72/96 [09:42<03:01,  7.56s/it, est. speed input: 14.00 toks/s, output: 1422.98 toks/s]
#566
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 276
Wall time 11.0 seconds
Output:
#567
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "276",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.7 seconds
Output:
#568
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 277
Wall time 11.0 seconds
Output:
#569
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "277",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.4 seconds
Output:
mpts:  88%|████████▊ | 84/96 [11:03<01:27,  7.31s/it, est. speed input: 19.80 toks/s, output: 1771.07 toks/s]
Processed prompts: 100%|██████████| 96/96 [11:27<00:00,  5.61s/it, est. speed input: 22.09 toks/s, output: 2160.49 toks/s]
Processed prompts: 100%|██████████| 96/96 [11:27<00:00,  5.61s/it, est. speed input: 22.09 toks/s, output: 2160.49 toks/s]
Processed prompts: 100%|██████████| 96/96 [11:27<00:00,  7.16s/it, est. speed input: 22.09 toks/s, output: 2160.49 toks/s]

Processed prompts:  57%|█████▋    | 48/84 [11:39<08:10, 13.63s/it, est. speed input: 6.31 toks/s, output: 943.32 toks/s]
#570
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 278
Wall time 11.0 seconds
Output:
#571
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "278",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 45.5 seconds
Output:

Processed prompts:  71%|███████▏  | 60/84 [12:15<03:54,  9.79s/it, est. speed input: 7.71 toks/s, output: 1315.12 toks/s]
Processed prompts:  25%|██▌       | 24/96 [12:26<37:51, 31.55s/it, est. speed input: 4.26 toks/s, output: 459.28 toks/s]
#572
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 279
Wall time 11.0 seconds
Output:
#573
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "279",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 46.7 seconds
Output:
prompts: 100%|██████████| 84/84 [13:17<00:00,  8.22s/it, est. speed input: 10.96 toks/s, output: 2017.63 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:17<00:00,  9.49s/it, est. speed input: 10.96 toks/s, output: 2017.63 toks/s]

Processed prompts:  29%|██▊       | 24/84 [13:27<31:07, 31.12s/it, est. speed input: 2.84 toks/s, output: 411.22 toks/s]
Processed prompts:  43%|████▎     | 36/84 [14:04<14:38, 18.31s/it, est. speed input: 4.12 toks/s, output: 682.39 toks/s]
Processed prompts:  57%|█████▋    | 48/84 [14:08<06:43, 11.20s/it, est. speed input: 5.52 toks/s, output: 1052.44 toks/s]
#574
narration

Two generation shards have completed; the remaining shards contain more max-length traces. The evaluation is healthy, just heterogeneous in runtime across AIME problems.

#575
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id: 58459, chars: \"\", yield_time_ms: 60000, max_output_tokens: 300}); text(r.output.slice(-600));\n"
}
exec result
Script running with cell ID 280
Wall time 11.0 seconds
Output: