gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 627
- duration
- 6h 30m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__RqGcDty
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "78",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 11.7 seconds
Output:
Processed prompts: 29%|██▊ | 24/84 [09:56<22:37, 22.63s/it, est. speed input: 3.02 toks/s, output: 474.84 toks/s]
Processed prompts: 38%|███▊ | 36/96 [09:59<19:10, 19.17s/it, est. speed input: 7.89 toks/s, output: 521.66 toks/s] encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 79
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "79",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 15.5 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 80
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "80",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 15.4 seconds
Output:
Processed prompts: 12%|█▎ | 12/96 [10:58<1:16:51, 54.90s/it, est. speed input: 2.13 toks/s, output: 327.20 toks/s]
Processed prompts: 43%|████▎ | 36/84 [11:21<12:25, 15.53s/it, est. speed input: 4.11 toks/s, output: 659.14 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 81
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "81",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 10.8 seconds
Output:
Processed prompts: 57%|█████▋ | 48/84 [11:34<05:53, 9.83s/it, est. speed input: 5.82 toks/s, output: 1073.29 toks/s]
Processed prompts: 14%|█▍ | 12/84 [11:46<1:10:38, 58.86s/it, est. speed input: 1.68 toks/s, output: 336.55 toks/s]
Processed prompts: 50%|█████ | 48/96 [12:03<12:33, 15.69s/it, est. speed input: 7.80 toks/s, output: 715.08 toks/s]
Processed prompts: 62%|██████▎ | 60/96 [12:04<06:02, 10.06s/it, est. speed input: 14.68 toks/s, output: 1154.44 toks/s]
Processed prompts: 71%|███████▏ | 60/84 [12:05<02:53, 7.23s/it, est. speed input: 7.81 toks/s, output: 1350.56 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 82
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "82",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 5.3 seconds
Output:
Processed prompts: 25%|██▌ | 24/96 [12:13<31:31, 26.27s/it, est. speed input: 4.33 toks/s, output: 496.21 toks/s]
Processed prompts: 62%|██████▎ | 60/96 [12:20<06:02, 10.06s/it, est. speed input: 14.68 toks/s, output: 1154.44 toks/s]
Processed prompts: 86%|████████▌ | 72/84 [12:23<01:03, 5.28s/it, est. speed input: 9.13 toks/s, output: 1792.00 toks/s]
Processed prompts: 29%|██▊ | 24/84 [12:41<26:57, 26.96s/it, est. speed input: 3.24 toks/s, output: 595.85 toks/s] The full evaluation is progressing normally, though several problems drive generations close to the frozen 38,912-token limit. That observation reinforces the rollout-coverage change: the baseline’s 1,024-token training traces often supervise only the opening of reasoning, while evaluation depends on much longer trajectories and correctly formatted conclusions.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 83
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "83",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 14.8 seconds
Output:
Processed prompts: 43%|████▎ | 36/84 [12:49<11:56, 14.92s/it, est. speed input: 4.53 toks/s, output: 825.66 toks/s]
Processed prompts: 38%|███▊ | 36/96 [12:53<15:47, 15.79s/it, est. speed input: 7.25 toks/s, output: 816.28 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00, 4.62s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00, 4.62s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00, 9.32s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]
Processed prompts: 75%|███████▌ | 72/96 [13:12<03:25, 8.58s/it, est. speed input: 14.74 toks/s, output: 1340.66 toks/s]
Processed prompts: 50%|█████ | 48/96 [13:24<08:28, 10.59s/it, est. speed input: 8.23 toks/s, output: 1111.25 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 84
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "84",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 9.3 seconds
Output:
Processed prompts: 57%|█████▋ | 48/84 [13:25<06:08, 10.23s/it, est. speed input: 6.11 toks/s, output: 1162.91 toks/s]
Processed prompts: 88%|████████▊ | 84/96 [13:27<01:14, 6.17s/it, est. speed input: 16.28 toks/s, output: 1622.43 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:27<00:00, 6.17s/it, est. speed input: 18.80 toks/s, output: 2080.49 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:27<00:00, 8.41s/it, est. speed input: 18.80 toks/s, output: 2080.49 toks/s]
Processed prompts: 71%|███████▏ | 60/84 [13:29<02:39, 6.66s/it, est. speed input: 9.34 toks/s, output: 1504.60 toks/s]
Processed prompts: 86%|████████▌ | 72/84 [13:49<00:59, 4.95s/it, est. speed input: 10.57 toks/s, output: 1776.12 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}exec result
Script running with cell ID 85
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "85",
"yield_time_ms": 30000,
"max_tokens": 1000
}wait result
Script completed
Wall time 10.6 seconds
Output:
Processed prompts: 100%|██████████| 84/84 [14:17<00:00, 4.12s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]
Processed prompts: 100%|██████████| 84/84 [14:17<00:00, 4.12s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]
Processed prompts: 100%|██████████| 84/84 [14:17<00:00, 10.21s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]
Processed prompts: 62%|██████▎ | 60/96 [14:20<05:04, 8.45s/it, est. speed input: 9.51 toks/s, output: 1321.29 toks/s]
Processed prompts: 75%|███████▌ | 72/96 [14:41<02:28, 6.17s/it, est. speed input: 10.97 toks/s, output: 1567.38 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}exec result
Script running with cell ID 86
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "86",
"yield_time_ms": 30000,
"max_tokens": 1000
}wait result
Script completed
Wall time 14.7 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}exec result
Script running with cell ID 87
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "87",
"yield_time_ms": 30000,
"max_tokens": 1000
}wait result
Script completed
Wall time 3.1 seconds
Output:
Processed prompts: 88%|████████▊ | 84/96 [15:49<01:11, 6.00s/it, est. speed input: 11.57 toks/s, output: 1818.57 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00, 4.36s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00, 4.36s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00, 10.00s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]
FINAL avg@12 accuracy: 52.50% (189/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/domain_only_full.json
The coarse replay control scored 52.50% (189/360), below the 54.22 baseline. That rules out simply flooding training with the mixed AMC/AIME source; it contains too much easy multiple-choice material, and constant-rate training drove the asymmetric clipped objective far past zero. I’m moving to the refined hard free-response mixture, exact prompt alignment, 1,536-token rollouts, and the paper-native linear decay.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n+max_completion_length=1536\\n+lr_scheduler_type=linear\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\n"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/refined_long bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens:2000});\ntext(r.output); if (r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}exec result
Script failed
Wall time 0.0 seconds
Output:
Script error:
SyntaxError: Invalid or unexpected tokenencrypted chain-of-thought — recorded, not exposed
args
{
"input": "const result = await tools.exec_command({\n cmd: \"SEED=42 OUT_DIR=/workspace/runs/refined_long bash /workspace/train_capped.sh\",\n workdir: \"/workspace\",\n yield_time_ms: 1000,\n max_output_tokens: 2000\n});\ntext(result.output);\nif (result.session_id) text(`SESSION_ID=${result.session_id}`);\n"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=61078encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:2500});\ntext(r.output);\n"
}exec result
Script running with cell ID 91
Wall time 11.0 seconds
Output: