Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#126
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "78",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 11.7 seconds
Output:

Processed prompts:  29%|██▊       | 24/84 [09:56<22:37, 22.63s/it, est. speed input: 3.02 toks/s, output: 474.84 toks/s]
Processed prompts:  38%|███▊      | 36/96 [09:59<19:10, 19.17s/it, est. speed input: 7.89 toks/s, output: 521.66 toks/s] 
#127
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 79
Wall time 11.0 seconds
Output:
#128
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "79",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 15.5 seconds
Output:
#129
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 80
Wall time 11.0 seconds
Output:
#130
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 15.4 seconds
Output:

Processed prompts:  12%|█▎        | 12/96 [10:58<1:16:51, 54.90s/it, est. speed input: 2.13 toks/s, output: 327.20 toks/s]
Processed prompts:  43%|████▎     | 36/84 [11:21<12:25, 15.53s/it, est. speed input: 4.11 toks/s, output: 659.14 toks/s]
#131
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 81
Wall time 11.0 seconds
Output:
#132
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "81",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 10.8 seconds
Output:

Processed prompts:  57%|█████▋    | 48/84 [11:34<05:53,  9.83s/it, est. speed input: 5.82 toks/s, output: 1073.29 toks/s]
Processed prompts:  14%|█▍        | 12/84 [11:46<1:10:38, 58.86s/it, est. speed input: 1.68 toks/s, output: 336.55 toks/s]
Processed prompts:  50%|█████     | 48/96 [12:03<12:33, 15.69s/it, est. speed input: 7.80 toks/s, output: 715.08 toks/s]
Processed prompts:  62%|██████▎   | 60/96 [12:04<06:02, 10.06s/it, est. speed input: 14.68 toks/s, output: 1154.44 toks/s]
Processed prompts:  71%|███████▏  | 60/84 [12:05<02:53,  7.23s/it, est. speed input: 7.81 toks/s, output: 1350.56 toks/s]
#133
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 82
Wall time 11.0 seconds
Output:
#134
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 5.3 seconds
Output:

Processed prompts:  25%|██▌       | 24/96 [12:13<31:31, 26.27s/it, est. speed input: 4.33 toks/s, output: 496.21 toks/s]  
Processed prompts:  62%|██████▎   | 60/96 [12:20<06:02, 10.06s/it, est. speed input: 14.68 toks/s, output: 1154.44 toks/s]
Processed prompts:  86%|████████▌ | 72/84 [12:23<01:03,  5.28s/it, est. speed input: 9.13 toks/s, output: 1792.00 toks/s]
Processed prompts:  29%|██▊       | 24/84 [12:41<26:57, 26.96s/it, est. speed input: 3.24 toks/s, output: 595.85 toks/s]  
#135
narration

The full evaluation is progressing normally, though several problems drive generations close to the frozen 38,912-token limit. That observation reinforces the rollout-coverage change: the baseline’s 1,024-token training traces often supervise only the opening of reasoning, while evaluation depends on much longer trajectories and correctly formatted conclusions.

#136
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 83
Wall time 11.0 seconds
Output:
#137
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "83",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 14.8 seconds
Output:

Processed prompts:  43%|████▎     | 36/84 [12:49<11:56, 14.92s/it, est. speed input: 4.53 toks/s, output: 825.66 toks/s]
Processed prompts:  38%|███▊      | 36/96 [12:53<15:47, 15.79s/it, est. speed input: 7.25 toks/s, output: 816.28 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00,  4.62s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00,  4.62s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]
Processed prompts: 100%|██████████| 84/84 [13:02<00:00,  9.32s/it, est. speed input: 11.16 toks/s, output: 2048.84 toks/s]

Processed prompts:  75%|███████▌  | 72/96 [13:12<03:25,  8.58s/it, est. speed input: 14.74 toks/s, output: 1340.66 toks/s]
Processed prompts:  50%|█████     | 48/96 [13:24<08:28, 10.59s/it, est. speed input: 8.23 toks/s, output: 1111.25 toks/s]
#138
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 84
Wall time 11.0 seconds
Output:
#139
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "84",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 9.3 seconds
Output:

Processed prompts:  57%|█████▋    | 48/84 [13:25<06:08, 10.23s/it, est. speed input: 6.11 toks/s, output: 1162.91 toks/s]
Processed prompts:  88%|████████▊ | 84/96 [13:27<01:14,  6.17s/it, est. speed input: 16.28 toks/s, output: 1622.43 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:27<00:00,  6.17s/it, est. speed input: 18.80 toks/s, output: 2080.49 toks/s]
Processed prompts: 100%|██████████| 96/96 [13:27<00:00,  8.41s/it, est. speed input: 18.80 toks/s, output: 2080.49 toks/s]

Processed prompts:  71%|███████▏  | 60/84 [13:29<02:39,  6.66s/it, est. speed input: 9.34 toks/s, output: 1504.60 toks/s]
Processed prompts:  86%|████████▌ | 72/84 [13:49<00:59,  4.95s/it, est. speed input: 10.57 toks/s, output: 1776.12 toks/s]
#140
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}
exec result
Script running with cell ID 85
Wall time 11.0 seconds
Output:
#141
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "85",
  "yield_time_ms": 30000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 10.6 seconds
Output:

Processed prompts: 100%|██████████| 84/84 [14:17<00:00,  4.12s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]
Processed prompts: 100%|██████████| 84/84 [14:17<00:00,  4.12s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]
Processed prompts: 100%|██████████| 84/84 [14:17<00:00, 10.21s/it, est. speed input: 11.60 toks/s, output: 2099.95 toks/s]

Processed prompts:  62%|██████▎   | 60/96 [14:20<05:04,  8.45s/it, est. speed input: 9.51 toks/s, output: 1321.29 toks/s]
Processed prompts:  75%|███████▌  | 72/96 [14:41<02:28,  6.17s/it, est. speed input: 10.97 toks/s, output: 1567.38 toks/s]
#142
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}
exec result
Script running with cell ID 86
Wall time 11.0 seconds
Output:
#143
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "86",
  "yield_time_ms": 30000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 14.7 seconds
Output:
#144
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:99639,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}
exec result
Script running with cell ID 87
Wall time 11.0 seconds
Output:
#145
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "87",
  "yield_time_ms": 30000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 3.1 seconds
Output:

Processed prompts:  88%|████████▊ | 84/96 [15:49<01:11,  6.00s/it, est. speed input: 11.57 toks/s, output: 1818.57 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00,  4.36s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00,  4.36s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]
Processed prompts: 100%|██████████| 96/96 [15:59<00:00, 10.00s/it, est. speed input: 13.63 toks/s, output: 2150.21 toks/s]

FINAL avg@12 accuracy: 52.50% (189/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/domain_only_full.json
#146
narration

The coarse replay control scored 52.50% (189/360), below the 54.22 baseline. That rules out simply flooding training with the mixed AMC/AIME source; it contains too much easy multiple-choice material, and constant-rate training drove the asymmetric clipped objective far past zero. I’m moving to the refined hard free-response mixture, exact prompt alignment, 1,536-token rollouts, and the paper-native linear decay.

#147
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n+max_completion_length=1536\\n+lr_scheduler_type=linear\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{}
#148
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/refined_long bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens:2000});\ntext(r.output); if (r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}
exec result
Script failed
Wall time 0.0 seconds
Output:
Script error:
SyntaxError: Invalid or unexpected token
#149
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const result = await tools.exec_command({\n  cmd: \"SEED=42 OUT_DIR=/workspace/runs/refined_long bash /workspace/train_capped.sh\",\n  workdir: \"/workspace\",\n  yield_time_ms: 1000,\n  max_output_tokens: 2000\n});\ntext(result.output);\nif (result.session_id) text(`SESSION_ID=${result.session_id}`);\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=61078
#150
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:2500});\ntext(r.output);\n"
}
exec result
Script running with cell ID 91
Wall time 11.0 seconds
Output: