gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 627
- duration
- 6h 30m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__RqGcDty
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "284",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 7.2 seconds
Output:
Processed prompts: 57%|█████▋ | 48/84 [14:52<08:24, 14.01s/it, est. speed input: 5.53 toks/s, output: 1076.40 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 285
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "285",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 12.6 seconds
Output:
Processed prompts: 62%|██████▎ | 60/96 [14:54<06:07, 10.20s/it, est. speed input: 9.07 toks/s, output: 1246.45 toks/s]
Processed prompts: 71%|███████▏ | 60/84 [14:59<03:39, 9.16s/it, est. speed input: 6.80 toks/s, output: 1365.75 toks/s]
Processed prompts: 86%|████████▌ | 72/84 [15:11<01:16, 6.39s/it, est. speed input: 9.61 toks/s, output: 1690.22 toks/s]
Processed prompts: 100%|██████████| 84/84 [15:13<00:00, 4.35s/it, est. speed input: 10.89 toks/s, output: 2063.97 toks/s]
Processed prompts: 100%|██████████| 84/84 [15:13<00:00, 4.35s/it, est. speed input: 10.89 toks/s, output: 2063.97 toks/s]
Processed prompts: 100%|██████████| 84/84 [15:13<00:00, 10.87s/it, est. speed input: 10.89 toks/s, output: 2063.97 toks/s]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 286
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "286",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 12.0 seconds
Output:
Processed prompts: 75%|███████▌ | 72/96 [15:34<03:09, 7.88s/it, est. speed input: 10.35 toks/s, output: 1471.13 toks/s]
Processed prompts: 88%|████████▊ | 84/96 [15:51<01:09, 5.76s/it, est. speed input: 11.54 toks/s, output: 1791.00 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.0 seconds
Output:
Processed prompts: 100%|██████████| 96/96 [16:10<00:00, 4.42s/it, est. speed input: 13.48 toks/s, output: 2130.15 toks/s]
Processed prompts: 100%|██████████| 96/96 [16:10<00:00, 4.42s/it, est. speed input: 13.48 toks/s, output: 2130.15 toks/s]
Processed prompts: 100%|██████████| 96/96 [16:10<00:00, 10.11s/it, est. speed input: 13.48 toks/s, output: 2130.15 toks/s]
FINAL avg@12 accuracy: 55.28% (199/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/refined_repeat2_full.json
The independent repeat scored 55.28% (199/360). Across two full retrains, the method averages 56.25%, +2.03 over the five-run baseline anchor; both runs beat 54.22 individually. I’m running one final replicate to reduce uncertainty further before finalizing.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/refined_repeat3 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output); if(r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=95533encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 289
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "289",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 8.3 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 22
Warning: truncated output (original token count: 2565)
Total output lines: 153
<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.
================================================================================
RUN CONFIGURATION
================================================================================
WandB Run Name: capped_seed42_lr5e-6_bs32
Output Directory: /workspace/runs/refined_repeat3/capped_seed42
================================================================================
===========================…27 tokens truncated…s connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0
Loading safetensors checkpoint shards: 0% Completed | 0/2 [00:00<?, ?it/s]
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00, 3.34it/s]
Loading safetensors checkpoint shards: 100% Completed | 2/2 [00:00<00:00, 3.34it/s]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:120});\ntext(r.output);\n"
}exec result
Script running with cell ID 290
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "290",
"yield_time_ms": 30000,
"max_tokens": 120
}wait result
Script completed
Wall time 2.7 seconds
Output:
Warning: truncated output (original token count: 147)
Total output lines: 5
Warning: truncated output (original token count: 1520)
Total output lines: 42
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/5 [00:00<?, ?it/s]
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 20%|█…27 tokens truncated…0:23<19:26, 11.90s/it]
{'loss': 0.0067, 'grad_norm': 0.15881772339344025, 'learning_rate': 4.95e-06, 'on_policy_loss': 0.0067, 'epoch': 0.02}
2%|▏ | 2/100 [00:24<19:26, 11.90s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 291
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "291",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 9.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 899)
Total output lines: 29
vLLM generation done - elapsed time: 5.04s, prompts: 4, total tokens: 4417, avg le…26 tokens truncated…66, avg length: 1266.5, speed: 1002.6 tok/s
vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5909, avg length: 1477.2, speed: 1164.9 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 292
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "292",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 5.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1947)
Total output lines: 145
vLLM generation done - elapsed time: 4.95s, prompts: 4, total tokens: 4691, avg …27 tokens truncated…4407, avg length: 1101.8, speed: 885.9 tok/s
vLLM generation done - elapsed time: 4.99s, prompts: 4, total tokens: 4734, avg length: 1183.5, speed: 948.1 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 293
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "293",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 0.6 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1555)
Total output lines: 178
vLLM generation done - elapsed time: 4.04s, prompts: 4, total tokens: 4175, avg …27 tokens truncated… 2590, avg length: 647.5, speed: 893.7 tok/s
vLLM generation done - elapsed time: 4.90s, prompts: 4, total tokens: 4130, avg length: 1032.5, speed: 842.5 tok/s
The third replicate is training normally with the same code and recipe. The final repository remains in the validated state: matched hard-integer replay, exact evaluator prompt, 1,536-token rollouts, and linear decay.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 294
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "294",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 12.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6
Warning: truncated output (original token count: 921)
Total output lines: 29
vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5546, avg le…26 tokens truncated…ration done - elapsed time: 5.08s, prompts: 4, total tokens: 5580, avg length: 1395.0, speed: 1097.9 tok/s
15%|█▌ | 15/100 [02:57<16:43, 11.80s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 295
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "295",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 9.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1027)
Total output lines: 36
vLLM generation done - elapsed time: 4.92s, prompts: 4, total tokens: 4525, avg l…27 tokens truncated…4527, avg length: 1131.8, speed: 910.1 tok/s
vLLM generation done - elapsed time: 4.98s, prompts: 4, total tokens: 4532, avg length: 1133.0, speed: 910.2 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:95533,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 296
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "296",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 2.8 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 2562)
Total output lines: 208
vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4710, avg …27 tokens truncated…70, avg length: 1342.5, speed: 1068.3 tok/s
vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5876, avg length: 1469.0, speed: 1156.0 tok/s