Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
428
duration
9h 21m
recorded
2026-08-02

run idopsd-aime24-4gpus__eNDoWrD

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol444 turns16 narration428 reasoning428 action428 observation1 final
#351
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "110",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 244)
Total output lines: 16

     | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]


 [02:52<20:09, 14.40s/it, est. speed input: 11.60 toks/s, output: 547.49 toks…164 tokens truncated…13:24<06:46, 11.28s/it, est. speed input: 7.92 toks/s, output: 1168.79 toks/s]
[14:34<03:16,  8.19s/it, est. speed input: 11.06 toks/s, output: 1479.72 toks/s]
#352
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<6;i++){const r=await tools.write_stdin({session_id:26063,chars:\"\",yield_time_ms:60000,max_output_tokens:70});out+=r.output.slice(-140)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 111
Wall time 11.0 seconds
Output:
#353
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "111",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 23.5 seconds
Output:
INAL avg@12 accuracy: 55.56% (200/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/dev_eval_1785709356.json

EXIT0
#354
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"cat /workspace/.timer/remaining_secs && python - <<'PY'\\nimport json\\nf='/workspace/eval_results/dev_eval_1785709356.json';d=json.load(open(f));print([(x['problem_id'],x['num_correct']) for x in d['results']])\\nPY\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
6529
[(60, 12), (61, 12), (62, 0), (63, 0), (64, 3), (65, 4), (66, 10), (67, 12), (68, 11), (69, 12), (70, 6), (71, 11), (72, 12), (73, 0), (74, 10), (75, 3), (76, 11), (77, 0), (78, 6), (79, 12), (80, 3), (81, 0), (82, 7), (83, 12), (84, 12), (85, 0), (86, 12), (87, 7), (88, 0), (89, 0)]
#355
narration

The reasoning-effort curriculum is the first full-run improvement: 55.56% (200/360), +1.11 over the matched 54.44 runs and +1.34 over the five-run baseline anchor. That is promising but within single-run training noise, so I’m repeating the exact same seed-42 train and full evaluation rather than tuning on this one outcome.

#356
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/hard3000_repeat PORT=12960 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=15173
#357
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<14;i++){const r=await tools.write_stdin({session_id:15173,chars:\"\",yield_time_ms:60000,max_output_tokens:40});out+=r.output.slice(-80)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 114
Wall time 11.0 seconds
Output:
#358
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#359
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#360
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#361
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#362
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#363
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#364
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 114
Wall time 61.0 seconds
Output:
#365
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 42.4 seconds
Output:
Warning: truncated output (original token count: 301)
Total output lines: 28

 3.42s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1198.6 tok/s

d…oss': 0.0056, 'epoch': 0.02}

 10%|█         | 10/100 [01:22<12:14,  8.16…221 tokens truncated…l tokens: 4096, avg length: 1024.0, speed: 1201.9 tok/s

94 tokens truncated…ch': 0.23}

 98%|█████████▊| 98/100 [13:22<00:16,  8.20s/it]
#366
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:15173,chars:\"\",yield_time_ms:60000,max_output_tokens:200});text(r.output);\n"
}
exec result
Script running with cell ID 115
Wall time 11.0 seconds
Output:
#367
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "115",
  "yield_time_ms": 60000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 3.4 seconds
Output:
Warning: truncated output (original token count: 936)
Total output lines: 24

vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3829, avg length: 957.2, speed: 1136.9 tok/s
vLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1204.4 tok/s
vLLM generation done - elapsed time: 3.41s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1201.1 tok/s
vLLM generation done - elapsed time: 3.41s, prompts: …736 tokens truncated…cs/stable/distributed.html#shutdown (function operator())
[rank1]:[W802 22:53:40.788583908 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/hard3000_repeat/capped_seed42/checkpoint-100
#368
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/hard3000_repeat/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":300});\ntext(r.output);if(r.session_id)text(`SESSION_ID=${r.session_id}`);"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=17634
#369
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<17;i++){const r=await tools.write_stdin({session_id:17634,chars:\"\",yield_time_ms:60000,max_output_tokens:40});out+=r.output.slice(-80)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 117
Wall time 11.0 seconds
Output:
#370
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output:
#371
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output:
#372
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output:
#373
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output:
#374
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output:
#375
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "117",
  "yield_time_ms": 60000,
  "max_tokens": 80
}
wait result
Script running with cell ID 117
Wall time 61.0 seconds
Output: