Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#351
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 175
Wall time 11.0 seconds
Output:
#352
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "175",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.3 seconds
Output:
Warning: truncated output (original token count: 310)
Total output lines: 1

{"chunk_id":"9f583d","wall_time_seconds":60.000828729,"session_id":85987,"original_token_count":5933,"output":"Warning: truncated output (original token count: 5933)\nTotal output lines: 154\n\n`torch_dtype` is deprecated! Use `dtype` instead!\nCheckpoint kind: lora\nMerging LoRA adapter into base weights for evaluation (CPU, once)...\n\rLoading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]\rLoading checkpoint shards: 100%|██████████| 2/2 [00:00<00:00, 51.85it/s]\n…60 tokens truncated… 100%|██████████| 8/8 [00:00<00:00, 315.87it/s]\n\rProcessed prompts:   0%|          | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]INFO 08-03 07:24:24 [llm.py:306] Supported_tasks: ['generate']\n\rAdding requests:   0%|          | 0/7 [00:00<?, ?it/s]\rAdding requests: 100%|██████████| 7/7 [00:00<00:00, 291.06it/s]\n\rProcessed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]"}
#353
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 176
Wall time 11.0 seconds
Output:
#354
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "176",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 43.6 seconds
Output:
{"chunk_id":"8b670f","wall_time_seconds":60.001433809,"session_id":85987,"original_token_count":0,"output":""}
#355
narration

The stronger-slope checkpoint is in full evaluation. If it only ties or modestly exceeds 56.11 once, I’ll still prefer the milder schedule’s two-run evidence; a slope change needs a clear margin to displace a replicated result.

#356
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 177
Wall time 11.0 seconds
Output:
#357
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "177",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 39.2 seconds
Output:
{"chunk_id":"923c11","wall_time_seconds":60.00099607,"session_id":85987,"original_token_count":32,"output":"\rProcessed prompts:  12%|█▎        | 12/96 [02:49<19:48, 14.15s/it, est. speed input: 11.02 toks/s, output: 491.58 toks/s]"}
#358
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 178
Wall time 11.0 seconds
Output:
#359
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "178",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 37.8 seconds
Output:
{"chunk_id":"7be078","wall_time_seconds":60.000836954,"session_id":85987,"original_token_count":0,"output":""}
#360
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 179
Wall time 11.0 seconds
Output:
#361
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "179",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 36.6 seconds
Output:
{"chunk_id":"531ff7","wall_time_seconds":60.00224391,"session_id":85987,"original_token_count":64,"output":"\rProcessed prompts:  25%|██▌       | 24/96 [04:11<11:46,  9.81s/it, est. speed input: 15.44 toks/s, output: 716.88 toks/s]\rProcessed prompts:  12%|█▎        | 12/96 [05:16<36:57, 26.40s/it, est. speed input: 5.61 toks/s, output: 360.03 toks/s]"}
#362
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 180
Wall time 11.0 seconds
Output:
#363
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "180",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 41.8 seconds
Output:
{"chunk_id":"69c320","wall_time_seconds":60.001111328,"session_id":85987,"original_token_count":0,"output":""}
#364
narration

Evaluation is proceeding normally, with one shard at 25% and another at 12% after their first long generations. No intermediate metric is available.

#365
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 181
Wall time 11.0 seconds
Output:
#366
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "181",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 41.8 seconds
Output:
{"chunk_id":"0a75b8","wall_time_seconds":60.001904435,"session_id":85987,"original_token_count":32,"output":"\rProcessed prompts:  14%|█▍        | 12/84 [07:38<45:52, 38.23s/it, est. speed input: 1.91 toks/s, output: 274.24 toks/s]"}
#367
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 182
Wall time 11.0 seconds
Output:
#368
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "182",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 40.7 seconds
Output:
{"chunk_id":"3492b5","wall_time_seconds":60.001944,"session_id":85987,"original_token_count":97,"output":"\rProcessed prompts:  38%|███▊      | 36/96 [08:09<14:23, 14.39s/it, est. speed input: 9.67 toks/s, output: 640.68 toks/s] \rProcessed prompts:  29%|██▊       | 24/84 [08:16<17:34, 17.58s/it, est. speed input: 3.77 toks/s, output: 583.23 toks/s]\rProcessed prompts:  43%|████▎     | 36/84 [08:43<08:28, 10.59s/it, est. speed input: 7.29 toks/s, output: 956.47 toks/s]"}
#369
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 183
Wall time 11.0 seconds
Output:
#370
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "183",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 43.8 seconds
Output:
{"chunk_id":"befeee","wall_time_seconds":60.001334526,"session_id":85987,"original_token_count":34,"output":"\rProcessed prompts:  57%|█████▋    | 48/84 [09:08<04:20,  7.24s/it, est. speed input: 8.64 toks/s, output: 1159.46 toks/s]"}
#371
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 184
Wall time 11.0 seconds
Output:
#372
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "184",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 38.8 seconds
Output:
Warning: truncated output (original token count: 256)
Total output lines: 1

{"chunk_id":"30f638","wall_time_seconds":60.000678315,"session_id":85987,"original_token_count":235,"output":"Warning: truncated output (original token count: 235)\nTotal output lines: 1\n\n\rProcessed prompts:  50%|█████     | 48/96 [10:16<10:19, 12.91s/it, est. speed input: 9.14 toks/s, output: 783.15 toks/s]\rProcessed prompts:  62%|██████▎   | 60/96 [10:30<05:11,  8.65s…56 tokens truncated…5%|███████▌  | 72/96 [10:43<02:26,  6.09s/it, est. speed input: 12.67 toks/s, output: 1373.37 toks/s]\rProcessed prompts:  71%|███████▏  | 60/84 [10:55<03:07,  7.83s/it, est. speed input: 9.14 toks/s, output: 1428.47 toks/s]\rProcessed prompts:  88%|████████▊ | 84/96 [11:06<00:56,  4.73s/it, est. speed input: 19.71 toks/s, output: 1829.70 toks/s]"}
#373
narration

Two shards have accelerated past 70–88%; the others are still working through longer problems. I’m preserving the full official settings for a meaningful comparison.

#374
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:85987,chars:\"\",yield_time_ms:60000,max_output_tokens:200});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 185
Wall time 11.0 seconds
Output:
#375
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "185",
  "yield_time_ms": 60000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 43.7 seconds
Output:
Warning: truncated output (original token count: 256)
Total output lines: 1

{"chunk_id":"c1ed52","wall_time_seconds":60.001304762,"session_id":85987,"original_token_count":208,"output":"Warning: truncated output (original token count: 208)\nTotal output lines: 1\n\n\rProcessed prompts:  38%|███▊      | 36/96 [11:48<17:08, 17.14s/it, est. speed input: 5.93 toks/s, output: 821.73 toks/s]\rProcessed prompts:  86%|████████▌ | 72/84 [11:55<01:22,  6.8…56 tokens truncated…██████| 96/96 [12:19<00:00,  5.16s/it, est. speed input: 20.53 toks/s, output: 2143.98 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [12:19<00:00,  5.16s/it, est. speed input: 20.53 toks/s, output: 2143.98 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [12:19<00:00,  7.70s/it, est. speed input: 20.53 toks/s, output: 2143.98 toks/s]\n"}