Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#351
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "190",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 666)
Total output lines: 21

vLLM generation done - elapsed time: 5.56s, prompts: 4, total tokens: 4290, avg le…26 tokens truncated…5643, avg length: 1410.8, speed: 841.4 tok/s
vLLM generation done - elapsed time: 6.72s, prompts: 4, total tokens: 6232, avg length: 1558.0, speed: 927.1 tok/s
#352
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 191
Wall time 11.0 seconds
Output:
#353
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "191",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.2 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 555)
Total output lines: 17

vLLM generation done - elapsed time: 3.76s, prompts: 4, total tokens: 3130, avg le…26 tokens truncated…405, avg length: 1351.2, speed: 864.1 tok/s
vLLM generation done - elapsed time: 6.84s, prompts: 4, total tokens: 7178, avg length: 1794.5, speed: 1049.3 tok/s
#354
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 192
Wall time 11.0 seconds
Output:
#355
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "192",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.3 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 814)
Total output lines: 29

vLLM generation done - elapsed time: 4.27s, prompts: 4, total tokens: 3590, avg le…26 tokens truncated…88223695755, 'learning_rate': 1.1500000000000002e-06, 'on_policy_loss': 0.0114, 'epoch': 0.68}

 78%|███████▊  | 78/100 [19:58<05:37, 15.35s/it]
#356
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 193
Wall time 11.0 seconds
Output:
#357
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "193",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 669)
Total output lines: 21

vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 4246, avg le…26 tokens truncated…5361, avg length: 1340.2, speed: 821.1 tok/s
vLLM generation done - elapsed time: 6.73s, prompts: 4, total tokens: 6063, avg length: 1515.8, speed: 900.8 tok/s
#358
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 194
Wall time 11.0 seconds
Output:
#359
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "194",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 873)
Total output lines: 31

vLLM generation done - elapsed time: 5.43s, prompts: 4, total tokens: 3260, avg le…26 tokens truncated…708, avg length: 1677.0, speed: 1005.0 tok/s
vLLM generation done - elapsed time: 6.69s, prompts: 4, total tokens: 5491, avg length: 1372.8, speed: 821.3 tok/s
#360
narration

The corrected loss run is at step 84 and remains healthy. I’ll evaluate it on all 360 samples, then compare against the already validated 57.22% candidate and retain the stronger method; the existing result is preserved independently of current source edits.

#361
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 195
Wall time 11.0 seconds
Output:
#362
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "195",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.0 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 3278)
Total output lines: 146

vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 4841, avg …27 tokens truncated…rm': 0.04028976336121559, 'learning_rate': 7.5e-07, 'on_policy_loss': 0.0105, 'epoch': 0.75}

 86%|████████▌ | 86/100 [22:00<03:34, 15.31s/it]
#363
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 196
Wall time 11.0 seconds
Output:
#364
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "196",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.0 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2329)
Total output lines: 265

vLLM generation done - elapsed time: 6.47s, prompts: 4, total tokens: 4984, avg …27 tokens truncated…elapsed time: 6.82s, prompts: 4, total tokens: 7330, avg length: 1832.5, speed: 1074.5 tok/s

 89%|████████▉ | 89/100 [22:46<02:47, 15.26s/it]
#365
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 197
Wall time 11.0 seconds
Output:
#366
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "197",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 767)
Total output lines: 28

vLLM generation done - elapsed time: 6.08s, prompts: 4, total tokens: 4685, avg le…26 tokens truncated…77, avg length: 1744.2, speed: 1022.8 tok/s
vLLM generation done - elapsed time: 4.81s, prompts: 4, total tokens: 4842, avg length: 1210.5, speed: 1006.7 tok/s
#367
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 198
Wall time 11.0 seconds
Output:
#368
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "198",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.0 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2410)
Total output lines: 300

vLLM generation done - elapsed time: 6.83s, prompts: 4, total tokens: 6883, avg …27 tokens truncated…lapsed time: 6.47s, prompts: 4, total tokens: 5364, avg length: 1341.0, speed: 828.8 tok/s

 95%|█████████▌| 95/100 [24:18<01:16, 15.40s/it]
#369
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 199
Wall time 11.0 seconds
Output:
#370
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "199",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 6.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 938)
Total output lines: 33

vLLM generation done - elapsed time: 5.59s, prompts: 4, total tokens: 4041, avg le…26 tokens truncated…3331909, 'learning_rate': 1.5000000000000002e-07, 'on_policy_loss': 0.0111, 'epoch': 0.86}

 98%|█████████▊| 98/100 [25:05<00:30, 15.42s/it]
#371
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:60664,chars:\"\",yield_time_ms:30000,max_output_tokens:1000});\ntext(r.output);\n"
}
exec result
Script running with cell ID 200
Wall time 11.0 seconds
Output:
#372
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "200",
  "yield_time_ms": 30000,
  "max_tokens": 1000
}
wait result
Script completed
Wall time 5.0 seconds
Output:
vLLM generation done - elapsed time: 4.91s, prompts: 4, total tokens: 4723, avg length: 1180.8, speed: 962.4 tok/s
vLLM generation done - elapsed time: 5.60s, prompts: 4, total tokens: 4846, avg length: 1211.5, speed: 865.9 tok/s
vLLM generation done - elapsed time: 5.94s, prompts: 4, total tokens: 4932, avg length: 1233.0, speed: 830.4 tok/s
vLLM generation done - elapsed time: 6.32s, prompts: 4, total tokens: 4996, avg length: 1249.0, speed: 790.2 tok/s
vLLM generation done - elapsed time: 3.84s, prompts: 4, total tokens: 3880, avg length: 970.0, speed: 1010.8 tok/s
vLLM generation done - elapsed time: 5.37s, prompts: 4, total tokens: 4428, avg length: 1107.0, speed: 825.2 tok/s
vLLM generation done - elapsed time: 6.50s, prompts: 4, total tokens: 5249, avg length: 1312.2, speed: 807.7 tok/s
vLLM generation done - elapsed time: 6.79s, prompts: 4, total tokens: 6813, avg length: 1703.2, speed: 1003.9 tok/s

 99%|█████████▉| 99/100 [25:20<00:15, 15.28s/it]vLLM generation done - elapsed time: 4.03s, prompts: 4, total tokens: 4207, avg length: 1051.8, speed: 1043.8 tok/s
vLLM generation done - elapsed time: 6.40s, prompts: 4, total tokens: 4750, avg length: 1187.5, speed: 742.7 tok/s
vLLM generation done - elapsed time: 6.69s, prompts: 4, total tokens: 5773, avg length: 1443.2, speed: 863.2 tok/s
vLLM generation done - elapsed time: 6.81s, prompts: 4, total tokens: 7456, avg length: 1864.0, speed: 1094.9 tok/s
vLLM generation done - elapsed time: 4.23s, prompts: 4, total tokens: 4240, avg length: 1060.0, speed: 1002.2 tok/s
vLLM generation done - elapsed time: 5.09s, prompts: 4, total tokens: 4481, avg length: 1120.2, speed: 880.4 tok/s
vLLM generation done - elapsed time: 5.91s, prompts: 4, total tokens: 4743, avg length: 1185.8, speed: 802.6 tok/s
vLLM generation done - elapsed time: 6.81s, prompts: 4, total tokens: 7059, avg length: 1764.8, speed: 1037.1 tok/s

100%|██████████| 100/100 [25:35<00:00, 15.33s/it]
                                                 
{'loss': 0.0116, 'grad_norm': 0.03328089416027069, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': 0.0116, 'epoch': 0.87}

100%|██████████| 100/100 [25:35<00:00, 15.33s/it]
                                                 
{'train_runtime': 1538.66, 'train_samples_per_second': 2.08, 'train_steps_per_second': 0.065, 'train_loss': 0.011702728271484376, 'epoch': 0.87}

100%|██████████| 100/100 [25:38<00:00, 15.33s/it]
100%|██████████| 100/100 [25:38<00:00, 15.39s/it]
[rank0]:[W802 20:10:26.456099936 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank2]:[W802 20:10:27.901958053 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank3]:[W802 20:10:27.932611673 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
[rank1]:[W802 20:10:27.932836959 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/robust_kl_2048/capped_seed42/checkpoint-100
#373
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/robust_kl_2048/capped_seed42/checkpoint-100 /workspace/eval_results/robust_kl_2048_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":500});\ntext(r.output); if(r.session_id) text(`SESSION_ID=${r.session_id}`);\n"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=79122
#374
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 202
Wall time 11.0 seconds
Output:
#375
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "202",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 13.5 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 16

Warning: truncated output (original token count: 5933)
Total output lines: 154

`torch_dtype` is deprecated! Use `dtype` instead!
Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...

Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards: 100%|██████████| 2/2 [00:00<00:00, 54.15it/s]
Generating with data_parallel_size=4 (TP=1 per engine) ...
INFO 08-02 20:11:07 [__init__.py:216] Automatically detected platform cuda.
INFO 08-02 20:11:07 [__init__.py:216] Automatically detected platform cuda.
INFO 08-0…27 tokens truncated…█████████| 7/7 [00:00<00:00, 284.09it/s]

Processed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s](EngineCore_DP0 pid=15907) INFO 08-02 20:11:23 [__init__.py:381] Cudagraph is disabled under eager mode
INFO 08-02 20:11:23 [llm.py:306] Supported_tasks: ['generate']

Adding requests:   0%|          | 0/7 [00:00<?, ?it/s]
Adding requests: 100%|██████████| 7/7 [00:00<00:00, 297.22it/s]

Processed prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]