Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
428
duration
9h 21m
recorded
2026-08-02

run idopsd-aime24-4gpus__eNDoWrD

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol444 turns16 narration428 reasoning428 action428 observation1 final
#201
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "77",
  "yield_time_ms": 60000,
  "max_tokens": 400
}
wait result
Script completed
Wall time 11.8 seconds
Output:
 time: 3.29s, prompts: 4, total tokens: 3200, avg length: 800.0, speed: 972.8 tok/s
vLLM generation done - elapsed time: 3.41s, prompts: 4, total tokens: 4004, avg length: 1001.0, speed: 1175.5 tok/s

 time: 3.37s, prompts: 4, total tokens: 3686, avg length: 921.5, speed: 1093.3 tok/s
vLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 3658, avg length: 914.5, speed: 1080.1 tok/s

 time: 3.36s, prompts: 4, total tokens: 3701, avg length: 925.2, speed: 1100.0 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3672, avg length: 918.0, speed: 1092.3 tok/s

 time: 3.36s, prompts: 4, total tokens: 3891, avg length: 972.8, speed: 1156.4 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3530, avg length: 882.5, speed: 1046.2 tok/s

time: 3.36s, prompts: 4, total tokens: 3370, avg length: 842.5, speed: 1003.2 tok/s
vLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 4064, avg length: 1016.0, speed: 1196.8 tok/s

time: 3.37s, prompts: 4, total tokens: 3897, avg length: 974.2, speed: 1157.1 tok/s
vLLM generation done - elapsed time: 3.41s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1200.3 tok/s

#202
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"rg \\\"\\\\{'loss'\\\" /workspace/runs/anchor02b/train_seed42.log | tail -5; tail -3 /workspace/runs/anchor02b/train_seed42.log; cat /workspace/.timer/remaining_secs\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:

 30%|███       | 30/100 [04:16<10:00,  8.57s/it]
                                                
{'loss': 0.0104, 'grad_norm': 0.050751786679029465, 'learning_rate': 3.5500000000000003e-06, 'on_policy_loss': 0.0104, 'opsd_kl': -0.0022087767720222473, 'anchor_ce': 0.5166015625, 'epoch': 0.03}

 32%|███▏      | 32/100 [04:33<09:43,  8.59s/it]
                                                
{'loss': 0.0073, 'grad_norm': 0.04762876406311989, 'learning_rate': 3.45e-06, 'on_policy_loss': 0.0073, 'opsd_kl': -0.00421142578125, 'anchor_ce': 0.509765625, 'epoch': 0.03}

 34%|███▍      | 34/100 [04:50<09:28,  8.61s/it]
                                                
{'loss': 0.0087, 'grad_norm': 0.05296391621232033, 'learning_rate': 3.3500000000000005e-06, 'on_policy_loss': 0.0087, 'opsd_kl': -0.0021543502807617188, 'anchor_ce': 0.57177734375, 'epoch': 0.04}

 36%|███▌      | 36/100 [05:07<09:07,  8.55s/it]
                                                
{'loss': 0.0082, 'grad_norm': 0.04530049115419388, 'learning_rate': 3.2500000000000002e-06, 'on_policy_loss': 0.0082, 'opsd_kl': -0.0029659271240234375, 'anchor_ce': 0.51025390625, 'epoch': 0.04}

 38%|███▊      | 38/100 [05:24<08:49,  8.53s/it]
                                                
{'loss': 0.0077, 'grad_norm': 0.0507129468023777, 'learning_rate': 3.1500000000000003e-06, 'on_policy_loss': 0.0077, 'opsd_kl': -0.001634359359741211, 'anchor_ce': 0.51953125, 'epoch': 0.04}
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3515, avg length: 878.8, speed: 1054.9 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3434, avg length: 858.5, speed: 1022.0 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3544, avg length: 886.0, speed: 1052.0 tok/s
14339
#203
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"rg \\\"\\\\{'loss'\\\" /workspace/runs/anchor02b/train_seed42.log | head -5\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":2000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:

  2%|▏         | 2/100 [00:17<14:16,  8.74s/it]
                                               
{'loss': 0.0279, 'grad_norm': 0.21670837700366974, 'learning_rate': 4.95e-06, 'on_policy_loss': 0.0279, 'opsd_kl': 0.0112457275390625, 'anchor_ce': 0.91015625, 'epoch': 0.0}

  4%|▍         | 4/100 [00:34<13:44,  8.59s/it]
                                               
{'loss': 0.0274, 'grad_norm': 0.1821713000535965, 'learning_rate': 4.85e-06, 'on_policy_loss': 0.0274, 'opsd_kl': 0.01204681396484375, 'anchor_ce': 0.8994140625, 'epoch': 0.0}

  6%|▌         | 6/100 [00:51<13:22,  8.54s/it]
                                               
{'loss': 0.0265, 'grad_norm': 0.1785186231136322, 'learning_rate': 4.75e-06, 'on_policy_loss': 0.0265, 'opsd_kl': 0.0104217529296875, 'anchor_ce': 0.83984375, 'epoch': 0.01}

  8%|▊         | 8/100 [01:08<12:59,  8.47s/it]
                                               
{'loss': 0.022, 'grad_norm': 0.16259250044822693, 'learning_rate': 4.65e-06, 'on_policy_loss': 0.022, 'opsd_kl': 0.0075225830078125, 'anchor_ce': 0.8642578125, 'epoch': 0.01}

 10%|█         | 10/100 [01:25<12:55,  8.61s/it]
                                                
{'loss': 0.0214, 'grad_norm': 0.1433873325586319, 'learning_rate': 4.5500000000000005e-06, 'on_policy_loss': 0.0214, 'opsd_kl': 0.00647735595703125, 'anchor_ce': 0.6943359375, 'epoch': 0.01}
#204
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<9;i++){const r=await tools.write_stdin({session_id:50612,chars:\"\",yield_time_ms:60000,max_output_tokens:60});out+=r.output.slice(-120)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 80
Wall time 11.0 seconds
Output:
#205
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#206
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#207
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#208
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#209
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#210
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#211
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 80
Wall time 61.0 seconds
Output:
#212
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "80",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 44.6 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 21

truncated… total tokens: 4042, avg length: 1010.5, speed: 1189.2 tok/s

 47%|████▋     | 47/100 [06:41<07:34,  8.57s/it]
s truncated…4817962646484375, 'anchor_ce': 0.603515625, 'epoch': …208 tokens truncated….1}

 96%|█████████▌| 96/100 [13:43<00:34,  8.64s/it]
table/distributed.html#shutdown (function operator())
TRAIN_CKPT=/workspace/runs/anchor02b/capped_seed42/checkpoint-100

EXIT0
#213
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/anchor02b/capped_seed42/checkpoint-100\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":500});\ntext(r.output); if(r.session_id) text(`SESSION_ID=${r.session_id}`);"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
SESSION_ID=4178
#214
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "let out=\"\";for(let i=0;i<12;i++){const r=await tools.write_stdin({session_id:4178,chars:\"\",yield_time_ms:60000,max_output_tokens:50});out+=r.output.slice(-100)+\"\\n\";if(r.exit_code!==undefined){out+=`EXIT${r.exit_code}`;break;}}text(out);\n"
}
exec result
Script running with cell ID 82
Wall time 11.0 seconds
Output:
#215
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#216
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#217
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#218
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#219
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#220
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#221
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#222
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#223
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#224
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script running with cell ID 82
Wall time 61.0 seconds
Output:
#225
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "82",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 43.5 seconds
Output:
Warning: truncated output (original token count: 242)
Total output lines: 12

 prompts:   0%|          | 0/84 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]


2%|█▎        | 12/96 [02:37<18:22, 13.12s/it, est. speed input: 11.89 toks/s, output: 520.86 …142 tokens truncated… | 72/84 [10:09<01:00,  5.08s/it, est. speed input: 12.49 toks/s, output: 1962.52 toks/s]
uncated…██▊ | 84/96 [11:11<01:11,  5.98s/it, est. speed input: 19.57 toks/s, output: 1812.79 toks/s]