Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#176
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 105
Wall time 11.0 seconds
Output:
#177
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "105",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.3 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1039)
Total output lines: 36

vLLM generation done - elapsed time: 4.99s, prompts: 4, total tokens: 4918, avg l…27 tokens truncated…44, avg length: 1536.0, speed: 1202.6 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 6001, avg length: 1500.2, speed: 1167.3 tok/s
#178
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"nvidia-smi --query-gpu=memory.used,memory.total --format=csv,noheader; tail -c 1000 /workspace/runs/refined_long/train_seed42.log | rg -o \\\"[0-9]+%\\\\|[^\\\\r]*|\\\\{'loss'.*\\\" | tail -3\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":1000});\ntext(r.output);\n"
}
exec result
Script completed
Wall time 0.2 seconds
Output:
109425 MiB, 143771 MiB
110993 MiB, 143771 MiB
110179 MiB, 143771 MiB
109637 MiB, 143771 MiB
32%|███▏      | 32/100 [06:19<13:35, 11.99s/it]vLLM generation done - elapsed time: 5.00s, prompts: 4, total tokens: 4384, avg length: 1096.0, speed: 877.6 tok/s
#179
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 107
Wall time 11.0 seconds
Output:
#180
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "107",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 15.2 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1058)
Total output lines: 37

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5376, avg l…27 tokens truncated….04333556815981865, 'learning_rate': 3.2500000000000002e-06, 'on_policy_loss': -0.0039, 'epoch': 0.31}

 36%|███▌      | 36/100 [07:06<12:43, 11.93s/it]
#181
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 108
Wall time 11.0 seconds
Output:
#182
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "108",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 783)
Total output lines: 25

vLLM generation done - elapsed time: 4.84s, prompts: 4, total tokens: 4211, avg le…26 tokens truncated…41, avg length: 1410.2, speed: 1109.0 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1196.1 tok/s
#183
narration

The 1,536-token configuration has substantial memory headroom (about 110/144 GiB per GPU), so rollout extension is not trading against batch size or stability. I’ll use the measured full score—not memory limits—to decide whether 1,536 or 2,048 tokens belongs in the final recipe.

#184
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 109
Wall time 11.0 seconds
Output:
#185
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "109",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 5

Warning: truncated output (original token count: 936)
Total output lines: 33


 39%|███▉      | 39/100 [07:42<12:04, 11.88s/it]vLLM generation done - el…26 tokens truncated…rad_norm': 0.05010228231549263, 'learning_rate': 2.95e-06, 'on_policy_loss': -0.0044, 'epoch': 0.37}

 42%|████▏     | 42/100 [08:18<11:28, 11.88s/it]
#186
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 110
Wall time 11.0 seconds
Output:
#187
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "110",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.3 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 911)
Total output lines: 29

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5468, avg le…26 tokens truncated…13, avg length: 1528.2, speed: 1196.5 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1192.5 tok/s
#188
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 111
Wall time 11.0 seconds
Output:
#189
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "111",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1061)
Total output lines: 37

vLLM generation done - elapsed time: 5.03s, prompts: 4, total tokens: 5371, avg l…27 tokens truncated… done - elapsed time: 5.16s, prompts: 4, total tokens: 6007, avg length: 1501.8, speed: 1164.0 tok/s

 49%|████▉     | 49/100 [09:41<10:11, 11.98s/it]
#190
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 112
Wall time 11.0 seconds
Output:
#191
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "112",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.0 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 928)
Total output lines: 33

vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5024, avg le…26 tokens truncated…75057831406593, 'learning_rate': 2.4500000000000003e-06, 'on_policy_loss': -0.0067, 'epoch': 0.45}

 52%|█████▏    | 52/100 [10:17<09:37, 12.02s/it]
#192
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 113
Wall time 11.0 seconds
Output:
#193
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "113",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 5.3 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2816)
Total output lines: 200

vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5243, avg …27 tokens truncated…19, avg length: 1404.8, speed: 1098.1 tok/s
vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 5948, avg length: 1487.0, speed: 1161.8 tok/s
#194
narration

Mid-run dynamics remain close to the paper’s stable linear-decay trajectory: at step 52 the clipped objective is −0.0067, compared with roughly −0.0057 in the supplied baseline log. That makes the upcoming score a fair test of data/coverage changes rather than a plainly unstable optimization run.

#195
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 114
Wall time 11.0 seconds
Output:
#196
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 14.4 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1058)
Total output lines: 37

vLLM generation done - elapsed time: 5.05s, prompts: 4, total tokens: 4846, avg l…27 tokens truncated…one - elapsed time: 5.16s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1191.2 tok/s

 59%|█████▉    | 59/100 [11:40<08:09, 11.94s/it]
#197
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 115
Wall time 11.0 seconds
Output:
#198
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "115",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 930)
Total output lines: 33

vLLM generation done - elapsed time: 4.78s, prompts: 4, total tokens: 3506, avg le…26 tokens truncated…328382372856, 'learning_rate': 1.9500000000000004e-06, 'on_policy_loss': -0.0082, 'epoch': 0.54}

 62%|██████▏   | 62/100 [12:16<07:34, 11.95s/it]
#199
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:61078,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 116
Wall time 11.0 seconds
Output:
#200
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "116",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 917)
Total output lines: 29

vLLM generation done - elapsed time: 5.07s, prompts: 4, total tokens: 5421, avg le…26 tokens truncated…84, avg length: 1371.0, speed: 1076.7 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 5741, avg length: 1435.2, speed: 1123.4 tok/s