gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 627
- duration
- 6h 30m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__RqGcDty
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:250});\ntext(r.output);\n"
}exec result
Script running with cell ID 35
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "35",
"yield_time_ms": 30000,
"max_tokens": 250
}wait result
Script completed
Wall time 10.9 seconds
Output:
Warning: truncated output (original token count: 277)
Total output lines: 11
Warning: truncated output (original token count: 1458)
Total output lines: 50
vLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 2787, avg length: 696.8, speed: 846.1 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3476, avg length: 869.0, speed: 1042.5 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3289, avg length: 822.2, speed: 982.3 tok/s
vLLM generation done - elapsed time: 3.41s, prompts: 4, total tokens: 4096, av…27 tokens truncated…tion done - elapsed time: 3.35s, prompts: 4, total tokens: 3229, avg length: 807.2, speed: 964.2 tok/s
vLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 3927, avg length: 981.8, speed: 1162.4 tok/s
40%|████ | 40/100 [05:27<08:09, 8.17s/it]
{'loss': -0.0051, 'grad_norm': 0.049328841269016266, 'learning_rate': 5e-06, 'on_policy_loss': -0.0051, 'epoch': 0.3}
40%|████ | 40/100 [05:27<08:09, 8.17s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:200});\ntext(r.output);\n"
}exec result
Script running with cell ID 36
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "36",
"yield_time_ms": 30000,
"max_tokens": 200
}wait result
Script completed
Wall time 15.6 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10
Warning: truncated output (original token count: 1408)
Total output lines: 49
vLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3287, avg length: 821.8, speed: 996.7 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1214.5 tok/s
vLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096, avg length: 102…27 tokens truncated…generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3832, avg length: 958.0, speed: 1140.5 tok/s
vLLM generation done - elapsed time: 3.38s, prompts: 4, total tokens: 3972, avg length: 993.0, speed: 1174.3 tok/s
vLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 4085, avg length: 1021.2, speed: 1199.7 tok/s
45%|████▌ | 45/100 [06:07<07:30, 8.19s/it]The run is behaving like the released estimator: its reported objective crossed below zero around step 20 even though it is labeled forward KL. I’m treating that as an estimator artifact rather than a convergence signal. The first evaluation will tell us whether target-domain replay benefits the existing dynamics; the next candidate will address rollout coverage and the malformed clipping separately so we do not conflate causes.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:200});\ntext(r.output);\n"
}exec result
Script running with cell ID 37
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "37",
"yield_time_ms": 30000,
"max_tokens": 200
}wait result
Script completed
Wall time 11.7 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10
Warning: truncated output (original token count: 1278)
Total output lines: 44
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3226, avg length: 806.5, speed: 968.5 tok/s
vLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3536, avg length: 884.0, speed: 1052.1 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024…27 tokens truncated…al tokens: 3041, avg length: 760.2, speed: 921.8 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3366, avg length: 841.5, speed: 1011.1 tok/s
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3468, avg length: 867.0, speed: 1039.3 tok/s
vLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 3787, avg length: 946.8, speed: 1116.1 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"tail -c 1600 /workspace/runs/domain_only/train_seed42.log | rg -o \\\"[0-9]+%\\\\|[^\\\\r]*|\\\\{'loss'.*\\\" | tail -3\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":500});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
50%|█████ | 50/100 [06:48<06:47, 8.15s/it]
{'loss': -0.0078, 'grad_norm': 0.059190534055233, 'learning_rate': 5e-06, 'on_policy_loss': -0.0078, 'epoch': 0.38}
50%|█████ | 50/100 [06:48<06:47, 8.15s/it]vLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 2472, avg length: 618.0, speed: 776.0 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"git diff > /tmp/domain_replay.patch; git status --short; cat /workspace/.timer/remaining_secs\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":1000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
M README.md
M opsd_train.py
20829
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"git diff -- README.md | head -60\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":2000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
diff --git a/README.md b/README.md
index d51e96b..e76afdd 100644
--- a/README.md
+++ b/README.md
@@ -18,8 +18,6 @@
(1) Fixed chat template and zero2 bugs (see [template issue](https://github.com/huggingface/trl/issues/5241)), we re-ran experiments with updated results (detailed results & ablations updated on arxiv/blog). The fixes yield improved OPSD performance, most notably on Qwen3-1.7B.
- (2) Added a new training stabilization strategy 🚀: per-token point-wise KL clipping. We find style tokens (such as 'wait', 'think') can exhibit 6–15× higher KL divergence than math-related tokens, and dominates the training signal. Clipping stablizes training and improves performance.
-
- **Mar 3, 2026**: Initial code release.
@@ -284,7 +282,7 @@ bash run_eval_nonthink.sh
| `--use_tinker_loss` | `False` | Use sampled-token policy-gradient objective instead of full-vocabulary JSD. More memory efficient. Currently no clipped implemented for this variant, could be unstable. |
| `--max_completion_length` | — | Student generation length for distillation. We use 1024 in our main experiments. |
| `--beta` | — | Interpolation weight for the JSD mixture distribution. Beta=0 means forward KL and 1 means reverse KL. |
-| `--jsd_token_clip` | 0.05 | Clip the JSD loss for each token to a maximum value. This can improve stability by preventing stylistic tokens from dominating the training signal. Note when clipping is applied, the loss can be negative due to positive KL summand being capped. |
+| `--jsd_token_clip` | 0.05 | Clip the JSD loss for each token to a maximum value. This can improve stability. |
| `--reason_first` | `False` | Prepend an explicit rationalization to the teacher context before distillation. |
| `--run_config` | `None` | Custom name suffix for the output directory and WandB run. |
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"tail -c 1000 /workspace/runs/domain_only/train_seed42.log | rg -o \\\"[0-9]+%\\\\|[^\\\\r]*|\\\\{'loss'.*\\\" | tail -3\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":500});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
54%|█████▍ | 54/100 [07:21<06:16, 8.19s/it]
{'loss': -0.0084, 'grad_norm': 0.05572548136115074, 'learning_rate': 5e-06, 'on_policy_loss': -0.0084, 'epoch': 0.41}
54%|█████▍ | 54/100 [07:21<06:16, 8.19s/it]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:150});\ntext(r.output);\n"
}exec result
Script running with cell ID 42
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "42",
"yield_time_ms": 30000,
"max_tokens": 150
}wait result
Script completed
Wall time 9.9 seconds
Output:
Warning: truncated output (original token count: 177)
Total output lines: 9
Warning: truncated output (original token count: 3788)
Total output lines: 236
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3483, avg length: 870.8, speed: 1047.1 tok/s
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 2999, avg length: 749.8, speed: 901…27 tokens truncated…1 tok/s
vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3472, avg length: 868.0, speed: 1037.6 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4014, avg length: 1003.5, speed: 1192.5 tok/s
59%|█████▉ | 59/100 [08:02<05:34, 8.15s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:120});\ntext(r.output);\n"
}exec result
Script running with cell ID 43
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "43",
"yield_time_ms": 30000,
"max_tokens": 120
}wait result
Script completed
Wall time 11.8 seconds
Output:
Warning: truncated output (original token count: 147)
Total output lines: 8
Warning: truncated output (original token count: 1719)
Total output lines: 58
vLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3552, avg length: 888.0, speed: 1067.7 tok/s
vLLM generation done - elapsed time: 3.34s, pr…27 tokens truncated…4, total tokens: 3603, avg length: 900.8, speed: 1072.4 tok/s
vLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3588, avg length: 897.0, speed: 1066.0 tok/s
65%|██████▌ | 65/100 [08:50<04:44, 8.12s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv,noheader\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":1000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.2 seconds
Output:
104025 MiB, 143771 MiB, 79 %
104205 MiB, 143771 MiB, 79 %
104063 MiB, 143771 MiB, 80 %
104061 MiB, 143771 MiB, 80 %
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"rg -n \\\"Please reason|Problem:|apply_chat_template|enable_thinking\\\" /workspace/eval_official.py | head -80\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
61: enable_thinking=True,
325: f"{ex['problem']}\n\nPlease reason step by step, and put your final answer within \\boxed{{}}."
327: text = tokenizer.apply_chat_template(
331: enable_thinking=FROZEN["enable_thinking"],
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:100});\ntext(r.output);\n"
}exec result
Script running with cell ID 46
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "46",
"yield_time_ms": 30000,
"max_tokens": 100
}wait result
Script completed
Wall time 11.4 seconds
Output:
Warning: truncated output (original token count: 127)
Total output lines: 6
Warning: truncated output (original token count: 2093)
Total output lines: 73
vLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3092, avg length: 773.0, speed: 930.0 tok/s
vLLM ge…27 tokens truncated…
{'loss': -0.0128, 'grad_norm': 0.0539618618786335, 'learning_rate': 5e-06, 'on_policy_loss': -0.0128, 'epoch': 0.55}
72%|███████▏ | 72/100 [09:48<03:48, 8.15s/it]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 47
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "47",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 10.1 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1288)
Total output lines: 44
vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3341, avg l…27 tokens truncated…96, avg length: 1024.0, speed: 1206.5 tok/s
vLLM generation done - elapsed time: 3.46s, prompts: 4, total tokens: 4057, avg length: 1014.2, speed: 1172.7 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"tail -c 700 /workspace/runs/domain_only/train_seed42.log | rg -o \\\"[0-9]+%\\\\|[^\\\\r]*|\\\\{'loss'.*\\\" | tail -2\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":300});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
77%|███████▋ | 77/100 [10:29<03:10, 8.27s/it]vLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3593, avg length: 898.2, speed: 1072.9 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:71812,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}exec result
Script running with cell ID 49
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "49",
"yield_time_ms": 30000,
"max_tokens": 80
}wait result
Script completed
Wall time 16.0 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6
Warning: truncated output (original token count: 1726)
Total output lines: 57
vLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 2837, avg l…27 tokens truncated…759, avg length: 939.8, speed: 1119.3 tok/s
vLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1205.7 tok/s
A second audit found a separate alignment issue: training prepends Problem: to the student message, while the frozen evaluator does not. More importantly, only about 48% of 1,024-token training rollouts reach a boxed answer. I’ll test an evaluation-aligned student prompt and a modestly longer rollout next; these directly increase useful teacher supervision without changing optimizer steps or global batch.