gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 596
- duration
- 7h 59m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__kuc95M7
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "19",
"yield_time_ms": 30000,
"max_tokens": 2000
}wait result
Script completed
Wall time 16.4 seconds
Output:
Warning: truncated output (original token count: 2078)
Total output lines: 1
{"chunk_id":"013f03","wall_time_seconds":30.001243242,"session_id":80127,"original_token_count":6903,"output":"Warning: truncated output (original token count: 6903)\nTotal output lines: 1\n\n\rTokenizing train dataset: 8%|▊ | 3638/43744 [00:21<03:40, 182.16 examples/s]\rTokenizing train dataset: 8%|▊ | 3669/43744 [00:22<03:34, 186.72 examples/s]\rTokenizing train dataset: 8%|▊ | 3694/43744 [00:22<03:47, 176.00 examples/s]\rTokenizing train dataset: 9%|▊ | 3720/43744 [00:22<03:56, 169.48 examples/s]\rTokenizing train dataset: 9%|▊ | 3740/43744 [00:22<03:53, 171.06 examples/s]\rTokenizing train dataset: 9%|▊ | 3768/43744 [00:22<03:50, 173.37 examples/s]\rTokenizing train dataset: 9%|▊ | 3786/43744 [00:22<03:52, 171.81 examples/s]\rTokenizing train dataset: 9%|▊ | 3805/43744 [00:22<03:50, 173.28 examples/s]\rTokenizing train dataset: 9%|▊ | 3823/43744 [00:22<03:54, 170.52 examples/s]\rTokenizing train dataset: 9%|▉ | 3841/43744 [00:23<03:59, 166.61 examples/s]\rTokenizing train dataset: 9%|▉ | 3859/43744 [00:23<03:56, 168.96 examples/s]\rTokenizing train dataset: 9%|▉ | 3877/43744 [00:23<03:56, 168.81 examples/s]\rTokenizing train dataset: 9%|▉ | 3897/43744 [00:23<03:47, 175.49 examples/s]\rTokenizing train dataset: 9%|▉ | 3917/43744 [00:23<03:42, 178.85 examples/s]\rTokenizing train dataset: 9%|▉ | 3937/43744 [00:23<03:39, 180.96 examples/s]\rTokenizing train dataset: 9%|▉ | 3965/43744 [00:23<03:44, 176.85 examples/s]\rTokenizing train dataset: 9%|▉ | 3983/43744 [00:23<03:49, 173.01 examples/s]\rTokenizing train dataset: 9%|▉ | 4008/43744 [00:24<07:25, 89.16 examples/s] \rTokenizing train dataset: 9%|▉ | 4023/43744 [00:24<06:50, 96.86 examples/s]\rTokenizing train dataset: 9%|▉ | 4043/43744 [00:24<05:50, 113.24 examples/s]\rTokenizing train dataset: 9%|▉ | 4061/43744 [00:24<05:17, 125.03 examples/s]\rTokenizing train dataset: 9%|▉ | 4081/43744 [00:24<04:45, 138.91 examples/s]\rTokenizing train dataset: 9%|▉ | 4099/43744 [00:24<04:33, 144.97 examples/s]\rTokenizing train dataset: 9%|▉ | 4122/43744 [00:25<04:05, 161.16 examples/s]\rTokenizing train dataset: 9%|▉ | 4141/43744 [00:25<04:01, 164.03 examples/s]\rTokenizing train dataset: 10%|▉ | 4159/43744 [00:25<04:00, 164.73 examples/s]\rTokenizing train dataset: 10%|▉ | 4179/43744 [00:25<03:48, 172.79 examples/s]\rTokenizing train dataset: 10%|▉ | 4197/43744 [00:25<03:47, 173.56 examples/s]\rTokenizing train dataset: 10%|▉ | 4217/43744 [00:25<03:41, 178.13 examples/s]\rTokenizing train dataset: 10%|▉ | 4236/43744 [00:25<03:43, 176.94 examples/s]\rTokenizing train dataset: 10%|▉ | 4258/43744 [00:25<03:29, 188.24 examples/s]\rTokenizing train dataset: 10%|▉ | 4282/43744 [00:25<03:19, 198.28 examples/s]\rTokenizing train dataset: 10%|▉ | 4303/43744 [00:26<03:16, 200.36 examples/s]\rTokenizing train dataset: 10%|▉ | 4333/43744 [00:26<03:19, 197.29 examples/s]\rTokenizing train dataset: 10%|▉ | 4362/43744 [00:26<03:26, 191.07 examples/s]\rTokenizing train dataset: 10%|█ | 4392/43744 [00:26<03:25, 191.08 examples/s]\rTokenizing train dataset: 10%|█ | 4412/43744 [00:26<03:24, 192.28 examples/s]\rTokenizing train dataset: 10%|█ | 4440/43744 [00:26<03:28, 188.88 examples/s]\rTokenizing train dataset: 10%|█ | 4468/43744 [00:26<03:31, 185.30 examples/s]\rTokenizing train dataset: 10%|█ | 4493/43744 [00:27<03:43, 175.55 examples/s]\rTokenizing train dataset: 10%|█ | 4515/43744 [00:27<03:34, 183.18 examples/s]\rTokenizing train dataset: 10%|█ | 4537/43744 [00:27<03:50, 169.78 examples/s]\rTokenizing train dataset: …78 tokens truncated… 21%|██ | 9109/43744 [00:56<03:54, 147.59 examples/s]\rTokenizing train dataset: 21%|██ | 9129/43744 [00:56<03:39, 157.89 examples/s]\rTokenizing train dataset: 21%|██ | 9153/43744 [00:56<03:40, 156.94 examples/s]\rTokenizing train dataset: 21%|██ | 9174/43744 [00:56<03:25, 168.25 examples/s]\rTokenizing train dataset: 21%|██ | 9199/43744 [00:57<03:29, 165.14 examples/s]\rTokenizing train dataset: 21%|██ | 9218/43744 [00:57<03:23, 169.93 examples/s]\rTokenizing train dataset: 21%|██ | 9236/43744 [00:57<03:23, 169.89 examples/s]\rTokenizing train dataset: 21%|██ | 9264/43744 [00:57<03:19, 172.69 examples/s]\rTokenizing train dataset: 21%|██ | 9282/43744 [00:57<03:18, 173.66 examples/s]\rTokenizing train dataset: 21%|██▏ | 9307/43744 [00:57<03:24, 168.57 examples/s]\rTokenizing train dataset: 21%|██▏ | 9330/43744 [00:57<03:11, 179.82 examples/s]\rTokenizing train dataset: 21%|██▏ | 9358/43744 [00:57<03:15, 175.50 examples/s]\rTokenizing train dataset: 21%|██▏ | 9379/43744 [00:58<03:10, 180.82 examples/s]\rTokenizing train dataset: 21%|██▏ | 9398/43744 [00:58<03:10, 180.19 examples/s]\rTokenizing train dataset: 22%|██▏ | 9418/43744 [00:58<03:08, 181.77 examples/s]\rTokenizing train dataset: 22%|██▏ | 9438/43744 [00:58<03:06, 183.92 examples/s]\rTokenizing train dataset: 22%|██▏ | 9457/43744 [00:58<03:10, 180.45 examples/s]\rTokenizing train dataset: 22%|██▏ | 9481/43744 [00:58<03:19, 171.45 examples/s]\rTokenizing train dataset: 22%|██▏ | 9500/43744 [00:58<03:16, 174.64 examples/s]\rTokenizing train dataset: 22%|██▏ | 9520/43744 [00:58<03:09, 180.45 examples/s]\rTokenizing train dataset: 22%|██▏ | 9544/43744 [00:58<03:22, 169.13 examples/s]\rTokenizing train dataset: 22%|██▏ | 9571/43744 [00:59<03:23, 168.11 examples/s]\rTokenizing train dataset: 22%|██▏ | 9592/43744 [00:59<03:12, 177.23 examples/s]\rTokenizing train dataset: 22%|██▏ | 9619/43744 [00:59<03:15, 174.74 examples/s]\rTokenizing train dataset: 22%|██▏ | 9640/43744 [00:59<03:08, 180.63 examples/s]\rTokenizing train dataset: 22%|██▏ | 9666/43744 [00:59<03:15, 174.23 examples/s]\rTokenizing train dataset: 22%|██▏ | 9684/43744 [00:59<03:17, 172.33 examples/s]\rTokenizing train dataset: 22%|██▏ | 9702/43744 [00:59<03:18, 171.37 examples/s]\rTokenizing train dataset: 22%|██▏ | 9721/43744 [00:59<03:14, 175.17 examples/s]\rTokenizing train dataset: 22%|██▏ | 9748/43744 [01:00<03:17, 172.27 examples/s]\rTokenizing train dataset: 22%|██▏ | 9767/43744 [01:00<03:15, 173.64 examples/s]\rTokenizing train dataset: 22%|██▏ | 9788/43744 [01:00<03:07, 181.49 examples/s]\rTokenizing train dataset: 22%|██▏ | 9810/43744 [01:00<02:59, 189.50 examples/s]\rTokenizing train dataset: 22%|██▏ | 9839/43744 [01:00<03:04, 183.97 examples/s]\rTokenizing train dataset: 23%|██▎ | 9858/43744 [01:00<03:06, 181.81 examples/s]\rTokenizing train dataset: 23%|██▎ | 9877/43744 [01:00<03:05, 182.11 examples/s]\rTokenizing train dataset: 23%|██▎ | 9896/43744 [01:00<03:05, 182.49 examples/s]\rTokenizing train dataset: 23%|██▎ | 9916/43744 [01:01<03:07, 180.69 examples/s]\rTokenizing train dataset: 23%|██▎ | 9935/43744 [01:01<03:05, 182.35 examples/s]\rTokenizing train dataset: 23%|██▎ | 9958/43744 [01:01<02:54, 193.95 examples/s]\rTokenizing train dataset: 23%|██▎ | 9981/43744 [01:01<02:49, 199.66 examples/s]\rTokenizing train dataset: 23%|██▎ | 10008/43744 [01:01<05:35, 100.54 examples/s]\rTokenizing train dataset: 23%|██▎ | 10028/43744 [01:01<04:52, 115.43 examples/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:1000});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 20
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "20",
"yield_time_ms": 60000,
"max_tokens": 1000
}wait result
Script completed
Wall time 41.3 seconds
Output:
Warning: truncated output (original token count: 1067)
Total output lines: 1
{"chunk_id":"887951","wall_time_seconds":60.001833265,"session_id":80127,"original_token_count":12387,"output":"Warning: truncated output (original token count: 12387)\nTotal output lines: 1\n\n\rTokenizing train dataset: 23%|██▎ | 10046/43744 [01:02<04:27, 125.75 examples/s]\rTokenizing train dataset: 23%|██▎ | 10067/43744 [01:02<03:58, 141.48 examples/s]\rTokenizing train dataset: 23%|██▎ | 10091/43744 [01:02<03:54, 143.43 examples/s]\rTokenizing train dataset: 23%|██▎ | 10113/43744 [01:02<03:34, 156.64 examples/s]\rTokenizing train dataset: 23%|██▎ | 10132/43744 [01:02<03:28, 161.04 examples/s]\rTokenizing train dataset: 23%|██▎ | 10151/43744 [01:02<03:21, 166.31 examples/s]\rTokenizing train dataset: 23%|██▎ | 10169/43744 [01:02<03:20, 167.06 examples/s]\rTokenizing train dataset: 23%|██▎ | 10188/43744 [01:02<03:16, 171.01 examples/s]\rTokenizing train dataset: 23%|██▎ | 10208/43744 [01:02<03:11, 175.21 examples/s]\rTokenizing train dataset: 23%|██▎ | 10226/43744 [01:03<03:13, 173.01 examples/s]\rTokenizing train dataset: 23%|██▎ | 10246/43744 [01:03<03:09, 176.34 examples/s]\rTokenizing train dataset: 23%|██▎ | 10265/43744 [01:03<03:11, 174.94 examples/s]\rTokenizing train dataset: 24%|██▎ | 10290/43744 [01:03<03:22, 164.87 examples/s]\rTokenizing train dataset: 24%|██▎ | 10311/43744 [01:03<03:13, 173.15 examples/s]\rTokenizing train dataset: 24%|██▎ | 10332/43744 [01:03<03:05, 180.44 examples/s]\rTokenizing train dataset: 24%|██▎ | 10357/43744 [01:03<03:14, 172.06 examples/s]\rTokenizing train dataset: 24%|██▎ | 10379/43744 [01:03<03:05, 179.73 examples/s]\rTokenizing train dataset: 24%|██▍ | 10406/43744 [01:04<03:12, 173.55 examples/s]\rTokenizing train dataset: 24%|██▍ | 10424/43744 [01:04<03:13, 172.64 examples/s]\…67 tokens truncated…164.99 examples/s]\rTokenizing train dataset: 47%|████▋ | 20534/43744 [02:06<02:30, 154.22 examples/s]\rTokenizing train dataset: 47%|████▋ | 20552/43744 [02:07<02:29, 154.83 examples/s]\rTokenizing train dataset: 47%|████▋ | 20568/43744 [02:07<02:30, 154.17 examples/s]\rTokenizing train dataset: 47%|████▋ | 20588/43744 [02:07<02:21, 163.30 examples/s]\rTokenizing train dataset: 47%|████▋ | 20605/43744 [02:07<02:21, 163.03 examples/s]\rTokenizing train dataset: 47%|████▋ | 20624/43744 [02:07<02:20, 164.16 examples/s]\rTokenizing train dataset: 47%|████▋ | 20644/43744 [02:07<02:13, 172.43 examples/s]\rTokenizing train dataset: 47%|████▋ | 20663/43744 [02:07<02:14, 171.80 examples/s]\rTokenizing train dataset: 47%|████▋ | 20686/43744 [02:07<02:22, 161.80 examples/s]\rTokenizing train dataset: 47%|████▋ | 20705/43744 [02:07<02:19, 165.41 examples/s]\rTokenizing train dataset: 47%|████▋ | 20722/43744 [02:08<02:19, 165.46 examples/s]\rTokenizing train dataset: 47%|████▋ | 20741/43744 [02:08<02:17, 167.21 examples/s]\rTokenizing train dataset: 47%|████▋ | 20762/43744 [02:08<02:10, 176.11 examples/s]\rTokenizing train dataset: 48%|████▊ | 20782/43744 [02:08<02:07, 180.79 examples/s]\rTokenizing train dataset: 48%|████▊ | 20802/43744 [02:08<02:04, 183.96 examples/s]\rTokenizing train dataset: 48%|████▊ | 20827/43744 [02:08<02:11, 174.79 examples/s]\rTokenizing train dataset: 48%|████▊ | 20845/43744 [02:08<02:12, 173.26 examples/s]\rTokenizing train dataset: 48%|████▊ | 20866/43744 [02:08<02:09, 176.75 examples/s]\rTokenizing train dataset: 48%|████▊ | 20884/43744 [02:08<02:09, 176.06 examples/s]\rTokenizing train dataset: 48%|████▊ | 20912/43744 [02:09<02:09, 175.99 examples/s]"}The replay mix is being built correctly (477 in-dataset AMC/AIME rows, about 33% effective sampling mass). Preprocessing is slower because the duplicated dataset is tokenized before training; once this run finishes I’ll replace physical duplication with a compact deterministic sampler if we keep the idea, so official execution stays lean.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:800});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 21
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "21",
"yield_time_ms": 60000,
"max_tokens": 800
}wait result
Script completed
Wall time 46.2 seconds
Output:
Warning: truncated output (original token count: 864)
Total output lines: 1
{"chunk_id":"95ae96","wall_time_seconds":60.000583453,"session_id":80127,"original_token_count":13365,"output":"Warning: truncated output (original token count: 13365)\nTotal output lines: 1\n\n\rTokenizing train dataset: 48%|████▊ | 20930/43744 [02:09<02:10, 174.96 examples/s]\rTokenizing train dataset: 48%|████▊ | 20953/43744 [02:09<02:01, 187.13 examples/s]\rTokenizing train dataset: 48%|████▊ | 20976/43744 [02:09<01:58, 192.75 examples/s]\rTokenizing train dataset: 48%|████▊ | 20996/43744 [02:09<01:58, 191.56 examples/s]\rTokenizing train dataset: 48%|████▊ | 21024/43744 [02:10<04:01, 94.13 examples/s] \rTokenizing train dataset: 48%|████▊ | 21047/43744 [02:10<03:36, 104.96 examples/s]\rTokenizing train dataset: 48%|████▊ | 21068/43744 [02:10<03:07, 120.81 examples/s]\rTokenizing train dataset: 48%|████▊ | 21088/43744 [02:10<02:49, 133.77 examples/s]\rTokenizing train dataset: 48%|████▊ | 21106/43744 [02:10<02:38, 142.75 examples/s]\rTokenizing train dataset: 48%|████▊ | 21129/43744 [02:10<02:21, 159.42 examples/s]\rTokenizing train dataset: 48%|████▊ | 21157/43744 [02:10<02:17, 164.64 examples/s]\rTokenizing train dataset: 48%|████▊ | 21182/43744 [02:11<02:20, 160.10 examples/s]\rTokenizing train dataset: 48%|████▊ | 21200/43744 [02:11<02:18, 163.01 examples/s]\rTokenizing train dataset: 49%|████▊ | 21222/43744 [02:11<02:09, 174.40 examples/s]\rTokenizing train d…64 tokens truncated…:59, 188.92 examples/s]\rTokenizing train dataset: 74%|███████▍ | 32548/43744 [03:16<00:59, 189.14 examples/s]\rTokenizing train dataset: 74%|███████▍ | 32577/43744 [03:16<01:01, 181.90 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32598/43744 [03:16<00:59, 186.76 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32618/43744 [03:16<00:59, 188.48 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32638/43744 [03:16<00:58, 190.75 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32659/43744 [03:17<00:57, 193.72 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32681/43744 [03:17<00:55, 198.94 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32704/43744 [03:17<00:54, 202.48 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32730/43744 [03:17<00:58, 188.31 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32750/43744 [03:17<00:58, 188.76 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32779/43744 [03:17<00:59, 185.31 examples/s]\rTokenizing train dataset: 75%|███████▍ | 32799/43744 [03:17<00:59, 183.87 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32820/43744 [03:17<00:58, 185.74 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32842/43744 [03:17<00:57, 191.25 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32862/43744 [03:18<00:57, 189.88 examples/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:1500});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 22
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "22",
"yield_time_ms": 60000,
"max_tokens": 1500
}wait result
Script completed
Wall time 43.9 seconds
Output:
Warning: truncated output (original token count: 1571)
Total output lines: 1
{"chunk_id":"26e527","wall_time_seconds":60.000934534,"session_id":80127,"original_token_count":13675,"output":"Warning: truncated output (original token count: 13675)\nTotal output lines: 2\n\n\rTokenizing train dataset: 75%|███████▌ | 32893/43744 [03:18<00:56, 193.45 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32914/43744 [03:18<00:55, 194.16 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32934/43744 [03:18<00:55, 195.28 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32954/43744 [03:18<00:55, 192.74 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32975/43744 [03:18<00:55, 193.58 examples/s]\rTokenizing train dataset: 75%|███████▌ | 32998/43744 [03:18<00:53, 201.17 examples/s]\rTokenizing train dataset: 75%|███████▌ | 33020/43744 [03:19<01:50, 97.24 examples/s] \rTokenizing train dataset: 76%|███████▌ | 33041/43744 [03:19<01:33, 115.02 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33062/43744 [03:19<01:20, 131.94 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33082/43744 [03:19<01:13, 144.34 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33104/43744 [03:19<01:07, 157.70 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33125/43744 [03:19<01:04, 165.01 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33145/43744 [03:19<01:02, 169.86 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33165/43744 [03:20<01:00, 175.58 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33190/43744 [03:20<01:03, 166.80 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33217/43744 [03:20<01:03, 166.78 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33243/43744 [03:20<01:02, 166.90 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33264/43744 [03:20<00:59, 175.95 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33286/43744 [03:20<00:57, 182.20 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33307/43744 [03:20<00:56, 185.37 examples/s]\rTokenizing train dataset: 76%|███████▌ | 33331/43744 [03:20<00:53, 196.01 examples/s]\rTokenizing train dataset: 76%|███████▋ | 33358/43744 [03:21<00:55, 187.60 examples/s]\rTokenizing train dataset: 76%|███████▋ | 33388/43744 [03:21<00:55, 186.85 examples/s]\rTokenizing train dataset: 76%|███████▋ | 33407/43744 [03:21<00:55, 186.14 examples/s]\rTokenizing train dataset: 76%|███████▋ | 33427/43744 [03:21<00:56, 184.20 examples/s]\rTokenizing train dataset: 76%|███████▋ | 33452/43744 [03:21<00:59, 172.91 examples/s]\rTokenizing train dataset: 77%|███████▋ | 33476/43744 [03…71 tokens truncated…████████▉| 43346/43744 [04:21<00:02, 181.95 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43365/43744 [04:21<00:02, 177.58 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43391/43744 [04:21<00:02, 171.55 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43410/43744 [04:21<00:01, 174.69 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43429/43744 [04:21<00:01, 174.60 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43456/43744 [04:22<00:01, 174.98 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43476/43744 [04:22<00:01, 178.09 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43495/43744 [04:22<00:01, 179.72 examples/s]\rTokenizing train dataset: 99%|█████████▉| 43516/43744 [04:22<00:01, 183.10 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43535/43744 [04:22<00:01, 182.64 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43562/43744 [04:22<00:01, 178.64 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43582/43744 [04:22<00:00, 178.67 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43601/43744 [04:22<00:00, 171.43 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43624/43744 [04:23<00:00, 162.68 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43641/43744 [04:23<00:00, 161.97 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43659/43744 [04:23<00:00, 163.89 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43677/43744 [04:23<00:00, 166.11 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43694/43744 [04:23<00:00, 162.43 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43711/43744 [04:23<00:00, 161.65 examples/s]\rTokenizing train dataset: 100%|█████████▉| 43729/43744 [04:23<00:00, 163.70 examples/s]\rTokenizing train dataset: 100%|██████████| 43744/43744 [04:24<00:00, 165.69 examples/s]\n\rTruncating train dataset: 0%| | 0/43744 [00:00<?, ? examples/s]\rTruncating train dataset: 5%|▍ | 2000/43744 [00:00<00:02, 18008.65 examples/s]\rTruncating train dataset: 9%|▉ | 4000/43744 [00:00<00:02, 18004.62 examples/s]\rTruncating train dataset: 14%|█▎ | 6000/43744 [00:00<00:02, 18406.89 examples/s]\rTruncating train dataset: 18%|█▊ | 8000/43744 [00:00<00:01, 18860.86 examples/s]\rTruncating train dataset: 23%|██▎ | 10000/43744 [00:00<00:01, 18777.21 examples/s]\rTruncating train dataset: 27%|██▋ | 12000/43744 [00:00<00:01, 18851.81 examples/s]\rTruncating train dataset: 34%|███▍ | 15000/43744 [00:00<00:01, 19440.82 examples/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:2000});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 23
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "23",
"yield_time_ms": 60000,
"max_tokens": 2000
}wait result
Script completed
Wall time 46.5 seconds
Output:
{"chunk_id":"389bd7","wall_time_seconds":60.002109953,"session_id":80127,"original_token_count":1806,"output":"\rTruncating train dataset: 39%|███▉ | 17000/43744 [00:00<00:01, 17882.68 examples/s]\rTruncating train dataset: 43%|████▎ | 19000/43744 [00:01<00:01, 17880.45 examples/s]\rTruncating train dataset: 48%|████▊ | 21000/43744 [00:01<00:01, 18201.72 examples/s]\rTruncating train dataset: 57%|█████▋ | 25000/43744 [00:01<00:00, 20694.00 examples/s]\rTruncating train dataset: 64%|██████▍ | 28000/43744 [00:01<00:00, 21113.62 examples/s]\rTruncating train dataset: 73%|███████▎ | 32000/43744 [00:01<00:00, 20749.48 examples/s]\rTruncating train dataset: 82%|████████▏ | 36000/43744 [00:01<00:00, 21336.15 examples/s]\rTruncating train dataset: 89%|████████▉ | 39000/43744 [00:01<00:00, 20393.02 examples/s]\rTruncating train dataset: 96%|█████████▌| 42000/43744 [00:02<00:00, 20375.56 examples/s]\rTruncating train dataset: 100%|██████████| 43744/43744 [00:02<00:00, 19735.20 examples/s]\nDetected kernel version 4.19.0, which is below the recommended minimum of 5.5.0; this can cause the process to hang. It is recommended to upgrade the kernel to the minimum version or higher.\n\n================================================================================\nFIXED TEACHER MODE ENABLED\nTeacher will use the initial policy (base model without LoRA adapters)\nStudent will update with LoRA adapters\n================================================================================\n\n\n================================================================================\nFIXED TEACHER MODE ENABLED\nTeacher will use the initial policy (base model without LoRA adapters)\nStudent will update with LoRA adapters\n================================================================================\n\n\n================================================================================\nFIXED TEACHER MODE ENABLED\nTeacher will use the initial policy (base model without LoRA adapters)\nStudent will update with LoRA adapters\n================================================================================\n\n\n================================================================================\nFIXED TEACHER MODE ENABLED\nTeacher will use the initial policy (base model without LoRA adapters)\nStudent will update with LoRA adapters\n================================================================================\n\n`torch_dtype` is deprecated! Use `dtype` instead!\n`torch_dtype` is deprecated! Use `dtype` instead!\n`torch_dtype` is deprecated! Use `dtype` instead!\n`torch_dtype` is deprecated! Use `dtype` instead!\n[Gloo] Rank 0 is connected to 3 peer ranks. Expected number of connected peer ranks is : 3\n[Gloo] Rank 2 is connected to 3 peer ranks. Expected number of connected peer ranks is : 3\n[Gloo] Rank 1 is connected to 3 peer ranks. Expected number of connected peer ranks is : 3\n[Gloo] Rank 3 is connected to 3 peer ranks. Expected number of connected peer ranks is : 3\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n[Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0\n\rLoading safetensors checkpoint shards: 0% Completed | 0/2 [00:00<?, ?it/s]\n\rLoading safetensors checkpoint shards: 100% Completed | 2/2 [00:01<00:00, 1.94it/s]\n\rLoading safetensors checkpoint shards: 100% Completed | 2/2 [00:01<00:00, 1.94it/s]\n\n\rCapturing CUDA graphs (mixed prefill-decode, PIECEWISE): 0%| | 0/5 [00:00<?, ?it/s]\rCapturing CUDA graphs (mixed prefill-decode, PIECEWISE): 40%|████ | 2/5 [00:00<00:00, 18.18it/s]\rCapturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 5/5 [00:00<00:00, 23.25it/s]\rCapturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|██████████| 5/5 [00:00<00:00, 22.47it/s]\n\rCapturing CUDA graphs (decode, FULL): 0%| | 0/4 [00:00<?, ?it/s]\rCapturing CUDA graphs (decode, FULL): 100%|██████████| 4/4 [00:00<00:00, 34.70it/s]\rCapturing CUDA graphs (decode, FULL): 100%|██████████| 4/4 [00:00<00:00, 34.63it/s]\nThe tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}.\nThe tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}.\nThe tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}.\nThe tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config. The model config and generation config were aligned accordingly, being updated with the tokenizer's values. Updated tokens: {'bos_token_id': None, 'pad_token_id': 151643}.\n[rank3]:W0803 04:48:06.705000 1049 torch/utils/cpp_extension.py:2425] TORCH_CUDA_ARCH_LIST is not set, all archs for visible cards are included for compilation. \n[rank3]:W0803 04:48:06.705000 1049 torch/utils/cpp_extension.py:2425] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:2000});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 24
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "24",
"yield_time_ms": 60000,
"max_tokens": 2000
}wait result
Script completed
Wall time 40.8 seconds
Output:
{"chunk_id":"22cede","wall_time_seconds":60.000864514,"session_id":80127,"original_token_count":1284,"output":"\r 0%| | 0/100 [00:00<?, ?it/s]vLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 2464, avg length: 616.0, speed: 748.5 tok/s\nvLLM generation done - elapsed time: 6.25s, prompts: 4, total tokens: 4070, avg length: 1017.5, speed: 651.5 tok/s\nvLLM generation done - elapsed time: 6.46s, prompts: 4, total tokens: 5296, avg length: 1324.0, speed: 819.3 tok/s\nvLLM generation done - elapsed time: 6.55s, prompts: 4, total tokens: 5165, avg length: 1291.2, speed: 788.6 tok/s\n/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None\n warnings.warn(\n/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None\n warnings.warn(\n/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None\n warnings.warn(\n/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None\n warnings.warn(\nCould not estimate the number of tokens of the input, floating-point operations will not be computed\nCould not estimate the number of tokens of the input, floating-point operations will not be computed\nCould not estimate the number of tokens of the input, floating-point operations will not be computed\nCould not estimate the number of tokens of the input, floating-point operations will not be computed\nvLLM generation done - elapsed time: 3.04s, prompts: 4, total tokens: 2592, avg length: 648.0, speed: 854.0 tok/s\nvLLM generation done - elapsed time: 3.96s, prompts: 4, total tokens: 3897, avg length: 974.2, speed: 984.0 tok/s\nvLLM generation done - elapsed time: 5.97s, prompts: 4, total tokens: 3871, avg length: 967.8, speed: 648.7 tok/s\nvLLM generation done - elapsed time: 6.38s, prompts: 4, total tokens: 4731, avg length: 1182.8, speed: 741.5 tok/s\n\r 1%| | 1/100 [00:13<22:33, 13.68s/it]vLLM generation done - elapsed time: 2.86s, prompts: 4, total tokens: 2381, avg length: 595.2, speed: 832.3 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 2706, avg length: 676.5, speed: 829.1 tok/s\nvLLM generation done - elapsed time: 6.31s, prompts: 4, total tokens: 4613, avg length: 1153.2, speed: 731.3 tok/s\nvLLM generation done - elapsed time: 6.54s, prompts: 4, total tokens: 5352, avg length: 1338.0, speed: 818.5 tok/s\nvLLM generation done - elapsed time: 2.73s, prompts: 4, total tokens: 2158, avg length: 539.5, speed: 789.1 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3004, avg length: 751.0, speed: 902.4 tok/s\nvLLM generation done - elapsed time: 3.45s, prompts: 4, total tokens: 2723, avg length: 680.8, speed: 788.2 tok/s\nvLLM generation done - elapsed time: 3.80s, prompts: 4, total tokens: 3525, avg length: 881.2, speed: 928.4 tok/s\n\r 2%|▏ | 2/100 [00:25<20:47, 12.73s/it]\r \r{'loss': 0.0097, 'grad_norm': 0.1635645180940628, 'learning_rate': 5e-06, 'on_policy_loss': 0.0097, 'epoch': 0.0}\n\r 2%|▏ | 2/100 [00:25<20:47, 12.73s/it]vLLM generation done - elapsed time: 2.63s, prompts: 4, total tokens: 2085, avg length: 521.2, speed: 793.8 tok/s\nvLLM generation done - elapsed time: 4.05s, prompts: 4, total tokens: 3766, avg length: 941.5, speed: 929.8 tok/s\nvLLM generation done - elapsed time: 4.40s, prompts: 4, total tokens: 3605, avg length: 901.2, speed: 819.8 tok/s\nvLLM generation done - elapsed time: 6.52s, prompts: 4, total tokens: 5265, avg length: 1316.2, speed: 807.1 tok/s\nvLLM generation done - elapsed time: 4.50s, prompts: 4, total tokens: 3023, avg length: 755.8, speed: 671.2 tok/s\nvLLM generation done - elapsed time: 6.14s, prompts: 4, total tokens: 3763, avg length: 940.8, speed: 613.1 tok/s\nvLLM generation done - elapsed time: 6.38s, prompts: 4, total tokens: 4568, avg length: 1142.0, speed: 715.8 tok/s\nvLLM generation done - elapsed time: 6.60s, prompts: 4, total tokens: 5927, avg length: 1481.8, speed: 897.8 tok/s\n\r 3%|▎ | 3/100 [00:40<22:08, 13.70s/it]vLLM generation done - elapsed time: 4.18s, prompts: 4, total tokens: 3122, avg length: 780.5, speed: 747.3 tok/s\nvLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 3134, avg length: 783.5, speed: 609.4 tok/s\nvLLM generation done - elapsed time: 6.30s, prompts: 4, total tokens: 4625, avg length: 1156.2, speed: 734.2 tok/s\nvLLM generation done - elapsed time: 6.46s, prompts: 4, total tokens: 5761, avg length: 1440.2, speed: 891.8 tok/s\nvLLM generation done - elapsed time: 4.36s, prompts: 4, total tokens: 3329, avg length: 832.2, speed: 763.9 tok/s\nvLLM generation done - elapsed time: 5.28s, prompts: 4, total tokens: 4135, avg length: 1033.8, speed: 782.4 tok/s\nvLLM generation done - elapsed time: 5.64s, prompts: 4, total tokens: 3649, avg length: 912.2, speed: 647.3 tok/s\nvLLM generation done - elapsed time: 6.62s, prompts: 4, total tokens: 6153, avg length: 1538.2, speed: 930.0 tok/s\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:800});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 25
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "25",
"yield_time_ms": 60000,
"max_tokens": 800
}wait result
Script completed
Wall time 43.6 seconds
Output:
Warning: truncated output (original token count: 865)
Total output lines: 1
{"chunk_id":"e727af","wall_time_seconds":60.000545461,"session_id":80127,"original_token_count":2888,"output":"Warning: truncated output (original token count: 2888)\nTotal output lines: 280\n\n\r 4%|▍ | 4/100 [00:55<22:37, 14.14s/it]\r \r{'loss': 0.0081, 'grad_norm': 0.1333799660205841, 'learning_rate': 5e-06, 'on_policy_loss': 0.0081, 'epoch': 0.0}\n\r 4%|▍ | 4/100 [00:55<22:37, 14.14s/it]vLLM generation done - elapsed time: 2.73s, prompts: 4, total tokens: 2831, avg length: 707.8, speed: 1036.5 tok/s\nvLLM generation done - elapsed time: 5.10s, prompts: 4, total tokens: 4180, avg length: 1045.0, speed: 820.2 tok/s\nvLLM generation done - elapsed time: 5.16s, prompts: 4, total tokens: 4794, avg length: 1198.5, speed: 929.8 tok/s\nvLLM generation done - elapsed time: 6.50s, prompts: 4, total tokens: 4658, avg length: 1164.5, speed: 716.8 tok/s\nvLLM generation done - elapsed time: 2.96s, prompts: 4, total tokens: 2473, avg length: 618.2, speed: 835.4 tok/s\nvLLM generation done - elapsed time: 4.97s, prompts: 4, total tokens: 4627, avg length: 1156.8, speed: 931.5 tok/s\nvLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 4256, avg length: 1064.0, speed: 832.5 tok/s\nvLLM generation done - elapsed time: 6.00s, prompts: 4, total tokens: 5967, avg length: 1491.8, speed: 995.2 tok/s\n\r 5%|▌ | 5/100 [01:09<22:25, 14.17s/it]vLLM generation done - elapsed time: 4.30s, prompts: 4, total tokens: 3268, avg length: 817.0, speed: 760.1 tok/s\nvLLM generation done - elapsed time: 4.99s, p…65 tokens truncated…===\n\n\r 7%|▋ | 7/100 [01:38<22:18, 14.39s/it]vLLM generation done - elapsed time: 2.86s, prompts: 4, total tokens: 2407, avg length: 601.8, speed: 841.5 tok/s\nvLLM generation done - elapsed time: 4.94s, prompts: 4, total tokens: 5370, avg length: 1342.5, speed: 1086.6 tok/s\nvLLM generation done - elapsed time: 5.93s, prompts: 4, total tokens: 4374, avg length: 1093.5, speed: 737.4 tok/s\nvLLM generation done - elapsed time: 6.51s, prompts: 4, total tokens: 5141, avg length: 1285.2, speed: 790.0 tok/s\nvLLM generation done - elapsed time: 2.89s, prompts: 4, total tokens: 2224, avg length: 556.0, speed: 769.9 tok/s\nvLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 3059, avg length: 764.8, speed: 899.0 tok/s\nvLLM generation done - elapsed time: 5.66s, prompts: 4, total tokens: 3480, avg length: 870.0, speed: 614.7 tok/s\nvLLM generation done - elapsed time: 5.82s, prompts: 4, total tokens: 2978, avg length: 744.5, speed: 511.9 tok/s\n\r 8%|▊ | 8/100 [01:52<21:53, 14.28s/it]\r \r{'loss': 0.0065, 'grad_norm': 0.15239836275577545, 'learning_rate': 5e-06, 'on_policy_loss': 0.0065, 'epoch': 0.01}\n\r 8%|▊ | 8/100 [01:52<21:53, 14.28s/it]vLLM generation done - elapsed time: 3.13s, prompts: 4, total tokens: 2439, avg length: 609.8, speed: 779.0 tok/s\nvLLM generation done - elapsed time: 6.15s, prompts: 4, total tokens: 3777, avg length: 944.2, speed: 614.0 tok/s\nvLLM generation done - elapsed time: 6.24s, prompts: 4, total tokens: 4152, avg length: 1038.0, speed: 665.0 tok/s\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python - <<'PY'\\nimport re,zlib\\nb=open('/workspace/paper.pdf','rb').read()\\nprint('streams',b.count(b'stream'))\\nfor i,m in enumerate(re.finditer(rb'stream\\\\r?\\\\n',b)):\\n e=b.find(b'endstream',m.end()); raw=b[m.end():e].rstrip(b'\\\\r\\\\n')\\n try: d=zlib.decompress(raw)\\n except: continue\\n if b'AIME' in d or b'OPSD' in d or b'KL' in d:\\n print('\\\\n---STREAM',i,len(d),'---')\\n for line in d.splitlines():\\n if any(k in line for k in [b'AIME',b'OPSD',b'KL',b'beta',b'Teacher',b'teacher',b'clip']): print(line[:1000])\\nPY\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":20000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
Warning: truncated output (original token count: 15555)
Total output lines: 139
streams 492
---STREAM 4 14909 ---
b'/F133 9.9626 Tf 1.019 0 0 1 75.366 571.88 Tm [(Kno)25(wledge)-244(distillation)-244(impro)14(v)15(es)-244(lar)18(ge)-244(language)]TJ 1.02 0 0 1 75.366 559.925 Tm [(model)-533(\\050LLM\\051)-532(reasoning)-533(by)-533(compressing)-533(the)]TJ 1.02 0 0 1 75.366 547.97 Tm [(kno)25(wledge)-419(of)-418(a)-419(teacher)-418(LLM)-418(to)-419(train)-418(smaller)]TJ 1.02 0 0 1 75.366 536.015 Tm [(LLMs.)-541(On-polic)15(y)-324(distillation)-324(adv)25(ances)-324(this)-324(ap-)]TJ 1.015 0 0 1 75.366 524.059 Tm [(proach)-246(by)-245(ha)19(ving)-245(the)-246(student)-246(sample)-246(its)-245(o)24(wn)-245(tra-)]TJ 1.02 0 0 1 75.366 512.104 Tm [(jectories)-391(while)-391(a)-392(teacher)-391(LLM)-391(pro)15(vides)-392(dense)]TJ 1.02 0 0 1 75.366 500.149 Tm [(tok)10(en-le)24(v)15(el)-271(supervision,)-277(addressing)-270(the)-271(distrib)20(u-)]TJ 1.02 0 0 1 75.366 488.194 Tm [(tion)-411(mismatch)-410(between)-411(training)-411(and)-411(inference)]TJ 1.02 0 0 1 75.366 476.239 Tm [(in)-270(of)24(f)1(-poli'
b'/F154 9.9626 Tf 1 0 0 1 236.638 237.135 Tm [(https:)]TJ -161.87 -11.955 Td [(//github.com/siyan-)-50(zhao/OPSD)]TJ'
b' [(\\051,)-246(or)-246(kno)25(wledge)-246(distillation,)-246(where)-246(recent)-246(w)10(ork)]TJ 1.02 0 0 1 307.44 505.058 Tm [(has)-265(sho)24(wn)-265(that)-266(distil)1(lation)-266(from)-265(adv)24(anced)-265(teacher)-265(models)]TJ 1.02 0 0 1 307.44 493.103 Tm [(can)-278(outperform)-277(RL)-278(in)-277(both)-278(performance)-277(and)-278(training)-277(ef)24(\\002)1(-)]TJ 1 0 0 1 307.44 481.147 Tm [(cienc)15(y)-250(\\050)]TJ'
b' [(\\051.)-309(T)35(raditional)-249(kno)25(wledge)-249(distillation)]TJ 0.98 0 0 1 307.44 319.753 Tm [(pro)15(vides)-204(dense)-204(tok)11(en-le)25(v)15(el)-204(s)1(up)-1(e)1(rvision)-204(from)-204(a)-204(teacher)-204(model)]TJ 1.02 0 0 1 307.44 307.797 Tm [(b)20(ut)-273(relies)-272(on)-272(of)25(f-polic)15(y)-272(data)-273(\\050)]TJ'
b' [(\\051.)-385(Recent)]TJ 1.009 0 0 1 307.44 295.842 Tm [(adv)25(ances)-248(in)-248(on-polic)15(y)-248(distillation\\227where)-248(a)-248(student)-247(model)]TJ 0.987 0 0 1 307.44 283.887 Tm [(samples)-253(its)-253(o)25(wn)-253(trajectories)-253(while)-253(a)-253(teacher)-253(polic)16(y)-253(pro)15(vides)]TJ 0.985 0 0 1 307.44 271.932 Tm [(dense)-252(tok)10(en-le)25(v)15(el)-252(supervision\\227ha)21(v)15(e)-253(demonstrated)-252(superior)]TJ 1.02 0 0 1 307.44 259.977 Tm [(sample)-255(ef)24(\\002)1(cienc)14(y)-255(by)-256(combining)-255(the)-256(distri)1(b)19(utional)-255(realism)]TJ 1.02 0 0 1 307.44 248.022 Tm [(of)-252(on-polic)14(y)-252(training)-252(with)-252(dense)-253(feedback)-252(\\050)]TJ'
b' [(\\051.)]TJ 0.995 0 0 1 306.972 218.134 Tm [(While)-253(on-polic)15(y)-252(distillation)-253(has)-253(sho)26(wn)-253(strong)-253(performance,)]TJ 0.988 0 0 1 307.44 206.179 Tm [(it)-254(relies)-253(on)-254(a)-254(distinct)-253(teacher)-254(model)-253(to)-254(supervise)-254(the)-253(student.)]TJ 1.009 0 0 1 307.44 194.223 Tm [(Gi)25(v)15(en)-247(that)-246(modern)-247(LLMs)-246(already)-246(e)14(xhibit)-246(strong)-246(reasoning)]TJ 1.02 0 0 1 307.44 182.268 Tm [(capabilities,)-363(we)-339(ask)-340(this)-339(research)-340(question:)]TJ/F145 9.9626 Tf [-492(can)-340(a)-339(model)]TJ 0.98 0 0 1 307.44 170.313 Tm [(ef)18(fectively)-244(serve)-244(as)-245(its)-244(own)-244(teac)15(her)-244(thr)46(ough)-244(self-distillation?)]TJ/F133 9.9626 Tf 0.98 0 0 1 307.44 158.358 Tm [(Our)-255(approach)-256(is)-255(inspired)-256(by)-255(human)-255(learning:)-317(after)-256(solving)-255(a)]TJ 0.987 0 0 1 307.44 146.403 Tm [(problem)-253(incorrectly)66(,)-253(a)-253(student)-253(can)-253(e)15(xamine)-25'
---STREAM 10 18970 ---
b'/F145 8.9664 Tf 1.005 0 0 1 54.893 576.227 Tm [(F)45(igur)37(e)-249(1.)]TJ/F127 8.9664 Tf [-309(Ov)10(er)10(view)-249(of)-249(On-P)20(olicy)-249(Self-Distillation)-249(\\050O)1(PSD\\051:)]TJ/F133 8.9664 Tf [-249(Gi)25(v)14(en)-248(a)-249(reasoning)-249(dataset)]TJ/F76 8.9664 Tf 1 0 0 1 372.222 576.227 Tm [(S)]TJ/F72 8.9664 Tf [-360(=)]TJ/F76 8.9664 Tf [-286(f)]TJ/F72 8.9664 Tf [(\\050)]TJ/F74 8.9664 Tf [(x)]TJ/F75 5.9776 Tf 31.952 -0.996 Td [(i)]TJ/F74 8.9664 Tf 3.162 0.996 Td [(;)-171(y)]TJ/F75 5.9776 Tf 8.956 3.809 Td [(?)]TJ -0.322 -5.474 Td [(i)]TJ/F72 8.9664 Tf 4.639 1.665 Td [(\\051)]TJ/F76 8.9664 Tf [(g)]TJ/F75 5.9776 Tf 8.191 3.809 Td [(N)]TJ 0 -5.474 Td [(i)]TJ/F73 5.9776 Tf [(=1)]TJ/F133 8.9664 Tf 1.005 0 0 1 441.205 576.227 Tm [(,)-249(we)-249(instantiate)-249(tw)10(o)-248(policies)]TJ 1.02 0 0 1 55.44 566.265 Tm [(from)-339(the)-340(same)-339(LLM:)-339(a)]TJ/F145 8.9664 Tf [-339(student)-340(policy)]TJ/F74 8.9664 Tf 1 0 0 1 199.594 566.265 Tm [(p)]TJ/F75 5.9776 Tf 4.626 -0.997 '
b'/F133 9.9626 Tf 0.997 0 0 1 55.44 503.5 Tm [(by)-251(this,)-250(we)-251(instantiate)-250(both)-251(the)-251(teacher)-250(and)-251(student)-251(policies)]TJ 1.02 0 0 1 55.44 491.545 Tm [(from)-331(a)-331(single)-330(LLM.)-331(The)-331(teacher)-330(polic)14(y)-331(is)-330(pro)14(vided)-330(with)]TJ 1 0 0 1 55.44 479.59 Tm [(pri)25(vile)15(ged)-251(information)]TJ/F35 9.9626 Tf [-251(y)]TJ/F34 6.9738 Tf 97.286 3.615 Td [(?)]TJ/F133 9.9626 Tf 4.58 -3.615 Td [(,)-251(such)-251(as)-251(the)-251(ground-truth)-251(answer)]TJ 1.02 0 0 1 55.44 467.635 Tm [(or)-322(a)-323(reference)-322(chain-of-thought,)-342(while)-322(the)-323(student)-322(polic)14(y)]TJ 1.02 0 0 1 55.44 455.679 Tm [(conditions)-248(only)-248(on)-248(the)-248(problem)]TJ/F35 9.9626 Tf 1 0 0 1 184.907 455.679 Tm [(x)]TJ/F133 9.9626 Tf 1.02 0 0 1 190.601 455.679 Tm [(.)-313(Concretely)64(,)-249(the)-248(teacher)]TJ 1.02 0 0 1 55.44 443.724 Tm [(polic)15(y)]TJ/F35 9.9626 Tf 1 0 0 1 84.095 443.724 Tm [(p)]TJ/F34 6.9738 Tf 5.012 -1.'
b' 0.98 0 0 1 63.44 136.715 Tm [(W)82(e)-236(introduce)-237(On-Polic)16(y)-236(Self-Distillation)-236(\\050OPSD\\051,)-236(a)-236(no)15(v)15(el)]TJ 1.02 0 0 1 63.909 124.76 Tm [(frame)25(w)9(ork)-383(that)-384(enables)-384(a)-384(single)-384(model)-383(to)-384(act)-384(as)-384(both)]TJ 1.02 0 0 1 63.909 112.805 Tm [(teacher)-248(and)-248(student,)-248(le)24(v)15(eraging)-248(ground-truth)-248(answers)-248(to)]TJ 0.988 0 0 1 63.909 100.85 Tm [(pro)15(vide)-253(dense)-253(tok)10(en-le)26(v)15(el)-253(supervision)-253(on)-253(student)-253(rollouts.)]TJ'
b' 1.016 0 0 1 63.44 88.894 Tm [(W)79(e)-245(introduce)-245(a)-245(per)19(-tok)10(en)-245(pointwise)-245(KL)-245(clipping)-245(mecha-)]TJ 0.981 0 0 1 63.909 76.939 Tm [(nism)-254(that)-254(stabilizes)-253(training)-254(and)-254(impro)15(v)15(es)-253(performance)-254(as)]TJ'
b' 0.996 0 0 1 315.44 479.59 Tm [(W)80(e)-251(e)25(v)25(aluate)-251(OPSD)-251(on)-251(three)-251(competition-le)25(v)15(el)-251(mathemat-)]TJ 1.02 0 0 1 315.909 467.635 Tm [(ical)-330(reasoning)-331(tasks,)-352(demonstrating)-330(that)-331(it)-330(matches)-331(the)]TJ 0.993 0 0 1 315.909 455.679 Tm [(performance)-251(of)-251(GRPO)-251(with)-251(signi\\002cantly)-252(im)1(pro)15(v)15(ed)-251(tok)10(en)]TJ 1 0 0 1 315.909 443.724 Tm [(ef)25(\\002cienc)15(y)-250(and)-250(outperform)-250(supervised)-250(\\002ne-tuning.)]TJ'
b" 1.02 0 0 1 315.44 431.769 Tm [(W)78(e)-430(analyze)-431(the)-431(impact)-430(of)-431(dif)24(f)1(erent)-431(di)25(v)14(er)18(gence)-431(objec-)]TJ 1.02 0 0 1 315.909 419.814 Tm [(ti)25(v)14(es,)-419(the)-385(ef)25(fect)-385(of)-384(student)-385(generation)-384(length,)-420(and)-385(stu-)]TJ 1 0 0 1 315.909 407.859 Tm [(dent\\226teacher)-250(generation)-250(styles.)]TJ/F127 11.9552 Tf -8.469 -34.977 Td [(2.)-250(Backgr)18(ound)]TJ/F127 9.9626 Tf 0 -19.373 Td [(2.1.)-250(Kno)10(wledge)-250(Distillation)-250(f)25(or)-250(A)50(utor)18(egr)18(essi)10(v)10(e)-250(Lar)10(ge)]TJ 17.435 -11.955 Td [(Language)-250(Models)]TJ/F133 9.9626 Tf 1.02 0 0 1 307.44 322.902 Tm [(Kno)24(wledge)-278(distillation)-278(transfers)-279(kno)25(wledge)-279(from)-278(a)-279(lar)18(ger)]TJ 1.02 0 0 1 307.44 310.947 Tm [(teacher)-306(model)-306(to)-305(a)-306(smaller)-306(student)-306(model)-306(by)-305(training)-306(the)]TJ 1.02 0 0 1 307.44 298.991 Tm [(student)-452(to)-452(mimic)-452(the)-452(teacher')54(s)-452(b"
b" [(\\051.)-628(The)-353(core)]TJ 1.02 0 0 1 307.44 275.081 Tm [(insight)-411(is)-412(that)-411(the)-411(teacher')54(s)-411(soft)-412(probability)-411(distrib)20(ution)]TJ 1.02 0 0 1 307.44 263.126 Tm [(o)15(v)14(er)-326(classes)-326(contains)-326(richer)-327(information)-326(than)-326(hard)-326(labels)]TJ 1.02 0 0 1 307.44 251.171 Tm [(alone,)-425(as)-390(it)-389(re)25(v)14(eals)-389(the)-389(teacher')54(s)-390(learned)-389(similarities)-389(be-)]TJ 1.007 0 0 1 307.44 239.216 Tm [(tween)-248(classes.)-308(F)14(or)-248(auto-re)15(gressi)25(v)15(e)-249(language)-248(models,)-248(gi)25(v)14(en)]TJ 0.986 0 0 1 307.44 227.26 Tm [(a)-254(dataset)]TJ/F38 9.9626 Tf 1 0 0 1 344.062 227.26 Tm [(S)]TJ/F32 9.9626 Tf [-353(=)]TJ/F38 9.9626 Tf [-278(f)]TJ/F32 9.9626 Tf [(\\050)]TJ/F35 9.9626 Tf [(x;)-166(y)]TJ/F34 6.9738 Tf 44.285 3.616 Td [(?)]TJ/F32 9.9626 Tf 4.58 -3.616 Td [(\\051)]TJ/F38 9.9626 Tf [(g)]TJ/F133 9.9626 Tf 0.986 0 0 1 404.274 227.26 Tm [(where)]TJ/F35 9.9626 Tf 1 0 0 1 430.763 227.26 "
---STREAM 12 20820 ---
b'/F145 8.9664 Tf 0.997 0 0 1 54.938 628.469 Tm [(T)92(able)-249(1.)]TJ/F133 8.9664 Tf [-310(Comparison)-249(of)-249(training)-250(methods)-249(for)-249(reasoning)-249(tasks.)-310(On-Polic)15(y)-249(Self-Distillation)-249(\\050OPSD\\051)-249(combines)-250(the)-249(adv)25(antages)-249(of)-249(on-polic)15(y)]TJ 1 0 0 1 55.44 618.506 Tm [(training)-250(with)-250(dense)-250(feedback)-250(without)-250(requiring)-250(an)-250(e)15(xternal)-250(teacher)-250(model.)]TJ'
b' [(\\051)-287(addresses)-286(this)-286(by)-286(training)-287(the)-286(student)-286(on)-286(its)]TJ 1.02 0 0 1 55.44 530.608 Tm [(o)24(wn)-368(generated)-369(sequences)]TJ/F32 9.9626 Tf 1 0 0 1 165.308 530.608 Tm [(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F38 9.9626 Tf [-547(\\030)]TJ/F35 9.9626 Tf [-512(p)]TJ/F34 6.9738 Tf 27.506 -1.494 Td [(S)]TJ/F32 9.9626 Tf 5.771 1.494 Td [(\\050)]TJ/F38 9.9626 Tf [(\\001j)]TJ/F35 9.9626 Tf [(x)]TJ/F32 9.9626 Tf [(\\051)]TJ/F133 9.9626 Tf 1.02 0 0 1 217.563 530.608 Tm [(,)-400(obtaining)-368(dense)]TJ 1.02 0 0 1 55.44 518.653 Tm [(tok)10(en-le)24(v)15(el)-297(feedback)-296(from)-297(the)-296(teacher)-297(on)-297(these)-296(on-polic)14(y)]TJ 1 0 0 1 55.44 506.698 Tm [(samples:)]TJ/F38 9.9626 Tf 0 -22.211 Td [(L)]TJ/F133 6.9738 Tf 6.872 -1.494 Td [(On-Polic)15(y)-250(Distillation)]TJ/F32 9.9626 Tf 62.195 1.494 Td [(\\050)]TJ/F35 9.9626 Tf [(\\022)]TJ/F32 9.9626 Tf [-28(\\051)-278(=)]TJ/F66 9.9626 Tf [-277(E)]TJ/F34 6.9738 Tf 32.627 -1.494 Td [(x)]TJ/F37 6.9738 Tf [(\\0'
b" [(\\051,)-372(where)-347(the)-347(student)-347(iterati)24(v)15(ely)-347(im-)]TJ 1.017 0 0 1 55.44 436.666 Tm [(pro)15(v)14(es)-245(by)-245(learning)-246(from)-245(the)-245(teacher')54(s)-246(guidance)-245(on)-245(its)-246(o)25(wn)]TJ 0.98 0 0 1 55.44 424.711 Tm [(outputs,)-214(combining)-204(the)-203(on-polic)15(y)-204(rele)26(v)25(ance)-204(of)-203(reinforcement)]TJ 1.02 0 0 1 55.44 412.756 Tm [(learning)-258(with)-258(the)-258(dense)-258(re)25(w)10(ard)-258(signal)-258(of)-258(supervised)-258(learn-)]TJ 1.02 0 0 1 55.44 400.801 Tm [(ing,)-388(thereby)-360(mitig)5(ating)-360(e)15(xposure)-360(bias)-360(while)-359(maintaining)]TJ 1 0 0 1 55.44 388.846 Tm [(computational)-250(ef)25(\\002cienc)15(y)65(.)]TJ/F127 9.9626 Tf 0 -25.134 Td [(2.2.)-250(Reinf)25(or)18(cement)-250(Lear)15(ning)-250(with)-250(V)100(eri\\002able)-250(Rewards)]TJ/F133 9.9626 Tf 1.02 0 0 1 55.44 345.06 Tm [(Reinforcement)-393(learning)-394(with)-393(v)14(eri\\002able)-393(re)25(w)9(ards)-393(\\050RL)98(VR\\051)]TJ 1.02 0"
b' 1.02 0 0 1 307.44 578.429 Tm [(at)-323(the)-322(sequence)-323(le)25(v)14(el.)-536(The)-322(GRPO)-323(objecti)25(v)14(e)-322(incorporates)]TJ 1.015 0 0 1 307.44 566.474 Tm [(a)-246(clipped)-246(surrog)5(ate)-247(loss)-246(to)-246(moderate)-246(polic)15(y)-247(updates,)-246(along)]TJ 1.02 0 0 1 307.082 554.519 Tm [(with)-313(a)-313(re)25(v)15(erse)-313(KL)-313(penalty)-313(to)-313(pre)25(v)15(ent)-313(e)15(xcessi)24(v)15(e)-313(de)24(viation)]TJ 1 0 0 1 307.44 542.564 Tm [(from)-250(a)-250(reference)-250(polic)15(y:)]TJ/F38 9.9626 Tf 24.313 -31.661 Td [(L)]TJ/F133 6.9738 Tf 6.872 -1.495 Td [(GRPO)]TJ/F32 9.9626 Tf 19.097 1.495 Td [(\\050)]TJ/F35 9.9626 Tf [(\\022)]TJ/F32 9.9626 Tf [-28(\\051)-278(=)]TJ/F66 9.9626 Tf [-277(E)]TJ/F34 6.9738 Tf 54.45 -1.334 Td [(x)]TJ/F37 6.9738 Tf [(\\030S)]TJ/F34 6.9738 Tf -21.822 -6.247 Td [(o)]TJ/F30 4.9813 Tf 3.932 -0.996 Td [(1)]TJ/F34 6.9738 Tf 3.888 0.996 Td [(;:::;o)]TJ/F33 4.9813 Tf 15.763 -1.002 Td [(G)]TJ/F37 6.9738 Tf 5.799 1.002 Td [(\\030)]TJ/F34'
b'/F38 9.9626 Tf 485.579 504.069 Td [(j)]TJ/F35 9.9626 Tf [(o)]TJ/F34 6.9738 Tf 7.596 -1.495 Td [(i)]TJ/F38 9.9626 Tf 3.317 1.495 Td [(j)]TJ/F37 6.9738 Tf 7.219 20.145 Td [(j)]TJ/F34 6.9738 Tf [(o)]TJ/F33 4.9813 Tf 6.299 -0.996 Td [(i)]TJ/F37 6.9738 Tf 3.156 0.996 Td [(j)]TJ/F25 9.9626 Tf -10.74 -3.847 Td [(X)]TJ/F34 6.9738 Tf -0.311 -21.099 Td [(n)]TJ/F31 6.9738 Tf [(=1)]TJ/F32 9.9626 Tf -136.63 -13.47 Td [(min)-167(\\050)]TJ/F35 9.9626 Tf [(\\032)]TJ/F34 6.9738 Tf 27.29 4.114 Td [(n)]TJ 0 -6.577 Td [(i)]TJ/F35 9.9626 Tf 5.423 2.463 Td [(A)]TJ/F34 6.9738 Tf 7.472 -1.494 Td [(i)]TJ/F35 9.9626 Tf 3.317 1.494 Td [(;)]TJ/F133 9.9626 Tf [-167(clip)]TJ/F32 9.9626 Tf [-166(\\050)]TJ/F35 9.9626 Tf [(\\032)]TJ/F34 6.9738 Tf 30.057 4.114 Td [(n)]TJ 0 -6.577 Td [(i)]TJ/F35 9.9626 Tf 5.423 2.463 Td [(;)]TJ/F32 9.9626 Tf [-167(1)]TJ/F38 9.9626 Tf [-222(\\000)]TJ/F35 9.9626 Tf [-222(";)]TJ/F32 9.9626 Tf [-167(1)-222(+)]TJ/F35 9.9626 Tf [-222(")]TJ/F32 9.9626 Tf [(\\051)]TJ/F35 9.9626 Tf [-167(A)]TJ/F34 6.9'
b'/F34 6.9738 Tf 360.418 423.107 Td [(\\031)]TJ/F33 4.9813 Tf 4.659 -1.057 Td [(\\022)]TJ/F133 4.9813 Tf 3.372 -1.681 Td [(old)]TJ/F31 6.9738 Tf 7.363 2.738 Td [(\\050)]TJ/F34 6.9738 Tf [(o)]TJ/F33 4.9813 Tf 7.046 2.402 Td [(n)]TJ 0 -4.674 Td [(i)]TJ/F37 6.9738 Tf 4.885 2.272 Td [(j)]TJ/F34 6.9738 Tf [(x;o)]TJ/F33 4.9813 Tf 13.183 2.905 Td [(<n)]TJ 0 -5.177 Td [(i)]TJ/F31 6.9738 Tf 10.282 2.272 Td [(\\051)]TJ/F133 9.9626 Tf 1.02 0 0 1 418.213 427.113 Tm [(is)-265(the)-266(importance)-265(ratio,)]TJ/F35 9.9626 Tf 1 0 0 1 515.13 427.113 Tm [(\\031)]TJ/F34 6.9738 Tf 5.679 -1.495 Td [(\\022)]TJ/F133 4.9813 Tf 3.795 -0.996 Td [(old)]TJ/F133 9.9626 Tf 1.02 0 0 1 534.663 427.113 Tm [(is)]TJ 1.02 0 0 1 307.44 412.521 Tm [(the)-331(polic)14(y)-331(before)-332(the)-331(update,)-353(and)]TJ/F35 9.9626 Tf 1 0 0 1 448.22 412.521 Tm [(")]TJ/F133 9.9626 Tf 1.02 0 0 1 456.234 412.521 Tm [(controls)-331(the)-332(clipping)]TJ 1 0 0 1 307.44 400.566 Tm [(range.)]TJ 1.005 0 0 1 306.972 382.633 Tm [(While)-250(RL)'
---STREAM 16 25803 ---
b'/F127 9.9626 Tf 55.082 715.286 Td [(Algorithm)-250(1)]TJ/F133 9.9626 Tf [-250(On-Polic)15(y)-250(Self-Distillation)-250(\\050OPSD\\051)]TJ'
b'/F133 9.9626 Tf [-2000(Calculate)-250(loss)]TJ/F38 9.9626 Tf [-250(L)]TJ/F31 6.9738 Tf 91.883 -1.494 Td [(OPSD)]TJ/F32 9.9626 Tf 22.368 1.494 Td [(\\050)]TJ/F35 9.9626 Tf [(\\022)]TJ/F32 9.9626 Tf [-28(\\051)]TJ/F38 9.9626 Tf [-278(\\040)]TJ/F31 6.9738 Tf 32.488 3.923 Td [(1)]TJ'
b'/F133 9.9626 Tf 1.003 0 0 1 55.44 519.988 Tm [(matches)-249(the)-249(teacher)-249(and)-249(student)-249(ne)15(xt-tok)10(en)-250(di)1(strib)19(utions)-249(at)]TJ 0.98 0 0 1 55.44 508.033 Tm [(each)-249(position.)-314(Gi)25(v)15(en)-249(a)-249(student-generated)-249(sequence)]TJ/F32 9.9626 Tf 1 0 0 1 256.136 508.033 Tm [(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F133 9.9626 Tf 0.98 0 0 1 260.694 508.033 Tm [(,)-250(de\\002ne)]TJ 1 0 0 1 55.44 496.078 Tm [(the)-250(trajectory-a)20(v)15(eraged,)-250(tok)10(en-wise)-250(di)25(v)15(er)18(gence)]TJ/F35 9.9626 Tf 8.353 -33.044 Td [(D)]TJ/F25 9.9626 Tf 8.525 8.07 Td [(\\000)]TJ/F35 9.9626 Tf 4.567 -8.07 Td [(p)]TJ/F34 6.9738 Tf 5.012 -1.494 Td [(T)]TJ/F38 9.9626 Tf 7.936 1.494 Td [(k)]TJ/F35 9.9626 Tf [-167(p)]TJ/F34 6.9738 Tf 11.655 -1.494 Td [(S)]TJ/F25 9.9626 Tf 5.771 9.564 Td [(\\001)]TJ/F32 9.9626 Tf 4.566 -8.07 Td [(\\050)-69(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F38 9.9626 Tf [-314(j)]TJ/F35 9.9626 Tf [-277(x)]TJ/F32 9.9626 Tf [(\\051)]TJ/F63 9.9626 Tf [-278(,'
b'/F38 9.9626 Tf 153.291 456.201 Td [(j)]TJ/F32 9.9626 Tf [-69(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F38 9.9626 Tf [-36(j)]TJ/F37 6.9738 Tf 16.627 20.145 Td [(j)]TJ/F31 6.9738 Tf [-84(^)]TJ/F34 6.9738 Tf [653(y)]TJ/F37 6.9738 Tf [-35(j)]TJ/F25 9.9626 Tf -2.684 -3.847 Td [(X)]TJ/F34 6.9738 Tf -0.31 -21.099 Td [(n)]TJ/F31 6.9738 Tf [(=1)]TJ/F35 9.9626 Tf 16.672 11.634 Td [(D)]TJ/F25 9.9626 Tf 8.525 14.048 Td [(\\022)]TJ/F35 9.9626 Tf 7.334 -14.048 Td [(p)]TJ/F34 6.9738 Tf 5.012 -1.494 Td [(T)]TJ/F32 9.9626 Tf 6.276 1.494 Td [(\\050)]TJ/F38 9.9626 Tf [(\\001)-278(j)]TJ/F35 9.9626 Tf [-278(x;)-166(y)]TJ/F34 6.9738 Tf 30.308 4.114 Td [(?)]TJ/F35 9.9626 Tf 4.58 -4.114 Td [(;)]TJ/F32 9.9626 Tf [-235(^)]TJ/F35 9.9626 Tf [568(y)]TJ/F34 6.9738 Tf 9.312 -1.494 Td [(<n)]TJ/F32 9.9626 Tf 11.65 1.494 Td [(\\051)]TJ/F25 9.9626 Tf -107.856 -22.593 Td [(\\015)]TJ 0 -5.977 Td [(\\015)]TJ/F35 9.9626 Tf 8.302 -2.491 Td [(p)]TJ/F34 6.9738 Tf 5.013 -1.494 Td [(S)]TJ/F32 9.9626 Tf 5.771 1.494 Td [(\\050)]TJ/F38 9.9626 Tf [('
b' 1.02 0 0 1 307.44 519.988 Tm [(training)-345(signal)-346(to)-345(be)-346(dominated)-345(by)-345(stylistic)-346(patterns.)-605(T)79(o)]TJ 0.985 0 0 1 307.44 508.033 Tm [(address)-254(this,)-253(we)-254(apply)-254(pointwise)-254(clipping)-253(to)-254(the)-254(v)21(ocab)20(ulary-)]TJ 1.02 0 0 1 307.44 496.078 Tm [(le)24(v)15(el)-253(di)24(v)15(er)18(gence)-254(contrib)20(utions.)-329(Let)]TJ/F35 9.9626 Tf 1 0 0 1 451.389 496.078 Tm [(D)]TJ/F34 6.9738 Tf 8.249 -1.494 Td [(f)]TJ/F32 9.9626 Tf 5.164 1.494 Td [(\\050)]TJ/F35 9.9626 Tf [(p)]TJ/F34 6.9738 Tf 8.887 -1.494 Td [(T)]TJ/F38 9.9626 Tf 6.276 1.494 Td [(k)]TJ/F35 9.9626 Tf [(p)]TJ/F34 6.9738 Tf 9.994 -1.494 Td [(S)]TJ/F32 9.9626 Tf 5.771 1.494 Td [(\\051)]TJ/F133 9.9626 Tf 1.02 0 0 1 502.18 496.078 Tm [(denote)-253(an)]TJ/F35 9.9626 Tf 1 0 0 1 307.44 484.123 Tm [(f)]TJ/F133 9.9626 Tf 1.02 0 0 1 313.39 484.123 Tm [(-di)24(v)15(er)18(gence.)-666(At)-365(each)-366(tok)10(en)-366(position)]TJ/F35 9.9626 Tf 1 0 0 1 468.953 484.123 Tm [(n)]TJ'
b'/F35 9.9626 Tf 441.315 439.185 Td [(p)]TJ/F34 6.9738 Tf 5.012 -1.494 Td [(T)]TJ/F32 9.9626 Tf 6.276 1.494 Td [(\\050)]TJ/F35 9.9626 Tf [(v)]TJ/F38 9.9626 Tf [-314(j)-277(\\001)]TJ/F32 9.9626 Tf [(\\051)]TJ/F25 9.9626 Tf 25.2 20.881 Td [(\\023)]TJ/F35 9.9626 Tf 8.994 -14.047 Td [(:)]TJ/F133 9.9626 Tf -179.825 -25.806 Td [(W)80(e)-250(compute)-250(the)-250(clipped)-250(di)25(v)15(er)18(gence:)]TJ/F35 9.9626 Tf 32.037 -31.24 Td [(D)]TJ/F31 6.9738 Tf 8.525 5.175 Td [(\\050)]TJ/F34 6.9738 Tf [(f)]TJ/F31 6.9738 Tf [-112(\\051)]TJ -0.277 -8.181 Td [(clip)]TJ/F32 9.9626 Tf 12.952 3.006 Td [(\\050)]TJ/F35 9.9626 Tf [(p)]TJ/F34 6.9738 Tf 8.887 -1.495 Td [(T)]TJ/F38 9.9626 Tf 6.276 1.495 Td [(k)]TJ/F35 9.9626 Tf [(p)]TJ/F34 6.9738 Tf 9.994 -1.495 Td [(S)]TJ/F32 9.9626 Tf 5.771 1.495 Td [(\\051)-278(=)]TJ 21.251 6.739 Td [(1)]TJ'
b' [(\\051,)-333(we)-316(form)-315(a)-316(sampled-)]TJ 1.02 0 0 1 307.44 310.819 Tm [(tok)10(en)-327(re)25(w)9(ard)-326(signal)-327(\\050a)-327(re)25(v)15(erse-KL)-327(signal)-327(on)-326(sampled)-327(ac-)]TJ 0.999 0 0 1 307.44 298.864 Tm [(tions\\051)-251(and)-251(optimize)-251(with)-251(polic)15(y)-251(gradient.)-313(F)15(or)-251(each)-251(position)]TJ/F35 9.9626 Tf 1 0 0 1 307.44 286.909 Tm [(n)]TJ/F133 9.9626 Tf [-250(in)-250(a)-250(sampled)-250(sequence)]TJ/F32 9.9626 Tf [-319(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F133 9.9626 Tf [-36(,)-250(de\\002ne)-250(the)-250(adv)25(antage)-250(term)]TJ/F35 9.9626 Tf 0 -20.505 Td [(A)]TJ/F34 6.9738 Tf 7.472 -1.495 Td [(n)]TJ/F32 9.9626 Tf 5.423 1.495 Td [(\\050)]TJ/F35 9.9626 Tf [(x;)]TJ/F32 9.9626 Tf [-235(^)]TJ/F35 9.9626 Tf [568(y)]TJ/F32 9.9626 Tf [-36(\\051)-277(=)-278(log)]TJ/F35 9.9626 Tf [-181(p)]TJ/F34 6.9738 Tf 55.937 -1.495 Td [(T)]TJ/F32 9.9626 Tf 6.276 1.495 Td [(\\050)-69(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F34 6.9738 Tf 8.759 -1.495 Td [(n)]TJ/F38 '
b'/F38 9.9626 Tf 453.427 207.823 Td [(j)]TJ/F32 9.9626 Tf [-69(^)]TJ/F35 9.9626 Tf [56…5555 tokens truncated…t)]TJ 1 0 0 1 55.44 708.003 Tm [(pass@8)-250(accurac)15(y)-250(on)-250(AIME25)-250(and)-250(HMMT25.)-310(Full-distrib)20(ution)-250(objecti)25(v)15(es)-250(\\050logit)-250(distillation\\051)-250(outperform)-250(sampled-tok)10(en)-250(objecti)25(v)15(es)1(.)]TJ'
b'/F127 9.9626 Tf 388.302 683.025 Td [(AIME25)-1200(HMMT25)]TJ'
b'/F133 9.9626 Tf 116.375 666.068 Td [(OPSD)-250(w/)-250(Full-v)20(ocab)20(ulary)-250(logit)-250(distillation)-250(\\050)]TJ'
b'/F127 9.9626 Tf 398.125 666.068 Td [(84.1)-3478(60.0)]TJ/F133 9.9626 Tf -281.75 -11.955 Td [(OPSD)-250(w/)-250(Sampled-tok)10(en)-250(distillation)-250(\\050)]TJ'
b" [(\\051,)-253(which)]TJ 0.98 0 0 1 55.44 527.022 Tm [(uses)-256(the)-256(same)-256(underlying)-256(model)-256(as)-256(both)-256(teacher)-256(and)-256(student)]TJ 1.02 0 0 1 55.44 515.066 Tm [(by)-269(pro)14(viding)-269(the)-269(teacher)-270(with)-269(pri)24(vile)15(ged)-269(conte)14(xt)-269(and)-269(then)]TJ 0.994 0 0 1 55.44 503.111 Tm [(SFT)-252(the)-253(student)-252(on)-252(the)-252(teacher')55(s)]TJ/F145 9.9626 Tf [-252(g)10(ener)15(ated)]TJ/F133 9.9626 Tf [-279(outputs)-253(without)]TJ 1.014 0 0 1 55.44 491.156 Tm [(conte)15(xt.)-306(This)-247(can)-247(be)-246(vie)24(wed)-247(as)]TJ/F145 9.9626 Tf [-246(of)18(f-policy)]TJ/F133 9.9626 Tf [(,)-247(where)-247(the)-246(learn-)]TJ 1.02 0 0 1 55.44 479.201 Tm [(ing)-318(signal)-318(is)-319(a)-318(discrete)-318(tok)10(en)-319(sequence.)-523(In)-318(the)-318(reasoning)]TJ 1.02 0 0 1 55.44 467.246 Tm [(domain,)-397(ReST)-367(\\050)]TJ"
b" [(\\051)-247(does)-247(on-polic)16(y)-247(sample)-247(from)]TJ 1.016 0 0 1 55.44 383.56 Tm [(student)-246(and)-246(sho)25(ws)-246(that)]TJ/F145 9.9626 Tf [-245(conte)19(xt-induced)]TJ/F133 9.9626 Tf [-272(kno)25(wledge)-246(can)-246(be)]TJ 0.997 0 0 1 55.44 371.604 Tm [(internalized)-251(via)-251(soft)-251(distillation)-251(by)-252(minimizing)-251(di)25(v)15(er)18(gences)]TJ 0.988 0 0 1 55.44 359.649 Tm [(and)-252(demonstrates)-252(this)-252(in)-253(kno)26(wledge)-253(editing)-252(settings.)-313(OPSD)]TJ 0.997 0 0 1 55.44 347.694 Tm [(dif)25(fers)-252(from)-252(these)-252(approaches)-252(in)-252(that)-251(we)-252(perform)]TJ/F145 9.9626 Tf [-252(on-policy)55(,)]TJ 1.006 0 0 1 55.44 335.739 Tm [(soft)-248(distillation)]TJ/F133 9.9626 Tf [-248(on)-247(the)-248(student')55(s)-248(o)25(wn)-248(rollouts)-248(for)-248(reasoning)]TJ 1.02 0 0 1 55.44 323.784 Tm [(tasks:)-524(the)-355(teacher')54(s)-355(supervision)-355(is)-355(per)19(-tok)10(en)-355(distrib)20(ution)]TJ 0.997 0 0 1 55.44 311.8"
b' [(\\051)-257(e)14(xplored)-257(on-polic)15(y)-257(self-)]TJ 1 0 0 1 55.44 228.142 Tm [(distillation)-250(on)-250(continual)-250(learning)-250(tasks.)]TJ/F127 9.9626 Tf 1.02 0 0 1 55.44 210.21 Tm [(On-P)20(olicy)-292(Distillation)]TJ/F133 9.9626 Tf [-292(methods)-292(train)-292(a)-291(student)-292(model)-292(di-)]TJ 1.02 0 0 1 55.44 198.254 Tm [(rectly)-310(on)-310(trajectori)1(es)-310(sampled)-310(from)-310(its)-310(o)25(wn)-310(polic)15(y)64(,)-326(while)]TJ 1.02 0 0 1 55.44 186.299 Tm [(a)-249(teacher)-248(model)-249(pro)15(vides)-249(per)20(-tok)10(en)-249(guidance)-248(through)-249(KL-)]TJ 1.02 0 0 1 55.44 174.344 Tm [(based)-294(re)15(gularization)-294(or)-295(rel)1(ated)-295(objecti)25(v)15(es)-294(\\050)]TJ'
b" [(\\051.)-380(These)-270(approaches)-270(miti-)]TJ 0.98 0 0 1 55.44 138.479 Tm [(g)5(ate)-221(distrib)20(ution)-221(shift)-222(by)-221(optimizing)-221(directly)-222(on)-221(the)-221(student')56(s)]TJ 1.02 0 0 1 55.191 126.523 Tm [(visitation)-259(distrib)20(ution,)-263(b)19(ut)-259(the)15(y)-259(typically)-259(rely)-259(on)-260(a)-259(distinct)]TJ 1.02 0 0 1 55.44 114.568 Tm [(and)-297(often)-298(lar)18(ger)-298(teacher)-297(model.)-461(In)-297(this)-298(w)10(ork,)-311(we)-297(e)15(xplore)]TJ 1.019 0 0 1 55.082 102.613 Tm [(whether)-244(an)-245(LLM)-244(can)-245(teach)-244(itself)-245(by)-244(conditioning)-245(on)-244(more)]TJ 1.02 0 0 1 55.44 90.658 Tm [(pri)24(vil)1(e)14(ged)-257(answer)-257(i)1(nformation)-257(and)-257(le)24(v)15(eraging)-257(its)-257(o)25(wn)-257(rea-)]TJ 0.991 0 0 1 55.44 78.703 Tm [(soning)-252(capability)-251(to)-252(guide)-252(a)-252(weak)10(er)-251(v)15(ersion)-252(of)-252(itself)-251(to)25(w)10(ard)]TJ"
b' [(\\051,)-253(where)-252(a)-252(human)-253(teacher)]TJ 1.007 0 0 1 307.44 582.26 Tm [(pro)15(vides)-249(correcti)25(v)14(e)-249(supervision)-249(on)-249(the)-249(states)-249(visited)-249(by)-250(the)]TJ 1 0 0 1 307.44 570.305 Tm [(student)-250(polic)15(y)65(.)]TJ/F127 9.9626 Tf 1.013 0 0 1 307.44 552.372 Tm [(Impr)18(o)10(ving)-248(LLM)-247(Reasoning)-247(thr)18(ough)-247(SFT)-248(and)-247(RL.)]TJ/F133 9.9626 Tf [-247(SFT)]TJ 1.02 0 0 1 307.44 540.417 Tm [(and)-247(RL)-246(are)-247(tw)10(o)-246(primary)-247(methods)-246(for)-247(impro)15(ving)-247(LLM)-246(rea-)]TJ 1.02 0 0 1 307.44 528.462 Tm [(soning)-321(abili)1(ty)63(.)-530(SFT)-320(on)-321(high-quality)-320(reasoning)-321(traces)-320(has)]TJ 0.994 0 0 1 307.44 516.507 Tm [(demonstrated)-251(strong)-251(performance)-252(\\050)]TJ'
b' [(\\051.)]TJ/F127 11.9552 Tf 0.578 -28.564 Td [(6.)-250(Conclusion)]TJ/F133 9.9626 Tf 1.017 0 0 1 306.972 313.151 Tm [(W)79(e)-248(introduced)-247(On-Polic)15(y)-247(Self-Distillation)-247(\\050OPSD\\051,)-247(a)-248(sim-)]TJ 0.984 0 0 1 307.44 301.196 Tm [(ple)-254(yet)-253(ef)25(fecti)25(v)16(e)-254(frame)26(w)10(ork)-254(for)-253(post-training)-254(lar)19(ge)-254(language)]TJ 1.02 0 0 1 307.44 289.241 Tm [(models)-254(on)-254(reasoning)-255(tas)1(ks.)-332(The)-254(intuition)-254(behind)-254(OPSD)-254(is)]TJ 1.02 0 0 1 307.44 277.286 Tm [(that)-257(a)-257(suf)24(\\002)1(ciently)-257(capable)-258(reasoning)-257(LLM)-257(can)-257(teach)-257(itself)]TJ 1.02 0 0 1 307.082 265.331 Tm [(when)-264(it)-265(has)-264(access)-265(to)-264(pri)25(vile)14(ged)-264(information)-264(about)-265(the)-264(an-)]TJ 0.98 0 0 1 307.44 253.376 Tm [(swer)-203(to)-203(a)-203(reasoning)-203(problem,)-213(utilizing)-203(its)-203(o)26(wn)-203(rationalization)]TJ 0.987 0 0 1 307.44 241.42 Tm [(ability)-253(to)-254(g'
---STREAM 32 10323 ---
b'/F127 11.9552 Tf 55.44 714.977 Td [(A.)-250(Limitations)-250(and)-250(Futur)18(e)-250(Dir)18(ections)]TJ/F133 9.9626 Tf 0.998 0 0 1 55.44 695.604 Tm [(Due)-250(to)-251(computational)-250(constraints,)-251(our)-250(e)15(xperiments)-250(are)-251(limited)-250(to)-250(models)-251(up)-250(to)-251(8B)-250(parameters.)-311(It)-250(remains)-250(an)-251(open)-250(question)]TJ 1.02 0 0 1 55.082 683.648 Tm [(whether)-249(this)-249(trend)-249(continues)-249(at)-249(scales)-249(be)15(yond)-249(8B)-249(parameters.)-316(Se)25(v)15(eral)-249(promising)-249(directions)-249(w)9(arrant)-249(further)-249(i)1(n)39(v)14(estig)5(ation.)]TJ 1 0 0 1 55.44 671.693 Tm [(First,)-249(our)-249(current)-248(frame)25(w)10(ork)-249(does)-249(not)-249(e)15(xplici)1(tly)-249(le)25(v)15(erage)-249(correctness)-249(v)15(eri\\002cation)-248(of)-249(generated)-249(answers;)-249(incorporating)-248(such)]TJ 1.02 0 0 1 55.44 659.738 Tm [(signals)-324(could)-324(pro)15(vide)-324(additional)-324(learning)-324(objecti)24(v)15(es)-'
b'/F145 8.9664 Tf 0.994 0 0 1 54.938 541.56 Tm [(T)93(able)-253(5.)]TJ/F127 8.9664 Tf [-315(P)20(er)37(-tok)11(en)-253(KL)-253(di)11(v)10(er)10(gence)-253(by)-252(tok)10(en)-253(category)-252(acr)18(oss)-253(generation)-252(styles.)]TJ/F133 8.9664 Tf [-315(Mean)-253(per)21(-tok)10(en)-253(KL)-252(di)25(v)15(er)18(gence)-253(brok)10(en)-252(do)25(wn)-252(by)-253(tok)10(en)]TJ 0.995 0 0 1 55.44 531.598 Tm [(cate)15(gory)-252(\\050see)-252(Appendix)]TJ'
b" [-251(for)-252(detailed)-252(de\\002nitions\\051,)-252(a)20(v)15(eraged)-252(o)15(v)15(er)-252(10)-251(problems.)-314(Thinking)-252(Mode)]TJ/F133 7.1731 Tf [-346(O)-62(FF)]TJ/F133 8.9664 Tf [-23(/)]TJ/F133 7.1731 Tf [-31(O)-62(N)]TJ/F133 8.9664 Tf [-276(indicates)-252(whether)-252(the)-252(student)-252(or)]TJ 1.003 0 0 1 55.44 521.635 Tm [(teacher)-249(LLM')55(s)-249(prompt)-249(format)-249(enables)-249(thinking)-249(mode.)-309(W)80(e)-249(\\002nd)-249(when)-249(student')55(s)-249(generation')55(s)-249(thinking)-249(mode)-249(is)-249(of)25(f)-249(and)-249(when)-249(the)-249(teacher')55(s)]TJ 1 0 0 1 55.44 511.672 Tm [(thinking)-250(mode)-250(is)-250(on,)-250(the)-250(KL)-250(signal)-250(on)-250(math)-250(related)-250(tok)10(ens)-250(are)-250(the)-250(highest.)-310(And)-250(we)-250(choose)-250(this)-250(setup)-250(for)-250(our)-250(e)15(xperiments.)]TJ"
b'/F133 9.9626 Tf 1.02 0 0 1 54.972 385.337 Tm [(W)78(e)-280(pro)15(vide)-280(the)-280(training)-280(and)-280(e)25(v)24(aluation)-280(con\\002gurations)-280(for)-280(our)-280(SFT)73(,)-280(GRPO)-280(and)-280(OPSD)-280(e)15(xperiments)-280(in)-280(T)78(ables)]TJ'
b' [(.)]TJ 0.993 0 0 1 55.44 373.382 Tm [(Note)-253(that)-253(we)-253(adopt)-254(the)-253(Thinking-Mode-of)26(f)-254(student)-253(/)-253(Thinking-Mode-on)-253(teacher)-253(con\\002guration)-253(for)-253(main)-254(OPSD)-253(e)15(xperiments.)]TJ 1.013 0 0 1 55.44 361.427 Tm [(F)15(or)-246(more)-245(e)15(xperiment)-246(details,)-245(please)-246(refer)-245(to)-245(our)-246(released)-245(training)-246(code)-245(in)]TJ'
b' [-245(https://github)39(.com/siyan-zhao/OPSD)]TJ'
b" [(.W)79(e)-245(didn')18(t)]TJ 0.982 0 0 1 55.44 349.471 Tm [(conduct)-253(tuning)-253(for)-254(the)-253(clipping)-253(parameter)]TJ/F35 9.9626 Tf 1 0 0 1 220.828 349.471 Tm [(\\034)]TJ/F133 9.9626 Tf 0.982 0 0 1 226.311 349.471 Tm [(,)-254(opti)1(mizing)-254(this)-253(h)5(yperparameter)-253(may)-253(yield)-253(further)-254(perform)1(ance)-254(g)5(ains)-253(within)-253(the)]TJ 1 0 0 1 55.44 337.516 Tm [(same)-250(100-step)-250(b)20(udget)-250(for)-250(lar)18(ger)-250(models.)]TJ"
b'/F145 8.9664 Tf 146.164 -28.121 Td [(T)92(able)-250(6.)]TJ/F133 8.9664 Tf [-310(T)35(raining)-250(Con\\002guration)-250(for)-250(GRPO)-250(and)-250(OPSD)]TJ'
b'/F127 9.9626 Tf 163.273 287.252 Td [(P)10(arameter)-11274(GRPO)-3708(OPSD)]TJ'
b'/F133 9.9626 Tf 163.273 171.603 Td [(Number)-250(of)-250(Generations)-250(per)-250(Prompt)-2848(8)-5985(1)]TJ 0 -11.955 Td [(Sampling)-250(T)70(emperature)-7514(1.2)-5235(1.1)]TJ 0 -11.955 Td [(KL)-250(Coef)25(\\002cient)-250(\\050)]TJ/F35 9.9626 Tf [(\\014)]TJ/F133 9.9626 Tf [-53(\\051)-9100(0.0)-5611(\\226)]TJ 0 -11.956 Td [(T)35(raining)-250(Steps)-10686(500)-4985(100)]TJ'
b' [(\\051)-275(optimizer)-276(and)-275(b\\003oat16)-275(precision)-276(for)-275(all)-276(training)-275(runs.)-395(F)14(or)]TJ 1 0 0 1 55.44 76.939 Tm [(OPSD,)-250(unless)-250(otherwise)-250(stated,)-250(we)-250(used)-250(full-v)20(ocab)20(ulary)-250(logit)-250(distillation.)]TJ'
---STREAM 34 9075 ---
b'/F127 11.9552 Tf 55.44 392.782 Td [(C.)-250(T)92(ok)10(en)-250(Category)-250(De\\002nitions)]TJ/F133 9.9626 Tf 0.984 0 0 1 54.972 373.408 Tm [(W)81(e)-254(cate)15(gorize)-254(tok)10(ens)-254(into)]TJ/F145 9.9626 Tf [-255(style)]TJ/F133 9.9626 Tf [-254(and)]TJ/F145 9.9626 Tf [-254(math)]TJ/F133 9.9626 Tf [-255(groups)-254(using)-254(prede\\002ned)-255(k)10(e)16(yw)10(ord)-255(lists.)-316(Thes)1(e)-255(k)10(e)16(yw)10(ord)-255(sets)-254(are)-254(used)-255(to)-254(analyze)-254(the)]TJ 1 0 0 1 55.44 361.453 Tm [(per)20(-tok)10(en)-250(KL)-250(di)25(v)15(er)18(gence)-250(stylistic)-250(tok)10(ens)-250(and)-250(mathematical)-250(kno)25(wledge)-250(tok)10(ens)-250(as)-250(in)-250(Section)]TJ'
b' [(.)]TJ/F127 9.9626 Tf 1.02 0 0 1 55.44 331.963 Tm [(Style)-313(T)90(ok)10(ens.)]TJ/F133 9.9626 Tf [-981(maybe,)-330(perhaps,)-331(probably)64(,)-330(p)-1(os)1(sibly)63(,)-330(let,)-331(okay)64(,)-330(ok,)-331(alright,)-330(hmm,)-331(w)10(ait,)-331(because,)-330(since,)-330(so,)-331(thus,)-330(hence,)]TJ 1.014 0 0 1 55.44 320.008 Tm [(therefore,)-247(b)20(ut,)-247(ho)25(we)24(v)15(er)40(,)-247(although,)-247(though,)-247(yet,)-247(or)40(,)-247(alternati)24(v)15(ely)64(,)-247(instead,)-246(otherwise,)-247(actually)64(,)-247(really)64(,)-247(just,)-247(sim)1(ply)64(,)-247(basically)64(,)]TJ 1.02 0 0 1 55.191 308.053 Tm [(v)15(ery)63(,)-309(quite,)-310(pretty)64(,)-310(rather)40(,)-310(f)10(airly)64(,)-310(no)24(w)64(,)-309(then,)-310(ne)15(xt,)-310(\\002rst,)-310(sec)1(ond)-1(,)-309(\\002nally)64(,)-310(try)64(,)-310(see,)-309(check,)-310(note,)-310(recall,)-309(think,)-310(idea,)-309(strate)14(gy)64(,)]TJ 0.999 0 0 1 55.44 296.098 Tm [(approach,)-249(method,)-249(w)10(ay)65(,)-249(w)10(o'
b" [(\\051)-251(can)-250(be)-250(interpreted)-251(as)-250(a)-251(polic)15(y-gradient)-250(update)-250(with)-251(a)]TJ/F145 9.9626 Tf [-250(dense)10(,)-251(tok)10(en-le)15(vel)]TJ/F133 9.9626 Tf [-251(re)25(w)10(ard)-251(signal)]TJ 0.982 0 0 1 55.44 170.849 Tm [(deri)25(v)16(ed)-256(from)-256(pri)26(vile)15(ged)-256(inform)1(ation.)-319(In)-256(this)-256(section,)-256(we)-255(sho)25(w:)-318(\\0501\\051)-255(OPSD)-256(can)-256(be)-255(seen)-256(as)-256(a)-255(dense-re)25(w)10(ard)-255(polic)15(y)-256(gradient,)-256(and)]TJ 1 0 0 1 55.112 158.894 Tm [(\\0502\\051)-250(we)-250(contrast)-250(OPSD)-250(with)-250(ST)80(aR,)-250(demonstrating)-250(that)-250(ST)80(aR')55(s)-250(learning)-250(signal)-250(is)]TJ/F145 9.9626 Tf [-250(sequence-le)15(vel)]TJ/F133 9.9626 Tf [-251(while)-250(OPSD)-250(is)]TJ/F145 9.9626 Tf [-250(tok)10(en-le)15(vel)]TJ/F133 9.9626 Tf [(.)]TJ/F127 9.9626 Tf 0.328 -25.134 Td [(D)20(.1.)-250(ST)92(aR)-250(as)-250(Sequence-Le)15(v)10(el)-250(P)20(olicy-Gradient)]TJ/F133 9.9626 T"
---STREAM 36 15944 ---
b'/F133 9.9626 Tf 1.011 0 0 1 55.082 714.977 Tm [(where)-247(the)-247(model)-248(\\002rst)-247(samples)-247(a)-247(latent)-248(rationale)]TJ/F35 9.9626 Tf 1 0 0 1 247.318 714.977 Tm [(r)]TJ/F133 9.9626 Tf 1.011 0 0 1 254.579 714.977 Tm [(before)-247(predicting)-247(the)-248(\\002nal)-247(answer)]TJ/F35 9.9626 Tf 1 0 0 1 392.338 714.977 Tm [(y)]TJ/F133 9.9626 Tf 1.011 0 0 1 397.58 714.977 Tm [(.)-307(Gi)25(v)15(en)-247(an)-248(indicator)-247(re)25(w)10(ard)]TJ/F35 9.9626 Tf 1 0 0 1 510.292 714.977 Tm [(R)]TJ/F32 9.9626 Tf [-8(\\050)]TJ/F35 9.9626 Tf [(y)]TJ/F32 9.9626 Tf [-36(\\051)-277(=)]TJ/F348 9.9626 Tf -455.425 -11.956 Td [(1)]TJ/F32 9.9626 Tf [(\\050)]TJ/F35 9.9626 Tf [(y)]TJ/F32 9.9626 Tf [-314(=)]TJ/F35 9.9626 Tf [-277(y)]TJ/F34 6.9738 Tf 33.371 3.616 Td [(?)]TJ/F32 9.9626 Tf 4.58 -3.616 Td [(\\051)]TJ/F133 9.9626 Tf [(,)-250(the)-250(e)15(xpected)-250(return)-250(across)-250(the)-250(dataset)]TJ/F38 9.9626 Tf [-250(S)]TJ/F32 9.9626 Tf [-353(=)]TJ/F38 9.9626 Tf [-277(f)]TJ/F32 9.9626 Tf'
b' [(\\051)-254(can)-254(also)-254(be)-254(vie)25(wed)-254(as)-254(a)-254(polic)15(y-gradient)-254(method,)-254(b)20(ut)-254(with)-254(a)-254(tok)10(en-le)25(v)15(el)-254(re)26(w)10(ard.)]TJ 1.015 0 0 1 55.44 469.311 Tm [(Fix)-245(a)-246(training)-245(pair)]TJ/F32 9.9626 Tf 1 0 0 1 130.531 469.311 Tm [(\\050)]TJ/F35 9.9626 Tf [(x;)-167(y)]TJ/F34 6.9738 Tf 19.238 3.615 Td [(?)]TJ/F32 9.9626 Tf 4.58 -3.615 Td [(\\051)]TJ/F133 9.9626 Tf 1.015 0 0 1 160.706 469.311 Tm [(and)-245(let)-246(the)-245(student)-246(generate)-245(a)-246(trajectory)]TJ/F32 9.9626 Tf 1 0 0 1 322.537 469.311 Tm [(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F38 9.9626 Tf [-314(\\030)]TJ/F35 9.9626 Tf [-278(p)]TJ/F34 6.9738 Tf 22.854 -1.495 Td [(S)]TJ/F32 9.9626 Tf 5.771 1.495 Td [(\\050)]TJ/F38 9.9626 Tf [(\\001)-278(j)]TJ/F35 9.9626 Tf [-278(x)]TJ/F32 9.9626 Tf [(\\051)]TJ/F133 9.9626 Tf 1.015 0 0 1 375.675 469.311 Tm [(.)-305(At)-246(each)-245(position)]TJ/F35 9.9626 Tf 1 0 0 1 449.963 469.311 Tm [(n)]TJ/F133 9.9626 Tf 1.015 0 0 1 455.9'
b'/F38 9.9626 Tf 336.326 284.494 Td [(j)]TJ/F32 9.9626 Tf [-69(^)]TJ/F35 9.9626 Tf [569(y)]TJ/F38 9.9626 Tf [-36(j)]TJ/F37 6.9738 Tf 16.627 20.145 Td [(j)]TJ/F31 6.9738 Tf [-84(^)]TJ/F34 6.9738 Tf [653(y)]TJ/F37 6.9738 Tf [-35(j)]TJ/F25 9.9626 Tf -2.684 -3.847 Td [(X)]TJ/F34 6.9738 Tf -0.31 -21.098 Td [(n)]TJ/F31 6.9738 Tf [(=1)]TJ/F35 9.9626 Tf 16.672 11.634 Td [(r)]TJ/F34 6.9738 Tf 4.495 -1.495 Td [(n)]TJ/F32 9.9626 Tf 5.423 1.495 Td [(\\050)]TJ/F35 9.9626 Tf [(x;)]TJ/F32 9.9626 Tf [-235(^)]TJ/F35 9.9626 Tf [568(y)]TJ/F32 9.9626 Tf [-36(\\051)]TJ/F25 9.9626 Tf 23.112 20.025 Td [(3)]TJ 0 -17.933 Td [(5)]TJ 6.642 17.933 Td [(3)]TJ 0 -17.933 Td [(5)]TJ/F35 9.9626 Tf 8.302 -2.092 Td [(:)]TJ/F133 9.9626 Tf 0.998 0 0 1 55.131 258.132 Tm [(This)-249(re)25(w)10(ard)-250(is)-249(dense:)-310(it)-250(pro)15(vides)-249(a)-250(learning)-249(signal)-250(at)-249(e)25(v)15(ery)-249(tok)10(en)-250(position,)-250(r)1(e)15(g)5(ardless)-250(of)-249(whether)-250(the)-249(\\002nal)-250(answer)-249(is)-250(corr'
---STREAM 54 28781 ---
b'[ (AIME24) ] TJ'
b'[ (AIME25) ] TJ'
b'[ (OPSD) ] TJ'
---STREAM 56 6858 ---
b'[ (A) 39.9305555556 (vg@12 AIME24 Accuracy \\(%\\)) ] TJ'
b'[ (w/o per-token KL Clipping) ] TJ'
b'[ (w/ per-token KL Clipping) ] TJ'
---STREAM 58 14412 ---
b'[ (AIME25 \\(Qwen3-1.7B\\)) ] TJ'
b'[ (AIME24 \\(Qwen3-1.7B\\)) ] TJ'
---STREAM 68 12757 ---
b'dup 12 /beta put'
b'B:\x86\xc5|\xba\xe8\xf9#\xbd\xcd\xa4\xcd\x9d\x8a\xedH:\xa9\x9c^\xfc\xfc\xfa\x0eK\xd9\xa8"\x03\xe6g\xb8\x8dO\x15%A*m\xef+\x7fP\xe3w\x12\x80d\xb6*_g\xed\xe31Y"D\xe7T72\xa2\xddj\xdeu\xfc!\xeb\xb7l\x08\xe5\x14z#\x8b\xcd\xadN\x8ek\xd1(\xea\xc6\xdc\xfa.U\xd3\x1c_\xf0\xe0\x7fW\x0f\x1d\xf6\x83\xa8\xbd\xa8\x08Ro\xfc\xa9|\xb2\xc1z+_\xa0V/\xc7#C\xb7\xe7\xa6I\x05G:\x99\x18YP\xb0F\x87T\x7f\xcb\x03w\x1a\xa0:\xcdH\xca\x9e\xfelKL~>\xce&\xd5\xf9\x90\x90/\xb6Q\x02\xa9\xbc\x98\x84\x98y\xe7\x04[\x177\x1d\xbf\x10\x04\xf5P\x08\x1b\xe9*\xe5\x11\xe3@\x86O-\x18\x1b\x83\x07=\xe2\xa3-\x8a<w"\x8d\x00\xf4!\x98\xe9\x97\xb1\xe07N\x08\xa1Q\xf9\xe2\xe3\xb9\xa9\x91\x15\x9bB\x8fb\xb3\x80\x0c\x96I\x17qj\x19%\xc5]\xae}\xc1\xa0\xc8Y\x8c\x0c\xb2\xeb\xf5-\xbeu\xbdB{S\xe2\xe8\xc1%%\x9f\xaaCwrEi\xb3\xc4Rp@\xe9\x13\xb9 -L\'d{\xb655\x17\xae\xa4\x94\x16+n\xa8\xe0\xd1\xc2\xcas\x16w\xf8 K\xe6\xb1\xf0[\xed\xb9B\x7f|\xdaT\x97\xddA\xc6i\xe27\xf9\x1b\xc1\xd1\xb2g_?['
---STREAM 72 9514 ---
b'dup 12 /beta put'
b'\x1b\'\xcd\x07[\xea\xd6/\x07d\x83z\x14\x01@\x94\x9d\xb7\x83\xc1IJ\xe25\xe1=O\x1e;n\x06\x03\xfa\t:!S,\x07:\x94\xeew\x86hd\x1boc\x98s\x06\xfc\x96\xc1E\xd1\xd5,\xfd\x0b\xe1Uw\xac\x88v\xd1\xc7\x02\xd2K2\t\x81\xc5\xf5\xda\x14\x8a\xa5\x8em\x1b\x83\xe0\xb9\x8f\xd9\xb6{\x8a\xbe\xe7\xa6\xe5\te$\xce\x94\x88\xac\xfa\xc8\xc2\x1dR\xca\xa6\xab\xc6\xde\xf2\xa3t\x14j3\x1aV\x8d}mel"@\xf6\x88\x83um\xf6\xb7h\x93\xbb\xebv\xb6\xcd\xc9\xcd\x11\xaf\xd7\xa9\x08\xb2\xf7\xd7\x8e^d)\xc5\xacEk/2\x1a\xd1I\x9cJ;\'\x99#KL\xa0\xd0\x9d\xf5\x8a\x88\xc1\xdf\xb6,\xe7\xfd\xecb <\xa4)\x14\xe2\xe0\x11\xce\xa8\xaf\\\x91)\xe1\xb9)\x82$\xbf\xeak\x81\x7f|\x1fy\xe1!\xbfB6\xf4Z<\xd2\xb0\x8d\x8e\x98\xd9\x9b\xe2\xbb\xd7g\xd4 \xe2\xfe\x10z\x04\xf3\t{\x80\xf0\x19`\xaaq\x92\xbc-\x95\x847A\xcez\x86\x0e\x02Yp\xdf\xb1m\x8cG\xa7\xf7\xd5l\xa7\xb5F'
---STREAM 140 9579 ---
b'\x86\xaa\xcf\xec|-\xb04\xafv KLA\xcb\xb41\xb8O\x1a\x9e=\xf1\x9b@\xda\x04\xd0\xbd\x08}UJ\x0b4\xbc'
---STREAM 184 10271 ---
b'T\xcdF\xe1!\xef\xb8\\\xff2\xc76\x880\xb4\xd5\xf4\xf1(\x8a\xa6\x06\xcd\xc7\xd0$^\x85\xff\xe8\xc1!\xca \x85\xd2\xcc\xb1\xc5b\xd7\xcf\xd5U\xba\x07L\xd3\x90E\xa1|\x90-\x8d~~\x03Y}-}?\xd3\x7f\xcd\x10\xcdS\xa7u\x085-\x16\xbd\xff\xe8UQ\xe2\xa8\xd8\x1d{)\x9d\xb3\xc9E\x9a\xa4\xb1\x16\x81?\xd5$\'KLoD\x0cG\x8d\xef\x8bz\x9f\xd9<\xb4\xdb\x072]\xbbN\xf8\xa9\x00\xfer{$6\xcf\xf6\x0bd\x8f \xfa7\x1a\xd8Hg_\xc8\x9d\xbf\x9e\xef\x98e\xfd\xee\x8b\xcc\xb5f\x85m8\x08=\xb4\x14\xc4\xef\x8d\x15r\xba\x12!\x8e\x02\xb3Y\x8d\xf8\xee\xa8C#\xa3\xfc\xee\x1c\xa1SNH\x19\xe5\x1f\x89/\xf7\x15\xfc\xaa\xc2\xd3D\x01\x87\xc8\xfe\x82\x96<\xcex,\x9d\xee\xfa\xbdz\x98N\xffz\xec\xb7*\xe5\x06\xfc\x16+\xd9l\xd0k<\x08\x95\x9b\xd6\t\x18\xf6F\x16\xef\xa6\xb1\xcf\x05\xban\xc6f\x05\xcf\xf6K\x0f\x95\xdd\xe9eQ\x0c<\x9e\xba\xce\x00G!$\x15\xea\xae\x87q\xa8M*\x19\xf30R5A \x1bk\x97m~\xa3\xdc\x92\x1fT\x15\xaf\xda\xafA\x07\xf5\x17\x98<\xc4\x9d\xc3+\x9d\xe3ui,\xd9\x18/\xfaS\xa0#1%\x1c\xaf\x0c\xddH`\xaa=7\xc7Ef\xad\x11\x1e\xa5E!o\xca\xb4\x19\xfe\xf5\xf0;\xcd\x0b>"\x14\xae\xb5\xdb\x02C2\x84\x12\x84L\xca\x0e4i\xbe?nRh_\x87\x91\'Cm\x1c\x81\x0bM{>\xbc\xbd{\xa9\xc9*\xf7\x88\xd1p~\x133Q\xfa\xcb=c\xb9\xde\xa8\x9a\x10'
---STREAM 190 7320 ---
b'E;8\x80\x7f\xc7\x9c)+\xc0\xa8U\xa9N_M\x87\xed`\x1fF\x1d\xd5$\x86\xb6\xa0Y\xc4\xc4\xd0\x1b-_m\xef\xd3\x14\xda\xe8\x1c\x0c\xd7ik\xd3\xeb\x1a\xab\x91[I]o|\x94i\x0eKL\xe1BL\xa9u<!\x95\xc0\xda\xf5 \x9cL\xc9\xb1Q\xbb\xcc\x08\xce\x91PU'
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python - <<'PY'\\nimport re,zlib\\nb=open('/workspace/paper.pdf','rb').read()\\ndef dec(s):\\n s=re.sub(rb'\\\\\\\\([()\\\\\\\\])',rb'\\\\1',s); s=s.replace(b'\\\\\\\\002',b'fi').replace(b'\\\\\\\\003',b'fl').replace(b'\\\\\\\\227',b'-').replace(b'\\\\\\\\226',b\\\"'\\\").replace(b'\\\\\\\\050',b'(').replace(b'\\\\\\\\051',b')').replace(b'\\\\\\\\%',b'%')\\n return s.decode('latin1','ignore')\\nfor i,m in enumerate(re.finditer(rb'stream\\\\r?\\\\n',b)):\\n e=b.find(b'endstream',m.end()); raw=b[m.end():e].rstrip(b'\\\\r\\\\n')\\n try: d=zlib.decompress(raw)\\n except: continue\\n if not (b'Tj' in d or b'TJ' in d): continue\\n chunks=[]\\n for line in d.splitlines():\\n if b'TJ' not in line and b'Tj' not in line: continue\\n parts=re.findall(rb'(?<!\\\\\\\\)\\\\((.*?)(?<!\\\\\\\\)\\\\)',line)\\n if parts: chunks.append(' '.join(dec(p) for p in parts))\\n text=' '.join(chunks)\\n if any(k in text for k in ['generation length','Generation Length','Ablation','AIME24','clipping parameter']):\\n print('\\\\n===== STREAM',i,'=====\\\\n',text[:16000])\\nPY\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":50000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.2 seconds
Output:
===== STREAM 10 =====
On-P olicy Self-Distillation f or Lar ge Language Models F igur e 1. Ov er view of On-P olicy Self-Distillation (O PSD): Gi v en a reasoning dataset S = f ( x i ; y ? i ) g N i =1 , we instantiate tw o policies from the same LLM: a student policy p S ( \001 j x ) and a teac her policy p T ( \001 j x; y ? ) . The student generates an on-polic y response ^ y \030 p S ( \001 j x ) . Both policies then e v aluate this trajectory to produce ne xt-tok en distrib utions p S ( \001 j x; ^ y <n ) and p T ( \001 j x; y ? ; ^ y <n ) at each step n . The learning objecti v e minimizes the per -tok en di v er gence D ( p T k p S ) along the student' s rollout. The di v er gence here can be forw ard KL, re v erse KL or JS D. Crucially , gradients backpropag ate only through the student' s logits, allo wing the model to self-distil. by this, we instantiate both the teacher and student policies from a single LLM. The teacher polic y is pro vided with pri vile ged information y ? , such as the ground-truth answer or a reference chain-of-thought, while the student polic y conditions only on the problem x . Concretely , the teacher polic y p T ( \001 j x; y ? ) conditions on both the problem and the pri vile ged answer , whereas the student polic y p S ( \001 j x ) observ es only the problem. W e preserv e the o n- p ol ic y t rain- ing paradigm by sampling traject ories ^ y e xclusi v ely from the student polic y , which then recei v es dense, tok en-l e v el supervision from the pri vile ged teacher polic y . W e therefore propose On-P oli cy Self-Distillation (OPSD) , a frame w ork in which a single model plays both teacher and student roles. The student samples its o wn trajectories ^ y \030 p S ( \001 j x ) ; we then compute the per -tok en di v er gence between the student and teacher distrib utions and minimize it o v er the student' s o wn rollouts. This formulation (i) uses on-polic y supervision (the student' s o wn trajectories), (ii) pro vides dense per -tok en feedback, (iii) e xploits ground- truth solutions y ? , and (i v) requires no separate teacher model. The learning process is captured by the loss L OPSD ( \022 ) = E ( x;y ? ) \030S E ^ y \030 p S ( \001j x ) j ^ y j X n =1 D \020 p T ( \001 j x; y ? ; ^ y <n ) \015 \015 \015 p S ( \001 j x ; ^ y <n ) \021 : (1) In summary , our contrib utions are as follo ws: \225 W e introduce On-Polic y Self-Distillation (OPSD), a no v el frame w ork that enables a single model to act as both teacher and student, le v eraging ground-truth answers to pro vide dense tok en-le v el supervision on student rollouts. \225 W e introduce a per -tok en pointwise KL clipping mecha- nism that stabilizes training and impro v es performance as we find stylis tic tok ens can dominate the training signal of math tok ens. \225 W e e v aluate OPSD on three competition-le v el mathemat- ical reasoning tasks, demonstrating that it matches the performance of GRPO with significantly im pro v ed tok en ef ficienc y and outperform supervised fine-tuning. \225 W e analyze the impact of dif f erent di v er gence objec- ti v es, the ef fect of student generation length, and stu- dent'teacher generation styles. 2. Backgr ound 2.1. Kno wledge Distillation f or A utor egr essi v e Lar ge Language Models Kno wledge distillation transfers kno wledge from a lar ger teacher model to a smaller student model by training the student to mimic the teacher' s beha vior ( Hinton et al. , 2015 ; Kim & Rush , 2016 ; Sanh et al. , 2019 ). The core insight is that the teacher' s soft probability distrib ution o v er classes contains richer information than hard labels alone, as it re v eals the teacher' s learned similarities be- tween classes. F or auto-re gressi v e language models, gi v en a dataset S = f ( x; y ? ) g where x denotes an input and y ? is the corre sponding reference output, both teacher p T and stu- dent p S define tok en-le v el distrib ut ions o v er v ocab ulary V . T raditional supervised distillation minimizes a di v er gence D between teacher and student distrib utions a v eraged o v er a fix ed dataset: L Supervised Distillation ( \022 ) = E ( x;y ) \030S [ D ( p T k p S )( y j x )] ; (2) where D ( p T k p S )( y j x ) = 1 j y j P j y j n =1 D ( p T ( \001j y <n ; x ) k p S ( \001j y <n ; x )) measures per -tok en discrepanc y . Ho we v er , this of f-polic y approach suf fers from distrib ution mismatch: the student encounters dif f erent partial sequences y <n during auto-re gressi v e 2
===== STREAM 18 =====
On-P olicy Self-Distillation f or Lar ge Language Models F igur e 3. T ok en Efficiency of OPSD. W e compare OPSD and GRPO on Qwen3-1.7B under t he same ef f ecti v e training batch size, reporting A vg@12 accurac y with training steps and total tok ens generated. Generation is capped at 1024 tok ens for OPSD and 16k for GRPO. At the same number of training steps, OPSD uses significantly fe wer tok e ns b ut outperforms GRPO on all benchmarks. Despite sampling more tok ens, GRPO only recei v es a binary outcome re w ard, and stagnates due to re w ard di v ersity collapse (rightmost plot): more than half of its ba tches ha v e zero re w ard standard de viation within 100 steps, yielding no gradient signal. OPSD sidesteps this disadv antage of outcome-based re w ards by learning from a dense distillation loss e v en with fe wer generated tok ens. OPSD as dense-r eward policy gradient and comparison to ST aR. The objecti v e in Equation ( 9 ) can be seen as pol- ic y gradient with dense, tok en-le v el re w ards. In Appendix Section D , we formalize this and contrast with ST aR ( Ze- likman et al . , 2022 ), a closely related method that also uses the same model to generate reasoning traces, then performs rejection sampling follo wed by SFT on correct traces. This procedure can be vie wed as polic y gradient with a sequence- le v el binary re w ard t h a t assigns identical credit to all tok ens and v anishes when samples are incorrect. In contrast, OPSD pro vides feedback at e v ery tok en position re g ardless of final- answer correctness. 4. Experiments W e conduct comprehensi v e e xperiments to answer the fol- lo wing research questions: (1) Ho w does OPSD compare to SFT and GRPO in rea- soning performance and sample ef ficienc y? (\247 4.2 ) (2) Ho w does per -tok en pointwise KL clipping in OPSD help stabilizing training? (\247 4.3.3 ) (3) What is the ef fect of generation style, generation length on performance? (\247 4.3.4 ) (4) Does full-v ocab ulary logit distillation pro vide benefits o v er sampled-tok en polic y gradient? (\247 4.3.5 ) 4.1. Experimental Setup Models and datasets. W e e xperiment with the Qwen3 ( T eam , 2025b ) model f amily at three scales: Qwen3- 1.7B, Qwen3-4B, and Qwen3-8B, using the instruc t-tuned v ersions. F or training data, we use the mathematical reason- ing subset of OpenThoughts ( Guha et al. , 2025 ), sampling up to 30K problem-solution pairs with chain-of-thought reasoning. W e e v aluate on competition-le v el mathematics benchmarks incl uding AIME 2024, AIME 2025, HMMT 2025. Baselines. W e compare ag ainst tw o methods trained on the same dataset: (1) SFT , standard supervised fine-tuning on e xpert trajectories, which can be seen as of f-polic y distilla- tion from a more po werful LLM that generated the reasoning traces; (2) GRPO ( Shao et al. , 2024 ), group relati v e polic y optimization with binary outcome re w ards v erified ag ainst ground-truth answers. The max generation length is set to 16k. Implementation details. W e fix the teacher polic y to be the initial polic y , rather than the currently updating learning polic y , as we find this helps stabili ze training and implicitly acts as re gularization to pre v ent e xcessi v e de viation from the initial polic y . W e use full-v ocab ulary logit distillation in our e xperiments. All e xperiments are conducted on A100 or H100 GPUs with LoRA ( Hu et al. , 2022 ). More e xperi- mental details are in Appendix B . 4.2. Main Results T able 2 reports results on competition-le v el mathematical reasoning benchmarks. OPSD consistently outperforms SFT and impro v es o v er the base model across all scales, match- ing or e xceeding GRPO in e v ery setting. Notably , OPSD achie v es these g ains using o nl y a single rollout per problem and con v er ges wit hin 100 steps, with each problem requir - ing only 1024 sampled tok ens, whereas GRPO requires 8 rollouts of 16k tok ens each and may e xhibit performance de gradation in later steps due to entrop y collapse-with most of re w ard standard de viations within a group being zero under this OpenThoughts dataset, yielding no learning signal and w asting sampling b udget. W e also observ e con- sistent performance de gradation under SFT across tasks and model scales when trained on the same datas et, which we attrib ute to the concise reasoning style of the ground truth solutions which has reduced reasoning lengths at test time. W e attrib ute OPSD' s tok en ef ficienc y to dense tok en-le v el supervision from the teac her distrib ution, and we h ypoth- 6
===== STREAM 20 =====
On-P olicy Self-Distillation f or Lar ge Language Models T able 2. Performance comparison on mathematical reasoning benchmarks for Qwen3 models. W e report A vg@12 under the sampling configuration recommended in the Qwen3 blog (temperature 1 : 0 , maximum generation length 38 k); full details are pro vided i n T able 8 . F or OPSD, we e v aluate checkpoints e v ery 20 steps up to 100 steps and report the best score. F or GRPO, we report the peak performance within 500 training steps, though we find GRPO performance to decrease for some tasks due to entrop y collapse in later steps. F or SFT , we train on the same number of samples as OPSD. SFT performance de grades due to fine-tuning on concise reasoning solutions and reduces generation length at test time, whereas OPSD transforms them into dense learning signal through rationaliz ation. Method AIME24 AIME25 HMMT25 A v erage Qwen3-8B Base (Instruct) 75.8 65.6 43.9 61.8 + SFT 72.3 64.2 42.9 59.8 + GRPO 76.4 68.9 46.7 64.0 + OPSD 77.8 70.8 45.8 64.8 Qwen3-4B Base (Instruct) 74.9 66.4 42.2 61.2 + SFT 70.2 62.3 43.4 58.6 + GRPO 75.6 68.1 44.4 62.7 + OPSD 76.4 68.3 46.1 63.6 Qwen3-1.7B Base (Instruct) 51.5 36.7 23.1 37.1 + SFT 48.4 36.3 22.7 35.8 + GRPO 51.1 38.3 23.7 37.7 + OPSD 57.2 43.9 29.2 43.4 esize that earlier tok ens may contrib ute more to ef fect i v e distillation as the y could represent more critical branching points in the reasoning process. As sho wn in Figure 3 , OPSD achie v es higher tok en learn- ing ef ficienc y within 100 steps of training as compared to GRPO. W ithin 100 steps, GRPO' s performance stagnates with less learning signal when the outcome re w ard within as sampling group remains the same, leading to zero gradient. These results suggest that OPSD may e xtract learning signal from the same reasoning datasets more ef fi ciently than both GRPO and SFT , while substantially reducing training time. 4.3. Ablation Studies & Discussions In this section, we conduct e xtensi v e ablations to study k e y design choices in OPSD, including (1) the di v er gence objecti v e, (2) the generation styles of the student and teacher (e.g., thinking-mode on/of f), (3) the ef fect of per -tok en KL clipping, (4) the impact of student generation le n gt h, and (5) comparison between full-v ocab ulary logit distillation with sampled-tok en distillation. 4 . 3 . 1 . E FF E C T O F D I V E R G E N C E O B J E C T I V E A k e y design choice in OPSD is the di v er gence used for per - tok en distrib ution matching between the pri vile ged teacher and the student. W e compare forw ard KL, re v erse KL, and JSD on AIME25 with Qwen3-1.7B in T able 3 . All ob- jecti v es are e v aluated under the same pointwise clipping scheme for stability . F orw ard KL consistently yields the strongest g ains, impro ving performance from 36.7 to 43.9 at step 50 and remaini ng abo v e the baseline at step 100. In contrast, re v erse KL and JSD pro vide limited or ne g a- ti v e impro v ements. W e therefore adopt forw ard KL in all remaining e xperiments. T able 3. Comparison of di v er gence objecti v es on AIME25 with Qwen3-1.7B. W e report A vg@12 at dif ferent training steps. F or - w ard KL significantly impro v es performance o v er the base model, while re v erse KL and JSD ( \014 = 0 : 5 ) sho w limited or ne g ati v e g ains. Method Base Step 50 Step 100 F orw ard KL ( KL( p T k p S ) ) 36.7 43.9 41.1 Re v erse KL ( KL( p S k p T ) ) 36.7 37.5 35.0 JSD ( \014 = 0 : 5 ) 36.7 36.9 39.0 4 . 3 . 2 . E FF E C T O F G E N E R A T I O N S T Y L E S A N D P E R - T O K E N K L C L I P P I N G Another k e y design choice in OPSD is the generation style of the student and teacher models, as it determines both which tok ens the student learns from and the style of super - vision pro vided by the teacher . Qwen3 models support tw o generation modes: Thinking Mode on (TM-on), in which the model produces self-reflecti v e chain-of-thought tok ens, and Thinki ng Mode of f (TM-of f ), in which it generates re- sponses directly . T o determine which combination yields 7
===== STREAM 22 =====
On-P olicy Self-Distillation f or Lar ge Language Models the most ef fecti v e learning signal, we analyze the forw ard KL di v er gence KL( p T k p S ) across all four student/teacher mode pairings, cate gorizing tok ens into three groups: math (numerals, operators, and mathematical k e yw ords), style (reasoning connecti v es), and other . T able 5 reports the mean per -tok en KL within each cate gory . Across all model sizes, the TM-of f student paired with a TM-on teacher yields the lar gest KL on math tok ens, in- dicating stronger supervision on mat hematically rele v ant tok ens. The reported KL v alues correspond to the e xpected di v er gence o v er the v ocab ulary at each position; as sho wn in T able 5 , this e xpectation is highly sk e wed, with stylistic tok ens contrib uting disproportionately lar ge v alues. This moti v ates our use of pointwise clipping to control such hea vy-tailed contrib utions. Empirically , this configuration achie v es the best do wnstream performance. W e therefore adopt the TM-of f student / TM-on teacher configuration. F igur e 4. Ef fect of Per -T ok en pointwise KL Clipping on Qwen3- 1.7B e v aluated on AIM E24. Clipping pre v ents performa nce col- lapse. 4 . 3 . 3 . E FF E C T O F P E R - T O K E N P O I N T W I S E C L I P P I N G As sho wn in T able 5 , stylistic tok ens can e xhibit higher KL di v er gence than math-related tok ens, causing them to dominate the training signal. W e mitig ate t his issue us- ing per -tok en pointwis e clipping. As sho wn in Figure 4 for Qwen3-1.7B, clipping stabilizes training and pre v ents performance de gradation, which is particularly important gi v en that OPSD con v er ges rapidly within a hundred steps of training. 4 . 3 . 4 . E FF E C T O F G E N E R A T I O N L E N G T H Since our objecti v e operates at t he tok en le v el (Eq. 6 ), the number of generated tok ens per sample directly determines the amount of supervision signa l a v ailable to the student. Longer sequences e xpose the student to more teacher feed- back, b ut the y also increase computational cost and may introduce noisy or uninformati v e continuations. T o study this trade-of f , we conduct an ablation on Qwen3-1.7B by v aryi ng the generation length of on-polic y sampled stu- dent responses among 1024 and 4096 tok ens and use full- F igur e 5. Ef f ect of Generation Length on Qwen3-1.7B. W e com- pare student generation length of 1024 vs 4096 on AIME25 and AIME24. v ocab ulary logit distillation. As sho wn in Figure 5 , in- creasing the generation length does not lead to consistent impro v ements across either task. W e attrib ute this t o early tok ens being more critical for learning: as the student gen- eration gro ws longer , later tok ens become increasingly pre- dictable to the teacher when conditioned on a suf ficiently long student prefix so less penalties are applied t o later to- k ens. This phenomenon is also noted in ( Lu & Lab , 2025 ). 4 . 3 . 5 . L E A R N I N G O B J E C T I V E C O M P A R I S O N : F U L L V O C A B U L A R Y L O G I T S D I S T I L L A T I O N V S . S A M P L E D - T O K E N D I S T I L L A T I O N Our objecti v e in Eq. 6 is defined as a per -tok en discrepanc y between the teacher and student distrib utions . In practice, OPSD can instantiate this objecti v e in tw o w ays. (1) Full- v ocab ulary logit distillation (as in GKD ( Ag arw al et al. , 2024 )): for each tok en position, we compute D ( p T k p S ) o v er the entire v ocab ulary via a full softmax, yiel ding a proper tok en-le v el f -di v er gence between the tw o policies. (2) Sampled-tok en adv antage policy-gradient objecti v e (as in the on-polic y distillation method of Lu & Lab ( 2025 )): we e v aluate teacher and student log-probabilities only at the tok en actually sampled by the student, ^ y n , and use the re v erse-KL term as a scalar adv antage inside a polic y- gradient-style loss. Thus, the first v ariant directly matches full tok en distrib utions, whereas the second optimizes an on- polic y RL objecti v e shaped by the teacher' s log-probabilities rather than a full-dist rib ution di v er gence. W e compare these v ariants on Qwen3-4B using a 2048-tok en generation b ud- get during distillation. T able 4 summarizes t he results. The full-v ocab ulary di v er gence objecti v e pro vides a consistent g ain o v er the sampled-tok en objecti v e. This suggests that e xposing the student to the full teacher distrib ution of f ers richer supervision than rel ying solely on per -tok en on-polic y shaping. Ho we v er , the full-v ocab ulary computation incurs higher peak memory usage due to storing v ocab ulary-sized logits at e v ery position, indicating a trade-of f between per - formance and ef ficienc y . 8
===== STREAM 24 =====
On-P olicy Self-Distillation f or Lar ge Language Models T able 4. Ablation on di v er gence computation strate gies for OPSD on Qwen3 - 4B with 2048 generation length for distillation. W e report pass@8 accurac y on AIME25 and HMMT25. Full-distrib ution objecti v es (logit distillation) outperform sampled-tok en objecti v es . Method V ariant AIME25 HMMT25 OPSD w/ Full-v ocab ulary logit distillation ( Ag arw al et al. , 2024 ) 84.1 60.0 OPSD w/ Sampled-tok en distillation ( Lu & Lab , 2025 ) 82.1 57.3 5. Related W ork LLM Self-T raining . Our w ork connects to a line of re- search sho wing that LLMs can impro v e by generating and e xploiting their o wn supervision signals ( Allen-Zhu & Li , 2020 ; Xu et al. , 2024b ; Chen et al. , 2024 ; W ang et al. , 2023 ; Sun et al. , 2023 ; Y uan et al. , 2024 ; Y ang et al. , 2024 ). Clos- est in spirit is conte xt distillation ( Snell et al. , 2022 ), which uses the same underlying model as both teacher and student by pro viding the teacher with pri vile ged conte xt and then SFT the student on the teacher' s g ener ated outputs without conte xt. This can be vie wed as of f-policy , where the learn- ing signal is a discrete tok en sequence. In the reasoning domain, ReST ( Gulcehre et al. , 2023 ) and ST aR ( Zelik- man et al. , 2022 ) similarly rely on iterati v e self-training loops-generate rationales conditioned on hints or answers, filter by re w ards or ground-truth answers, and fine-tune on successful trajectories-ag ain yielding hard distillation; Mitra & Ulukus ( 2025 ) e xtends this to soft dist illation. In- conte xt editing ( Qi et al. , 2025 ) does on-polic y sample from student and sho ws that conte xt-induced kno wledge can be internalized via soft distillation by minimizing di v er gences and demonstrates this in kno wledge editing settings. OPSD dif fers from these approaches in that we perform on-policy , soft distillation on the student' s o wn rollouts for reasoning tasks: the teacher' s supervision is per -tok en distrib ution matching rather than generating a rationale for SFT . OPSD frames reasoning impro v em ent as learning a conditional distrib ution induced jointly by the dataset' s ground-truth so- lutions and the model' s o wn reasoning ability . Concurrently , SDPO ( H \250 ubotter et al. , 2026 ) e xplored similar algorithm with en vironment feedbacks as pri villedged information and SDFT ( Shenfeld et al. , 2026 ) e xplored on-polic y self- distillation on continual learning tasks. On-P olicy Distillation methods train a student model di- rectly on trajectori es sampled from its o wn polic y , while a teacher model pro vides per -tok en guidance through KL- based re gularization or rel ated objecti v es ( Ag a rw al et al. , 2024 ; Xu et al. , 2024a ; Gu et al. , 2024 ; Lu & Lab , 2025 ; Xiaomi , 2026 ; Y ang et al. , 2025 ). These approaches miti- g ate distrib ution shift by optimizing directly on the student' s visitation distrib ution, b ut the y typically rely on a distinct and often lar ger teacher model. In this w ork, we e xplore whether an LLM can teach itself by conditioning on more pri vil e ged answer i nformation and le v eraging its o wn rea- soning capability to guide a weak er v ersion of itself to w ard impro v ed reasoning. On-polic y training paradigms are also widely used in robotics and deep reinforcement learning, such as D Agger ( Ross et al. , 2011 ), where a human teacher pro vides correcti v e supervision on the states visited by the student polic y . Impr o ving LLM Reasoning thr ough SFT and RL. SFT and RL are tw o primary methods for impro ving LLM rea- soning abili ty . SFT on high-quality reasoning traces has demonstrated strong performance ( Y u et al. , 2023 ; LI et al. , 2024 ; P aster et al. , 2023 ; T eam , 2025a ; Y e et al. , 2025 ; Muennighof f et al. , 2025 ; Zhou et al. , 2023 ). Ho we v er , prior w ork sho ws that SFT can rely on memorization rather than rob ust generalization ( Chu e t al. , 2025 ). In contrast, RL optimizes directly for outcome-based objecti v es can e x- hibit better generalization ( Huan e t al. , 2025 ). More recent algorithms such as GRPO ( Guo et al. , 2025 ; Shao et al. , 2024 ) enable scalable RL by estimating adv antages from group-le v el re w a rds wit hout requi ring an e xplicit critic as in PPO ( Schulman et al. , 2017 ). Building on this line of w ork, a gro wing body of research highlights the ef fecti v eness of RL VR for reasoning tasks ( Y u et al. , 2025 ; Liu et al. , 2025 ; Y ue et al. , 2025 ; An et al. , 2025 ; Zheng et al. , 2025 ). 6. Conclusion W e introduced On-Polic y Self-Distillation (OPSD), a sim- ple yet ef fecti v e frame w ork for post-training lar ge language models on reasoning tas ks. The intuition behind OPSD is that a suf fi ciently capable reasoning LLM can teach itself when it has access to pri vile ged information about the an- swer to a reasoning problem, utilizing its o wn rationalization ability to grade its weak er self without access to the ground truth. W e e xperimentally demonstrated that OPSD achie v es better performance than of f-polic y distillation/SFT , and per - forms on par with or better than GRPO, while e xhibiting significantly better sample ef ficienc y than GRPO. 7. Impact Statement This paper presents w ork whose goal is to adv ance the field of machine learning. Our method impro v es the ef fi cienc y of training language models for reasoning tasks, reducing computational costs compared to e xisting reinforcement learning approaches. W e do not foresee specific ne g at i v e societal consequences. 9
===== STREAM 32 =====
On-P olicy Self-Distillation f or Lar ge Language Models A. Limitations and Futur e Dir ections Due to computational constraints, our e xperiments are limited to models up to 8B parameters. It remains an open question whether this trend continues at scales be yond 8B parameters. Se v eral promising directions w arrant further i n v estig ation. First, our current frame w ork does not e xplici tly le v erage correctness v erification of generated answers; incorporating such signals could pro vide additional learning objecti v es be yond distrib ution matching. Finally , problem dif ficulty plays a crucial role in self-distillation: if reasoning problems e xceed the model' s comprehension threshold, the teacher polic y cannot pro vide meaningful supervision e v en with access to ground-truth solutions. This suggests that curriculum learning strate gies-gradually increasing problem dif ficulty as the model impro v es-could enhance training ef f ecti v eness. Exploring adapti v e curricula that maintain problems at the front ier of model capabilities represents an important direction for scaling OPSD to more challenging reasoning tasks. B. Experimental Details T able 5. P er -tok en KL di v er gence by tok en category acr oss generation styles. Mean per -tok en KL di v er gence brok en do wn by tok en cate gory (see Appendix C for detailed definitions), a v eraged o v er 10 problems. Thinking Mode O FF / O N indicates whether the student or teacher LLM' s prompt format enables thinking mode. W e find when student' s generation' s thinking mode is of f and when the teacher' s thinking mode is on, the KL signal on math related tok ens are the highest. And we choose this setup for our e xperiments. Qwen3-1.7B Qwen3-4B Qwen3-8B Student T eacher Style Math Other Style Math Other Style Math Other TM-of f TM-of f 0.68 0.12 0.11 0.61 0.06 0.10 0.56 0.05 0.11 TM-on TM-of f 0.51 0.10 0.17 0.41 0.05 0.18 0.33 0.05 0.15 TM-on TM-on 0.51 0.09 0.08 0.50 0.04 0.09 0.42 0.04 0.08 TM-off TM-on 0.85 0.14 0.25 0.92 0.10 0.29 0.79 0.06 0.25 W e pro vide the training and e v aluation configurations for our SFT , GRPO and OPSD e xperiments in T ables 7 , 6 and 8 . Note that we adopt the Thinking-Mode-of f student / Thinking-Mode-on teacher configuration for main OPSD e xperiments. F or more e xperiment details, please refer to our released training code in https://github .com/siyan-zhao/OPSD .W e didn' t conduct tuning for the clipping parameter \034 , opti mizing this h yperparameter may yield further perform ance g ains within the same 100-step b udget for lar ger models. T able 6. T raining Configuration for GRPO and OPSD P arameter GRPO OPSD Learning Rate 5 fi 10 \000 6 5 fi 10 \000 6 Ef fecti v e Batch Size 32 32 LoRA Rank ( r ) 64 64 LoRA Alpha ( \013 ) 128 128 LoRA T ar get Modules q proj, k proj, v proj, o proj, g ate proj, up proj, do wn proj Max Completion Length 16,000 1024 Number of Generations per Prompt 8 1 Sampling T emperature 1.2 1.1 KL Coef ficient ( \014 ) 0.0 ' T raining Steps 500 100 All e xperiments were conducted using 8 A100 or H100 GPUs with gradient checkpointing and Flash Attention 2 for memory ef fi cienc y . W e use the AdamW ( Loshchilo v & Hutter , 2017 ) optimizer and bfloat16 precision for all training runs. F or OPSD, unless otherwise stated, we used full-v ocab ulary logit distillation. 13
===== STREAM 54 =====
0 25 50 75 100 Gradient Update Steps 48 50 52 54 56 58 A vg@12 Accuracy (%) AIME24 0 25 50 75 100 Gradient Update Steps 34 36 38 40 42 44 46 A vg@12 Accuracy (%) AIME25 0 25 50 75 100 Gradient Update Steps 22 24 26 28 30 32 A vg@12 Accuracy (%) HMMT25 0 10 20 T o k e n s G e n e r a t e d ( × 1 0 6 ) 37 38 39 40 41 42 A vg@12 Accuracy (%) A verage 10 30 50 70 90 Gradient Update Steps 0.0 0.2 0.4 0.6 0.8 1.0 Fraction of Batches Zero Reward Std Frac. of GRPO GRPO OPSD
===== STREAM 56 =====
0 25 50 75 100 Gradient Update Steps 48 50 52 54 56 58 A vg@12 AIME24 Accuracy (%) w/o per-token KL Clipping w/ per-token KL Clipping
===== STREAM 58 =====
0 25 50 75 100 Gradient Update Steps 34 36 38 40 42 44 46 A vg@12 Accuracy (%) AIME25 (Qwen3-1.7B) Gen Length 4096 Gen Length 1024 0 25 50 75 100 Gradient Update Steps 50 52 54 56 58 60 A vg@12 Accuracy (%) AIME24 (Qwen3-1.7B) Gen Length 4096 Gen Length 1024
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"command -v convert || true; command -v magick || true; command -v pdftoppm || true; command -v chromium || command -v google-chrome || true; cat /workspace/.timer/remaining_secs\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":2000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
20798
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python - <<'PY'\\nimport re,zlib\\nb=open('/workspace/paper.pdf','rb').read(); streams=[]\\nfor m in re.finditer(rb'stream\\\\r?\\\\n',b):\\n e=b.find(b'endstream',m.end()); raw=b[m.end():e].rstrip(b'\\\\r\\\\n')\\n try: streams.append(zlib.decompress(raw))\\n except: streams.append(b'')\\nd=streams[58]\\nprint(d.decode('latin1','ignore')[:20000])\\nPY\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":22000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
/DeviceRGB CS
/DeviceRGB cs
1 j
1 g 0 j 0 w 1 G 1 g
0 0 m
392.711558 0 l
392.711558 178.548125 l
0 178.548125 l
h
f
/A1 gs 0 G
37.23125 34.66875 m
192.79187 34.66875 l
192.79187 156.98875 l
37.23125 156.98875 l
h
f
q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
44.302187 34.66875 m
44.302187 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
44.302187 34.66875 m
44.302187 31.16875 l
B
1 w
q
1 0 -0 1 41.7553122652 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (0) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
79.656874 34.66875 m
79.656874 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
79.656874 34.66875 m
79.656874 31.16875 l
B
1 w
q
1 0 -0 1 74.5631235913 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (25) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
115.01156 34.66875 m
115.01156 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
115.01156 34.66875 m
115.01156 31.16875 l
B
1 w
q
1 0 -0 1 109.9178099174 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (50) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
150.366246 34.66875 m
150.366246 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
150.366246 34.66875 m
150.366246 31.16875 l
B
1 w
q
1 0 -0 1 145.2724962434 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (75) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
185.720933 34.66875 m
185.720933 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
185.720933 34.66875 m
185.720933 31.16875 l
B
1 w
q
1 0 -0 1 178.0803075695 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (100) ] TJ
ET
Q
q
1 0 -0 1 62.5584349174 9.075 cm
BT
/F1 9 Tf
0 0 Td
[ (Gradient Update Steps) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 34.66875 m
192.79187 34.66875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 34.66875 m
33.73125 34.66875 l
B
1 w
q
1 0 -0 1 20.04375 31.6296875 cm
BT
/F1 8 Tf
0 0 Td
[ (34) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 55.055417 m
192.79187 55.055417 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 55.055417 m
33.73125 55.055417 l
B
1 w
q
1 0 -0 1 20.04375 52.0163541667 cm
BT
/F1 8 Tf
0 0 Td
[ (36) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 75.442083 m
192.79187 75.442083 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 75.442083 m
33.73125 75.442083 l
B
1 w
q
1 0 -0 1 20.04375 72.4030208333 cm
BT
/F1 8 Tf
0 0 Td
[ (38) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 95.82875 m
192.79187 95.82875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 95.82875 m
33.73125 95.82875 l
B
1 w
q
1 0 -0 1 20.04375 92.7896875 cm
BT
/F1 8 Tf
0 0 Td
[ (40) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 116.215417 m
192.79187 116.215417 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 116.215417 m
33.73125 116.215417 l
B
1 w
q
1 0 -0 1 20.04375 113.1763541667 cm
BT
/F1 8 Tf
0 0 Td
[ (42) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 136.602083 m
192.79187 136.602083 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 136.602083 m
33.73125 136.602083 l
B
1 w
q
1 0 -0 1 20.04375 133.5630208333 cm
BT
/F1 8 Tf
0 0 Td
[ (44) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
37.23125 156.98875 m
192.79187 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
37.23125 156.98875 m
33.73125 156.98875 l
B
1 w
q
1 0 -0 1 20.04375 153.9496875 cm
BT
/F1 8 Tf
0 0 Td
[ (46) ] TJ
ET
Q
q
0 1 -1 0 14.04375 45.375625 cm
BT
/F1 9 Tf
0 0 Td
[ (A) 39.9305555556 (vg@12 Accuracy \(%\)) ] TJ
ET
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A3 gs 1 j 1.5 w
[ 5.55 2.4 ] 0 d 0.8392156863 0.1529411765 0.1568627451 RG /DeviceRGB cs
44.302187 62.19075 m
79.656874 87.674083 l
115.01156 89.71275 l
150.366246 124.370083 l
185.720933 113.157417 l
S
0.8392156863 0.1529411765 0.1568627451 rg 1 w [ ] 0 d 0.8392156863
0.1529411765 0.1568627451 rg
q
1 0 0 1 44.3021872652 62.19075 cm /M0 Do
1 0 0 1 35.3546863261 25.4833333333 cm /M0 Do
1 0 0 1 35.3546863261 2.0386666667 cm /M0 Do
1 0 0 1 35.3546863261 34.6573333333 cm /M0 Do
1 0 0 1 35.3546863261 -11.2126666667 cm /M0 Do
Q
Q q 37.23125 34.66875 155.5606198347 122.32 re W n /A3 gs 2 J 1 j 1.5 w
0.1215686275 0.4666666667 0.7058823529 RG /DeviceRGB cs
44.302187 62.19075 m
79.656874 121.312083 l
115.01156 135.58275 l
150.366246 101.94475 l
185.720933 107.041417 l
S
0 J 0.1215686275 0.4666666667 0.7058823529 rg 0 j 1 w 0.1215686275
0.4666666667 0.7058823529 rg
q
1 0 0 1 44.3021872652 62.19075 cm /M1 Do
1 0 0 1 35.3546863261 59.1213333333 cm /M1 Do
1 0 0 1 35.3546863261 14.2706666667 cm /M1 Do
1 0 0 1 35.3546863261 -33.638 cm /M1 Do
1 0 0 1 35.3546863261 5.0966666667 cm /M1 Do
Q
Q q /A3 gs 2 J 0.8 w 0 G /DeviceRGB cs
37.23125 34.66875 m
37.23125 156.98875 l
S
192.79187 34.66875 m
192.79187 156.98875 l
S
37.23125 34.66875 m
192.79187 34.66875 l
S
37.23125 156.98875 m
192.79187 156.98875 l
S
0 J 0 g 1 j 1 w 0 g
q
1 0 -0 1 53.4334349174 162.98875 cm
BT
/F1 10.8 Tf
0 0 Td
[ (AIME25 \(Qwen3-1.7B\)) ] TJ
ET
Q
/A4 gs 1 g 0 j 0.8 G 1 g
91.138745 38.66875 m
187.19187 38.66875 l
188.258537 38.66875 188.79187 39.202083 188.79187 40.26875 c
188.79187 63.1875 l
188.79187 64.254167 188.258537 64.7875 187.19187 64.7875 c
91.138745 64.7875 l
90.072078 64.7875 89.538745 64.254167 89.538745 63.1875 c
89.538745 40.26875 l
89.538745 39.202083 90.072078 38.66875 91.138745 38.66875 c
h
B
/A3 gs 1 j 1.5 w [ 5.55 2.4 ] 0 d 0.8392156863 0.1529411765 0.1568627451
RG /DeviceRGB cs
92.738745 58.309375 m
100.738745 58.309375 l
108.738745 58.309375 l
S
0.8392156863 0.1529411765 0.1568627451 rg 1 w [ ] 0 d 0.8392156863
0.1529411765 0.1568627451 rg
100.738745 55.809375 m
101.401753 55.809375 102.037695 56.072791 102.506512 56.541608 c
102.975329 57.010425 103.238745 57.646367 103.238745 58.309375 c
103.238745 58.972383 102.975329 59.608325 102.506512 60.077142 c
102.037695 60.545959 101.401753 60.809375 100.738745 60.809375 c
100.075737 60.809375 99.439795 60.545959 98.970978 60.077142 c
98.502161 59.608325 98.238745 58.972383 98.238745 58.309375 c
98.238745 57.646367 98.502161 57.010425 98.970978 56.541608 c
99.439795 56.072791 100.075737 55.809375 100.738745 55.809375 c
h
B
0 g 0 G 0 g
q
1 0 -0 1 115.1387448347 55.509375 cm
BT
/F1 8 Tf
0 0 Td
[ (Gen Length 4096) ] TJ
ET
Q
2 J 1.5 w 0.1215686275 0.4666666667 0.7058823529 RG /DeviceRGB cs
92.738745 46.45 m
100.738745 46.45 l
108.738745 46.45 l
S
0 J 0.1215686275 0.4666666667 0.7058823529 rg 0 j 1 w 0.1215686275
0.4666666667 0.7058823529 rg
98.238745 43.95 m
103.238745 43.95 l
103.238745 48.95 l
98.238745 48.95 l
h
B
0 g 1 j 0 G 0 g
q
1 0 -0 1 115.1387448347 43.65 cm
BT
/F1 8 Tf
0 0 Td
[ (Gen Length 1024) ] TJ
ET
Q
/A1 gs 1 g 0 j 0 w 0 G 1 g
229.38125 34.66875 m
384.94187 34.66875 l
384.94187 156.98875 l
229.38125 156.98875 l
h
f
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
236.452187 34.66875 m
236.452187 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
236.452187 34.66875 m
236.452187 31.16875 l
B
1 w
q
1 0 -0 1 233.9053122652 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (0) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
271.806874 34.66875 m
271.806874 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
271.806874 34.66875 m
271.806874 31.16875 l
B
1 w
q
1 0 -0 1 266.7131235913 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (25) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
307.16156 34.66875 m
307.16156 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
307.16156 34.66875 m
307.16156 31.16875 l
B
1 w
q
1 0 -0 1 302.0678099174 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (50) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
342.516246 34.66875 m
342.516246 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
342.516246 34.66875 m
342.516246 31.16875 l
B
1 w
q
1 0 -0 1 337.4224962434 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (75) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
377.870933 34.66875 m
377.870933 156.98875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
377.870933 34.66875 m
377.870933 31.16875 l
B
1 w
q
1 0 -0 1 370.2303075695 21.590625 cm
BT
/F1 8 Tf
0 0 Td
[ (100) ] TJ
ET
Q
q
1 0 -0 1 254.7084349174 9.075 cm
BT
/F1 9 Tf
0 0 Td
[ (Gradient Update Steps) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 44.862083 m
384.94187 44.862083 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 44.862083 m
225.88125 44.862083 l
B
1 w
q
1 0 -0 1 212.19375 41.8230208333 cm
BT
/F1 8 Tf
0 0 Td
[ (50) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 65.24875 m
384.94187 65.24875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 65.24875 m
225.88125 65.24875 l
B
1 w
q
1 0 -0 1 212.19375 62.2096875 cm
BT
/F1 8 Tf
0 0 Td
[ (52) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 85.635417 m
384.94187 85.635417 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 85.635417 m
225.88125 85.635417 l
B
1 w
q
1 0 -0 1 212.19375 82.5963541667 cm
BT
/F1 8 Tf
0 0 Td
[ (54) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 106.022083 m
384.94187 106.022083 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 106.022083 m
225.88125 106.022083 l
B
1 w
q
1 0 -0 1 212.19375 102.9830208333 cm
BT
/F1 8 Tf
0 0 Td
[ (56) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 126.40875 m
384.94187 126.40875 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 126.40875 m
225.88125 126.40875 l
B
1 w
q
1 0 -0 1 212.19375 123.3696875 cm
BT
/F1 8 Tf
0 0 Td
[ (58) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A2 gs 1 j 0.5 w
[ 0.5 0.825 ] 0 d 0.6901960784 G /DeviceRGB cs
229.38125 146.795417 m
384.94187 146.795417 l
S
Q q /A3 gs 0 g 1 j 0.8 w 0 G 0 g
229.38125 146.795417 m
225.88125 146.795417 l
B
1 w
q
1 0 -0 1 212.19375 143.7563541667 cm
BT
/F1 8 Tf
0 0 Td
[ (60) ] TJ
ET
Q
q
0 1 -1 0 206.19375 45.375625 cm
BT
/F1 9 Tf
0 0 Td
[ (A) 39.9305555556 (vg@12 Accuracy \(%\)) ] TJ
ET
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A3 gs 1 j 1.5 w
[ 5.55 2.4 ] 0 d 0.8392156863 0.1529411765 0.1568627451 RG /DeviceRGB cs
236.452187 60.152083 m
271.806874 78.500083 l
307.16156 132.52475 l
342.516246 103.983417 l
377.870933 101.94475 l
S
0.8392156863 0.1529411765 0.1568627451 rg 1 w [ ] 0 d 0.8392156863
0.1529411765 0.1568627451 rg
q
1 0 0 1 236.4521872652 60.1520833333 cm /M2 Do
1 0 0 1 35.3546863261 18.348 cm /M2 Do
1 0 0 1 35.3546863261 54.0246666667 cm /M2 Do
1 0 0 1 35.3546863261 -28.5413333333 cm /M2 Do
1 0 0 1 35.3546863261 -2.0386666667 cm /M2 Do
Q
Q q 229.38125 34.66875 155.5606198347 122.32 re W n /A3 gs 2 J 1 j 1.5 w
0.1215686275 0.4666666667 0.7058823529 RG /DeviceRGB cs
236.452187 60.152083 m
271.806874 59.13275 l
307.16156 73.403417 l
342.516246 89.71275 l
377.870933 118.254083 l
S
0 J 0.1215686275 0.4666666667 0.7058823529 rg 0 j 1 w 0.1215686275
0.4666666667 0.7058823529 rg
q
1 0 0 1 236.4521872652 60.1520833333 cm /M3 Do
1 0 0 1 35.3546863261 -1.0193333333 cm /M3 Do
1 0 0 1 35.3546863261 14.2706666667 cm /M3 Do
1 0 0 1 35.3546863261 16.3093333333 cm /M3 Do
1 0 0 1 35.3546863261 28.5413333333 cm /M3 Do
Q
Q q /A3 gs 2 J 0.8 w 0 G /DeviceRGB cs
229.38125 34.66875 m
229.38125 156.98875 l
S
384.94187 34.66875 m
384.94187 156.98875 l
S
229.38125 34.66875 m
384.94187 34.66875 l
S
229.38125 156.98875 m
384.94187 156.98875 l
S
0 J 0 g 1 j 1 w 0 g
q
1 0 -0 1 245.5834349174 162.98875 cm
BT
/F1 10.8 Tf
0 0 Td
[ (AIME24 \(Qwen3-1.7B\)) ] TJ
ET
Q
/A4 gs 1 g 0 j 0.8 G 1 g
283.288745 38.66875 m
379.34187 38.66875 l
380.408537 38.66875 380.94187 39.202083 380.94187 40.26875 c
380.94187 63.1875 l
380.94187 64.254167 380.408537 64.7875 379.34187 64.7875 c
283.288745 64.7875 l
282.222078 64.7875 281.688745 64.254167 281.688745 63.1875 c
281.688745 40.26875 l
281.688745 39.202083 282.222078 38.66875 283.288745 38.66875 c
h
B
/A3 gs 1 j 1.5 w [ 5.55 2.4 ] 0 d 0.8392156863 0.1529411765 0.1568627451
RG /DeviceRGB cs
284.888745 58.309375 m
292.888745 58.309375 l
300.888745 58.309375 l
S
0.8392156863 0.1529411765 0.1568627451 rg 1 w [ ] 0 d 0.8392156863
0.1529411765 0.1568627451 rg
292.888745 55.809375 m
293.551753 55.809375 294.187695 56.072791 294.656512 56.541608 c
295.125329 57.010425 295.388745 57.646367 295.388745 58.309375 c
295.388745 58.972383 295.125329 59.608325 294.656512 60.077142 c
294.187695 60.545959 293.551753 60.809375 292.888745 60.809375 c
292.225737 60.809375 291.589795 60.545959 291.120978 60.077142 c
290.652161 59.608325 290.388745 58.972383 290.388745 58.309375 c
290.388745 57.646367 290.652161 57.010425 291.120978 56.541608 c
291.589795 56.072791 292.225737 55.809375 292.888745 55.809375 c
h
B
0 g 0 G 0 g
q
1 0 -0 1 307.2887448347 55.509375 cm
BT
/F1 8 Tf
0 0 Td
[ (Gen Length 4096) ] TJ
ET
Q
2 J 1.5 w 0.1215686275 0.4666666667 0.7058823529 RG /DeviceRGB cs
284.888745 46.45 m
292.888745 46.45 l
300.888745 46.45 l
S
0 J 0.1215686275 0.4666666667 0.7058823529 rg 0 j 1 w 0.1215686275
0.4666666667 0.7058823529 rg
290.388745 43.95 m
295.388745 43.95 l
295.388745 48.95 l
290.388745 48.95 l
h
B
0 g 1 j 0 G 0 g
q
1 0 -0 1 307.2887448347 43.65 cm
BT
/F1 8 Tf
0 0 Td
[ (Gen Length 1024) ] TJ
ET
Q
Q
The paper’s own plotted ablation confirms 4,096 tokens peak early but fall to about 55.6 at step 100, while 1,024 rises to 57.2; later-token supervision is not uniformly useful. I’ll treat 2,048 as a diagnostic rather than assume “more is better.” The more novel opportunity supported by the paper is a difficulty-aware contest curriculum, which I’ll test after isolating this run’s result.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:600});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 30
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "30",
"yield_time_ms": 60000,
"max_tokens": 600
}wait result
Script completed
Wall time 42.2 seconds
Output:
Warning: truncated output (original token count: 663)
Total output lines: 1
{"chunk_id":"21dff6","wall_time_seconds":60.001636593,"session_id":80127,"original_token_count":2940,"output":"Warning: truncated output (original token count: 2940)\nTotal output lines: 204\n\nvLLM generation done - elapsed time: 6.47s, prompts: 4, total tokens: 4931, avg length: 1232.8, speed: 762.7 tok/s\nvLLM generation done - elapsed time: 2.15s, prompts: 4, total tokens: 1744, avg length: 436.0, speed: 810.0 tok/s\nvLLM generation done - elapsed time: 4.47s, prompts: 4, total tokens: 3528, avg length: 882.0, speed: 788.4 tok/s\nvLLM generation done - elapsed time: 4.71s, prompts: 4, total tokens: 3828, avg length: 957.0, speed: 813.4 tok/s\nvLLM generation done - elapsed time: 4.84s, prompts: 4, total tokens: 3620, avg length: 905.0, speed: 747.4 tok/s\n\r 9%|▉ | 9/100 [02:05<21:07, 13.93s/it]vLLM generation done - elapsed time: 1.39s, prompts: 4, total tokens: 1159, avg length: 289.8, speed: 834.7 tok/s\nvLLM generation done - elapsed time: 3.96s, prompts: 4, total tokens: 4222, avg length: 1055.5, speed: 1066.4 tok/s\nvLLM generation done - elapsed time: 6.47s, prompts: 4, total tokens: 5702, avg length: 1425.5, speed: 880.8 tok/s\nvLLM generation done - elapsed …63 tokens truncated… total tokens: 4701, avg length: 1175.2, speed: 909.1 tok/s\nvLLM generation done - elapsed time: 6.49s, prompts: 4, total tokens: 4877, avg length: 1219.2, speed: 752.0 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n /workspace/runs/contest2048/capped_seed42/generations/generations_step_15.json\n================================================================================\n\n\r 16%|█▌ | 16/100 [03:47<20:26, 14.60s/it]\r \r{'loss': 0.0018, 'grad_norm': 0.06540733575820923, 'learning_rate': 5e-06, 'on_policy_loss': 0.0018, 'epoch': 0.01}\n\r 16%|█▌ | 16/100 [03:48<20:26, 14.60s/it]vLLM generation done - elapsed time: 4.54s, prompts: 4, total tokens: 3945, avg length: 986.2, speed: 869.6 tok/s\nvLLM generation done - elapsed time: 6.00s, prompts: 4, total tokens: 3164, avg length: 791.0, speed: 527.0 tok/s\nvLLM generation done - elapsed time: 6.19s, prompts: 4, total tokens: 4006, avg length: 1001.5, speed: 646.8 tok/s\nvLLM generation done - elapsed time: 6.23s, prompts: 4, total tokens: 3739, avg length: 934.8, speed: 599.8 tok/s\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:500});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 31
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "31",
"yield_time_ms": 60000,
"max_tokens": 500
}wait result
Script completed
Wall time 46.4 seconds
Output:
Warning: truncated output (original token count: 562)
Total output lines: 1
{"chunk_id":"b8d2b6","wall_time_seconds":60.002767984,"session_id":80127,"original_token_count":2583,"output":"Warning: truncated output (original token count: 2583)\nTotal output lines: 239\n\nvLLM generation done - elapsed time: 4.88s, prompts: 4, total tokens: 3445, avg length: 861.2, speed: 706.0 tok/s\nvLLM generation done - elapsed time: 6.23s, prompts: 4, total tokens: 4424, avg length: 1106.0, speed: 709.7 tok/s\nvLLM generation done - elapsed time: 6.49s, prompts: 4, total tokens: 5581, avg length: 1395.2, speed: 860.2 tok/s\nvLLM generation done - elapsed time: 6.64s, prompts: 4, total tokens: 6275, avg length: 1568.8, speed: 944.3 tok/s\n\r 17%|█▋ | 17/100 [04:02<20:13, 14.62s/it]vLLM generation done - elapsed time: 4.80s, prompts: 4, total tokens: 4278, avg length: 1069.5, speed: 891.8 tok/s\nvLLM generation done - elapsed time: 6.09s, prompts: 4, total tokens: 3841, avg length: 960.2, speed: 630.4 tok/s\nvLLM generation done - elapsed time: 6.22s, prompts: 4, to…62 tokens truncated…ime: 5.89s, prompts: 4, total tokens: 4316, avg length: 1079.0, speed: 732.4 tok/s\nvLLM generation done - elapsed time: 6.43s, prompts: 4, total tokens: 5094, avg length: 1273.5, speed: 792.6 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n /workspace/runs/contest2048/capped_seed42/generations/generations_step_20.json\n================================================================================\n\n\r 21%|██ | 21/100 [04:57<18:52, 14.34s/it]vLLM generation done - elapsed time: 3.13s, prompts: 4, total tokens: 2848, avg length: 712.0, speed: 908.5 tok/s\nvLLM generation done - elapsed time: 5.56s, prompts: 4, total tokens: 3317, avg length: 829.2, speed: 596.8 tok/s\nvLLM generation done - elapsed time: 6.15s, prompts: 4, total tokens: 3307, avg length: 826.8, speed: 537.8 tok/s\nvLLM generation done - elapsed time: 6.34s, prompts: 4, total tokens: 4608, avg length: 1152.0, speed: 726.5 tok/s\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:80127,chars:\"\",yield_time_ms:60000,max_output_tokens:400});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 32
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "32",
"yield_time_ms": 60000,
"max_tokens": 400
}wait result
Script completed
Wall time 45.3 seconds
Output:
Warning: truncated output (original token count: 461)
Total output lines: 1
{"chunk_id":"6a5725","wall_time_seconds":60.000925646,"session_id":80127,"original_token_count":1367,"output":"Warning: truncated output (original token count: 1367)\nTotal output lines: 46\n\nvLLM generation done - elapsed time: 4.93s, prompts: 4, total tokens: 3938, avg length: 984.5, speed: 799.4 tok/s\nvLLM generation done - elapsed time: 5.46s, prompts: 4, total tokens: 4681, avg length: 1170.2, speed: 857.4 tok/s\nvLLM generation done - elapsed time: 6.64s, prompts: 4, total tokens: 5694, avg length: 1423.5, speed: 857.8 tok/s\nvLLM generation done - elapsed time: 6.70s, prompts: 4, total tokens: 5555, avg length: 1388.8, speed: 829.4 tok/s\n\r 22%|██▏ | 22/100 [05:12<18:48, 14.47s/it]\r \r{'loss': -0.0023, 'grad_norm': 0.043…61 tokens truncated…nvLLM generation done - elapsed time: 6.71s, prompts: 4, total tokens: 5236, avg length: 1309.0, speed: 780.1 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n /workspace/runs/contest2048/capped_seed42/generations/generations_step_25.json\n================================================================================\n\n\r 26%|██▌ | 26/100 [06:13<18:38, 15.12s/it]\r \r{'loss': -0.0028, 'grad_norm': 0.05375239625573158, 'learning_rate': 5e-06, 'on_policy_loss': -0.0028, 'epoch': 0.02}\n\r 26%|██▌ | 26/100 [06:14<18:38, 15.12s/it]vLLM generation done - elapsed time: 3.82s, prompts: 4, total tokens: 3260, avg length: 815.0, speed: 854.1 tok/s\n"}