Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
596
duration
7h 59m
recorded
2026-08-03

run idopsd-aime24-4gpus__kuc95M7

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol669 turns73 narration596 reasoning596 action596 observation1 final
#201
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 103
Wall time 11.0 seconds
Output:
#202
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "103",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 43.8 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"e639ee","wall_time_seconds":60.00177412,"session_id":89905,"original_token_count":4258,"output":"Warning: truncated output (original token count: 4258)\nTotal output lines: 342\n\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3578, avg length: 894.5, speed: 1094.8 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3728, avg length: 932.0, speed: 1139.8 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3764,…58 tokens truncated… | 12/100 [01:33<11:19,  7.73s/it]vLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 3079, avg length: 769.8, speed: 966.7 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3746, avg length: 936.5, speed: 1148.4 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3449, avg length: 862.2, speed: 1055.2 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1239.7 tok/s\n"}
#203
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 104
Wall time 11.0 seconds
Output:
#204
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "104",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 43.2 seconds
Output:
Warning: truncated output (original token count: 309)
Total output lines: 1

{"chunk_id":"b3fc33","wall_time_seconds":60.00113627,"session_id":89905,"original_token_count":3008,"output":"Warning: truncated output (original token count: 3008)\nTotal output lines: 170\n\nvLLM generation done - elapsed time: 3.09s, prompts: 4, total tokens: 2576, avg length: 644.0, speed: 834.0 tok/s\nvLLM generation done - elapsed time: 3.12s, prompts: 4, total tokens: 2388, avg length: 597.0, speed: 764.5 tok/s\nvLLM generation done - elapsed time: 3.20s, prompts: 4, total tokens: 2662, a…59 tokens truncated…vg length: 861.0, speed: 1026.8 tok/s\nvLLM generation done - elapsed time: 3.46s, prompts: 4, total tokens: 2783, avg length: 695.8, speed: 805.0 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_linear/capped_seed42/generations/generations_step_20.json\n================================================================================\n\n\r 21%|██        | 21/100 [02:44<10:32,  8.01s/it]"}
#205
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 105
Wall time 11.0 seconds
Output:
#206
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "105",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.2 seconds
Output:
Warning: truncated output (original token count: 309)
Total output lines: 1

{"chunk_id":"d1647d","wall_time_seconds":60.001167559,"session_id":89905,"original_token_count":2547,"output":"Warning: truncated output (original token count: 2547)\nTotal output lines: 84\n\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3027, avg length: 756.8, speed: 911.9 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 3735, avg length: 933.8, speed: 1110.8 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 3981, …59 tokens truncated…: 3.36s, prompts: 4, total tokens: 3968, avg length: 992.0, speed: 1181.6 tok/s\nvLLM generation done - elapsed time: 3.40s, prompts: 4, total tokens: 3719, avg length: 929.8, speed: 1094.0 tok/s\n\r 30%|███       | 30/100 [03:57<09:21,  8.02s/it]\r                                                \r{'loss': -0.0001, 'grad_norm': 0.056843049824237823, 'learning_rate': 3.5500000000000003e-06, 'on_policy_loss': -0.0001, 'epoch': 0.03}\n\r 30%|███       | 30/100 [03:57<09:21,  8.02s/it]"}
#207
narration

Linear decay is confirmed in the live log (3.55e−6 near step 30 versus 5e−6 constant previously). The loss is less negative at the same stage, consistent with avoiding the increasingly aggressive late updates seen under the constant schedule.

#208
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 106
Wall time 11.0 seconds
Output:
#209
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "106",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 42.1 seconds
Output:
Warning: truncated output (original token count: 309)
Total output lines: 1

{"chunk_id":"0909fa","wall_time_seconds":60.001281185,"session_id":89905,"original_token_count":2332,"output":"Warning: truncated output (original token count: 2332)\nTotal output lines: 81\n\nvLLM generation done - elapsed time: 3.23s, prompts: 4, total tokens: 2951, avg length: 737.8, speed: 914.0 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4029, avg length: 1007.2, speed: 1188.8 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096,…59 tokens truncated…30s, prompts: 4, total tokens: 3641, avg length: 910.2, speed: 1104.0 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 4060, avg length: 1015.0, speed: 1216.0 tok/s\n\r 38%|███▊      | 38/100 [05:01<08:12,  7.95s/it]\r                                                \r{'loss': -0.0019, 'grad_norm': 0.05336622893810272, 'learning_rate': 3.1500000000000003e-06, 'on_policy_loss': -0.0019, 'epoch': 0.04}\n\r 38%|███▊      | 38/100 [05:01<08:12,  7.95s/it]"}
#210
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 107
Wall time 11.0 seconds
Output:
#211
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "107",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 45.1 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"0cf69d","wall_time_seconds":60.002018763,"session_id":89905,"original_token_count":2560,"output":"Warning: truncated output (original token count: 2560)\nTotal output lines: 88\n\nvLLM generation done - elapsed time: 3.22s, prompts: 4, total tokens: 3394, avg length: 848.5, speed: 1053.2 tok/s\nvLLM generation done - elapsed time: 3.28s, prompts: 4, total tokens: 3443, avg length: 860.8, speed: 1050.5 tok/s\nvLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3233,…58 tokens truncated…gth: 924.0, speed: 1102.3 tok/s\nvLLM generation done - elapsed time: 3.13s, prompts: 4, total tokens: 2510, avg length: 627.5, speed: 801.9 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1235.3 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1231.0 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3892, avg length: 973.0, speed: 1169.4 tok/s\n"}
#212
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 108
Wall time 11.0 seconds
Output:
#213
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "108",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 42.7 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"6bd4fc","wall_time_seconds":60.002582333,"session_id":89905,"original_token_count":3199,"output":"Warning: truncated output (original token count: 3199)\nTotal output lines: 246\n\n\r 47%|████▋     | 47/100 [06:14<07:10,  8.11s/it]vLLM generation done - elapsed time: 3.19s, prompts: 4, total tokens: 2705, avg length: 676.2, speed: 849.1 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3329, avg length: 832.2, speed: 1018.1 tok/s\nvLLM generation…58 tokens truncated…   | 55/100 [07:18<06:02,  8.06s/it]vLLM generation done - elapsed time: 3.20s, prompts: 4, total tokens: 2695, avg length: 673.8, speed: 842.3 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 2841, avg length: 710.2, speed: 863.7 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3436, avg length: 859.0, speed: 1032.4 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 3860, avg length: 965.0, speed: 1139.4 tok/s\n"}
#214
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 109
Wall time 11.0 seconds
Output:
#215
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "109",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 43.9 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"73eff9","wall_time_seconds":60.001227791,"session_id":89905,"original_token_count":2645,"output":"Warning: truncated output (original token count: 2645)\nTotal output lines: 89\n\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 2962, avg length: 740.5, speed: 901.6 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3092, avg length: 773.0, speed: 936.0 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, a…58 tokens truncated…  | 64/100 [08:30<04:47,  7.99s/it]vLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3147, avg length: 786.8, speed: 961.9 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3474, avg length: 868.5, speed: 1042.6 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3372, avg length: 843.0, speed: 1014.1 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3583, avg length: 895.8, speed: 1070.1 tok/s\n"}
#216
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 110
Wall time 11.0 seconds
Output:
#217
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "110",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 42.2 seconds
Output:
Warning: truncated output (original token count: 308)
Total output lines: 1

{"chunk_id":"01921f","wall_time_seconds":60.001439721,"session_id":89905,"original_token_count":2712,"output":"Warning: truncated output (original token count: 2712)\nTotal output lines: 92\n\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3704, avg length: 926.0, speed: 1109.3 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3768, avg length: 942.0, speed: 1124.6 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096,…58 tokens truncated…gth: 889.0, speed: 1070.0 tok/s\nvLLM generation done - elapsed time: 3.32s, prompts: 4, total tokens: 3708, avg length: 927.0, speed: 1117.3 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3819, avg length: 954.8, speed: 1143.2 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1208.3 tok/s\nvLLM generation done - elapsed time: 3.42s, prompts: 4, total tokens: 3996, avg length: 999.0, speed: 1168.6 tok/s\n"}
#218
narration

The linear run is around 70% complete and remains smooth. Because the scheduler only changes optimization dynamics—not data, teacher, rollout length, or loss definition—the upcoming evaluation is a clean comparison against the 56.11 constant-schedule result.

#219
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 111
Wall time 11.0 seconds
Output:
#220
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "111",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 44.9 seconds
Output:
Warning: truncated output (original token count: 309)
Total output lines: 1

{"chunk_id":"7641f8","wall_time_seconds":60.001547653,"session_id":89905,"original_token_count":2439,"output":"Warning: truncated output (original token count: 2439)\nTotal output lines: 82\n\n\r 74%|███████▍  | 74/100 [09:50<03:27,  7.98s/it]\r                                                \r{'loss': -0.006, 'grad_norm': 0.04503677412867546, 'learning_rate': 1.3500000000000002e-06, 'on_policy_loss': -0.006, 'epoch': 0.08}\n\r 74%|███████▍  | 74/100 [09:50<03:27,…59 tokens truncated…tal tokens: 3433, avg length: 858.2, speed: 1037.2 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3412, avg length: 853.0, speed: 1023.7 tok/s\n\r 82%|████████▏ | 82/100 [10:54<02:23,  7.95s/it]\r                                                \r{'loss': -0.0075, 'grad_norm': 0.053328584879636765, 'learning_rate': 9.500000000000001e-07, 'on_policy_loss': -0.0075, 'epoch': 0.09}\n\r 82%|████████▏ | 82/100 [10:54<02:23,  7.95s/it]"}
#221
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 112
Wall time 11.0 seconds
Output:
#222
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "112",
  "yield_time_ms": 60000,
  "max_tokens": 250
}
wait result
Script completed
Wall time 40.4 seconds
Output:
Warning: truncated output (original token count: 309)
Total output lines: 1

{"chunk_id":"df1bec","wall_time_seconds":60.001394362,"session_id":89905,"original_token_count":3924,"output":"Warning: truncated output (original token count: 3924)\nTotal output lines: 405\n\nvLLM generation done - elapsed time: 3.22s, prompts: 4, total tokens: 2882, avg length: 720.5, speed: 896.4 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3468, avg length: 867.0, speed: 1053.5 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096,…59 tokens truncated… prompts: 4, total tokens: 3496, avg length: 874.0, speed: 1047.3 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4020, avg length: 1005.0, speed: 1197.3 tok/s\n\r 90%|█████████ | 90/100 [11:57<01:19,  7.95s/it]\r                                                \r{'loss': -0.0072, 'grad_norm': 0.05461088940501213, 'learning_rate': 5.5e-07, 'on_policy_loss': -0.0072, 'epoch': 0.1}\n\r 90%|█████████ | 90/100 [11:57<01:19,  7.95s/it]"}
#223
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:60000,max_output_tokens:800});\ntext(JSON.stringify(r));\n"
}
exec result
Script running with cell ID 113
Wall time 11.0 seconds
Output:
#224
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "113",
  "yield_time_ms": 60000,
  "max_tokens": 800
}
wait result
Script completed
Wall time 40.9 seconds
Output:
Warning: truncated output (original token count: 865)
Total output lines: 1

{"chunk_id":"f96c44","wall_time_seconds":60.001496785,"session_id":89905,"original_token_count":3398,"output":"Warning: truncated output (original token count: 3398)\nTotal output lines: 242\n\nvLLM generation done - elapsed time: 3.25s, prompts: 4, total tokens: 2948, avg length: 737.0, speed: 906.0 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3282, avg length: 820.5, speed: 993.1 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3677, avg length: 919.2, speed: 1103.3 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1218.6 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 3091, avg length: 772.8, speed: 948.3 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3607, avg length: 901.8, speed: 1076.9 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3867, avg length: 966.8, speed: 1152.8 tok/s\nvLLM generation done - elapsed time: 3.39s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1209.2 tok/s\n\n================================================================================\nSaved 40 generation outputs to:\n  /workspace/runs/earlycredit_linear/capped_seed42/generations/generations_step_90.json\n================================================================================\n\n\r 91%|█████████ | 91/100 [12:05<01:11,  7.97s/it]vLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3019, avg length: 754.8, speed: 912.5 tok/s\nvLLM g…65 tokens truncated…vg length: 953.8, speed: 1144.7 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1223.9 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3807, avg length: 951.8, speed: 1129.5 tok/s\n\r 98%|█████████▊| 98/100 [13:01<00:15,  7.97s/it]\r                                                \r{'loss': -0.0079, 'grad_norm': 0.05285836011171341, 'learning_rate': 1.5000000000000002e-07, 'on_policy_loss': -0.0079, 'epoch': 0.11}\n\r 98%|█████████▊| 98/100 [13:01<00:15,  7.97s/it]vLLM generation done - elapsed time: 3.31s, prompts: 4, total tokens: 3536, avg length: 884.0, speed: 1066.9 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3663, avg length: 915.8, speed: 1096.7 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3692, avg length: 923.0, speed: 1101.7 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 3963, avg length: 990.8, speed: 1176.3 tok/s\nvLLM generation done - elapsed time: 3.30s, prompts: 4, total tokens: 3150, avg length: 787.5, speed: 954.2 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3574, avg length: 893.5, speed: 1071.8 tok/s\nvLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3798, avg length: 949.5, speed: 1138.3 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1216.6 tok/s\n\r 99%|█████████▉| 99/100 [13:09<00:07,  7.98s/it]"}
#225
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:89905,chars:\"\",yield_time_ms:30000,max_output_tokens:800});\ntext(JSON.stringify(r));\n"
}
exec result
Script completed
Wall time 10.8 seconds
Output:
{"chunk_id":"b34d76","wall_time_seconds":10.817909539,"exit_code":0,"original_token_count":686,"output":"vLLM generation done - elapsed time: 3.34s, prompts: 4, total tokens: 3636, avg length: 909.0, speed: 1089.2 tok/s\nvLLM generation done - elapsed time: 3.35s, prompts: 4, total tokens: 3709, avg length: 927.2, speed: 1106.9 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1215.9 tok/s\nvLLM generation done - elapsed time: 3.37s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1215.3 tok/s\nvLLM generation done - elapsed time: 3.21s, prompts: 4, total tokens: 2634, avg length: 658.5, speed: 820.7 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3182, avg length: 795.5, speed: 971.7 tok/s\nvLLM generation done - elapsed time: 3.29s, prompts: 4, total tokens: 3259, avg length: 814.8, speed: 990.0 tok/s\nvLLM generation done - elapsed time: 3.36s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1220.0 tok/s\n\r100%|██████████| 100/100 [13:17<00:00,  7.98s/it]\r                                                 \r{'loss': -0.0084, 'grad_norm': 0.06052081659436226, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': -0.0084, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:17<00:00,  7.98s/it]\r                                                 \r{'train_runtime': 800.9326, 'train_samples_per_second': 3.995, 'train_steps_per_second': 0.125, 'train_loss': -0.0028223273158073424, 'epoch': 0.11}\n\r100%|██████████| 100/100 [13:20<00:00,  7.98s/it]\r100%|██████████| 100/100 [13:20<00:00,  8.01s/it]\n[rank0]:[W803 06:19:55.260838312 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank2]:[W803 06:19:56.617538867 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank1]:[W803 06:19:56.070159300 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\n[rank3]:[W803 06:19:56.161389618 ProcessGroupNCCL.cpp:1538] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())\nTRAIN_CKPT=/workspace/runs/earlycredit_linear/capped_seed42/checkpoint-100\n"}