gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}exec result
Script running with cell ID 135
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "135",
"yield_time_ms": 50000,
"max_tokens": 200
}wait result
Script completed
Wall time 25.7 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 10
Warning: truncated output (original token count: 4202)
Total output lines: 299
vLLM generation done - elapsed time: 3.65s, prompts: 4, total tokens: 3255, avg length: 813.8, speed: 890.8 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3615, avg length: 903.8, speed: 978.5 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3501, avg length: 87…27 tokens truncated…okens: 3340, avg length: 835.0, speed: 914.0 tok/s
vLLM generation done - elapsed time: 3.67s, prompts: 4, total tokens: 3429, avg length: 857.2, speed: 934.7 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3816, avg length: 954.0, speed: 1033.2 tok/s
vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 4028, avg length: 1007.0, speed: 1087.8 tok/s
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}exec result
Script running with cell ID 136
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "136",
"yield_time_ms": 50000,
"max_tokens": 500
}wait result
Script completed
Wall time 33.7 seconds
Output:
Warning: truncated output (original token count: 527)
Total output lines: 15
Warning: truncated output (original token count: 2071)
Total output lines: 64
vLLM generation done - elapsed time: 3.64s, prompts: 4, total tokens: 3356, avg length: 839.0, speed: 921.9 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3723, avg length: 930.8, speed: 1009.7 tok/s
vLLM generation done - elapsed time: 3.72s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1100.6 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1097.4 tok/s
94%|█████████▍| 94/100 [13:58<00:52, 8.78s/it]
{'loss': -0.0012, 'grad_norm': 0.054118622094392776, 'learning_rate': 3.5000000000000004e-07, 'on_policy_loss': -0.0012, 'rollout_accuracy': 0.1875, 'epoch': 0.34}
94%|█████████▍| 94/100 [13:58<00:52, 8.78s/it]vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 3387, avg length: 846.8, speed: 9…27 tokens truncated… prompts: 4, total tokens: 3584, avg length: 896.0, speed: 978.6 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3573, avg length: 893.2, speed: 967.5 tok/s
vLLM generation done - elapsed time: 3.73s, prompts: 4, total tokens: 4096, avg length: 1024.0, speed: 1098.5 tok/s
100%|██████████| 100/100 [14:51<00:00, 8.80s/it]
{'loss': -0.0058, 'grad_norm': 0.06921099871397018, 'learning_rate': 5.0000000000000004e-08, 'on_policy_loss': -0.0058, 'rollout_accuracy': 0.1875, 'epoch': 0.36}
100%|██████████| 100/100 [14:51<00:00, 8.80s/it]
{'train_runtime': 894.9095, 'train_samples_per_second': 3.576, 'train_steps_per_second': 0.112, 'train_loss': 0.001304372102022171, 'epoch': 0.36}
100%|██████████| 100/100 [14:54<00:00, 8.80s/it]
100%|██████████| 100/100 [14:54<00:00, 8.95s/it]
v2 completed all 100 steps in 14:54. I’m launching its full official-shape evaluation now. If it clears the current two-run v1 mean of 55.14 by a useful margin, I’ll repeat v2; otherwise I’ll revert to the already-positive v1 method.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/frontier_v2/capped_seed42/checkpoint-100 /workspace/eval_results/frontier_v2_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":400,\"tty\":true});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"3daf82","wall_time_seconds":1.001777257,"session_id":1051,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 138
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "138",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 34.8 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 14
Warning: truncated output (original token count: 5926)
Total output lines: 146
Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...
`torch_dtype` is deprecated! Use `dtype` instead!
Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards: 100%|█████████████████| 2/2 [00:00<00:00, 49.77it/s]
Generating with data_parallel_size=4 (TP=1 per engine) ...
INFO 08-03 06:18:32 [__init__.py:216] Automatically detected platform cuda.
INFO 08-03 06:18:32 [__init__.py:216] Autom…27 tokens truncated… | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 8/8 [00:00<00:00, 259.37it/s]
Processed prompts: 0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,INFO 08-03 06:18:52 [llm.py:306] Supported_tasks: ['generate']
Adding requests: 0%| | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 8/8 [00:00<00:00, 255.28it/s]
Processed prompts: 0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 139
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "139",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 32.7 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 140
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "140",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 26.2 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 141
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "141",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 27.9 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 142
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "142",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 32.3 seconds
Output:
Processed prompts: 12%|▏| 12/96 [03:37<25:23, 18.14s/it, est. speed input: 8.6encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 143
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "143",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 34.7 seconds
Output:
v2 evaluation is in the long-trace phase. I’m holding all selection decisions until the final 360-sample score; the run remains healthy and identical to official inference settings.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 144
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "144",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 34.2 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 145
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "145",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 32.9 seconds
Output:
Processed prompts: 12%|▏| 12/96 [06:12<43:30, 31.07s/it, est. speed input: 4.7encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:1051,chars:\"\",yield_time_ms:50000,max_output_tokens:250});\ntext(r.output);"
}exec result
Script running with cell ID 146
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "146",
"yield_time_ms": 50000,
"max_tokens": 250
}wait result
Script completed
Wall time 29.2 seconds
Output:
Processed prompts: 14%|▏| 12/84 [07:05<42:30, 35.42s/it, est. speed input: 2.0
Processed prompts: 25%|▎| 24/96 [07:14<21:44, 18.11s/it, est. speed input: 8.9