gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 627
- duration
- 6h 30m
- recorded
- 2026-08-02
run idopsd-aime24-4gpus__RqGcDty
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "261",
"yield_time_ms": 30000,
"max_tokens": 150
}wait result
Script completed
Wall time 10.7 seconds
Output:
Warning: truncated output (original token count: 177)
Total output lines: 10
Warning: truncated output (original token count: 5933)
Total output lines: 154
`torch_dtype` is deprecated! Use `dtype` instead!
Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...
Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s]
Loading check…27 tokens truncated…:48:15 [llm.py:306] Supported_tasks: ['generate']
Adding requests: 0%| | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████| 8/8 [00:00<00:00, 322.14it/s]
Processed prompts: 0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 262
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "262",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 12.6 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 263
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "263",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 13.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 264
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "264",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 10.6 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 265
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "265",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 13.3 seconds
Output:
The replicated full evaluation is healthy but currently on long-generation problems; as before, shard progress arrives in 12-sample increments. I’m waiting for the merged 360-sample score.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 266
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "266",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 10.8 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 267
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "267",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 12.9 seconds
Output:
Processed prompts: 12%|█▎ | 12/96 [03:20<23:24, 16.73s/it, est. speed input: 9.33 toks/s, output: 434.98 toks/s]
Processed prompts: 25%|██▌ | 24/96 [03:32<08:57, 7.46s/it, est. speed input: 18.24 toks/s, output: 852.20 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 268
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "268",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 13.9 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 269
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "269",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 13.9 seconds
Output:
Processed prompts: 14%|█▍ | 12/84 [05:01<30:08, 25.12s/it, est. speed input: 3.34 toks/s, output: 449.96 toks/s]encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 270
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "270",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 13.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 271
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "271",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 10.1 seconds
Output:
Processed prompts: 14%|█▍ | 12/84 [05:51<35:10, 29.31s/it, est. speed input: 2.63 toks/s, output: 370.48 toks/s]
Processed prompts: 38%|███▊ | 36/96 [06:10<10:04, 10.07s/it, est. speed input: 12.76 toks/s, output: 861.15 toks/s]All four shards are now making progress. The repeat’s generation-length profile is similar to the first validated run, so there is no obvious behavioral shift before scoring.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:50469,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}exec result
Script running with cell ID 272
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "272",
"yield_time_ms": 30000,
"max_tokens": 300
}wait result
Script completed
Wall time 10.2 seconds
Output:
Processed prompts: 12%|█▎ | 12/96 [06:30<45:35, 32.57s/it, est. speed input: 4.54 toks/s, output: 310.04 toks/s]