gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗The 0.02 run completed, but late rollout accuracy weakened and the optimization regime stayed much more negative than v1. I’m still running the full benchmark because training diagnostics are not sufficient; this candidate must exceed 56% on the fixed evaluation to earn a replication.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"bash /workspace/dev_eval.sh /workspace/runs/frontier_v4_clip002/capped_seed42/checkpoint-100 /workspace/eval_results/frontier_v4_clip002_full.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":250,\"tty\":true});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"e683b5","wall_time_seconds":1.001353446,"session_id":4422,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:180});\ntext(r.output);"
}exec result
Script running with cell ID 271
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "271",
"yield_time_ms": 50000,
"max_tokens": 180
}wait result
Script completed
Wall time 29.0 seconds
Output:
Warning: truncated output (original token count: 207)
Total output lines: 10
Warning: truncated output (original token count: 5926)
Total output lines: 146
Checkpoint kind: lora
Merging LoRA adapter into base weights for evaluation (CPU, once)...
`torch_dtype` is deprecated! Use `dtype` instead!
Loading checkpoint shards: 0%| | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards: 100%|███████…27 tokens truncated…s,INFO 08-03 08:22:45 [llm.py:306] Supported_tasks: ['generate']
Adding requests: 0%| | 0/8 [00:00<?, ?it/s]
Adding requests: 100%|██████████████████████████| 8/8 [00:00<00:00, 226.03it/s]
Processed prompts: 0%| | 0/96 [00:00<?, ?it/s, est. speed input: 0.00 toks/s,encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 272
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "272",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 29.6 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 273
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "273",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 28.3 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 274
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "274",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 29.7 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 275
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "275",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 29.5 seconds
Output:
Processed prompts: 12%|▏| 12/96 [03:23<23:45, 16.96s/it, est. speed input: 9.8
Processed prompts: 25%|▎| 24/96 [03:48<09:51, 8.22s/it, est. speed input: 16.The strong-clipping evaluation is generating normally. I’ll apply the predeclared >56% replication threshold; anything below it reverts jsd_token_clip to the replicated 0.05 setting before final cleanup.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 276
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "276",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 32.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 277
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "277",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 31.8 seconds
Output:
Processed prompts: 12%|▏| 12/96 [05:26<38:03, 27.19s/it, est. speed input: 5.4
Processed prompts: 14%|▏| 12/84 [05:52<35:17, 29.41s/it, est. speed input: 2.6encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 278
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "278",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 27.7 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 279
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "279",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 21.4 seconds
Output:
Processed prompts: 29%|▎| 24/84 [07:38<17:16, 17.27s/it, est. speed input: 3.9
Processed prompts: 38%|▍| 36/96 [08:08<14:20, 14.34s/it, est. speed input: 9.6encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 280
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "280",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 26.9 seconds
Output:
Processed prompts: 25%|▎| 24/96 [09:10<26:36, 22.18s/it, est. speed input: 5.7encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:4422,chars:\"\",yield_time_ms:50000,max_output_tokens:150});\ntext(r.output);"
}exec result
Script running with cell ID 281
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "281",
"yield_time_ms": 50000,
"max_tokens": 150
}wait result
Script completed
Wall time 32.2 seconds
Output: