gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 636
- duration
- 8h 10m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__3tyLi4e
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 98
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "98",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 35.6 seconds
Output:
Repeat-two full evaluation is now generating. The method and seed are identical to the 56.39 run; only unavoidable vLLM/CUDA training nondeterminism differs. A second score near or above 56 would be strong evidence that the method—not a lucky checkpoint—is driving the gain.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 99
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "99",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 35.1 seconds
Output:
Processed prompts: 12%|▏| 12/96 [03:24<23:53, 17.06s/it, est. speed input: 9.1encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 100
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "100",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 35.4 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 101
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "101",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 34.6 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 102
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "102",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 32.4 seconds
Output:
Processed prompts: 12%|▏| 12/96 [05:58<41:48, 29.86s/it, est. speed input: 4.9
Processed prompts: 25%|▎| 24/96 [06:09<18:07, 15.11s/it, est. speed input: 10.
Processed prompts: 14%|▏| 12/84 [06:27<38:44, 32.29s/it, est. speed input: 2.2encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 103
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "103",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 34.8 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 104
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "104",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 33.6 seconds
Output:
Processed prompts: 29%|▎| 24/84 [08:09<18:17, 18.29s/it, est. speed input: 3.8encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 105
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "105",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 35.5 seconds
Output:
Processed prompts: 43%|▍| 36/84 [08:28<08:32, 10.68s/it, est. speed input: 5.5
Processed prompts: 25%|▎| 24/96 [08:44<24:31, 20.44s/it, est. speed input: 6.0encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 106
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "106",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 32.9 seconds
Output:
Processed prompts: 38%|▍| 36/96 [09:24<15:37, 15.63s/it, est. speed input: 8.3
Processed prompts: 50%|▌| 48/96 [09:50<08:15, 10.32s/it, est. speed input: 9.7The repeat evaluation is progressing normally. While it runs, the main refinement I’m considering is leaving early-token weighting on the dense KL but making the verified correct-trajectory term uniform across the full response, so final answer/boxing tokens are not discounted. I’ll only spend a run on that change if repeat two confirms the base candidate.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 107
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "107",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 34.9 seconds
Output:
Processed prompts: 57%|▌| 48/84 [10:32<06:19, 10.53s/it, est. speed input: 7.4
Processed prompts: 71%|▋| 60/84 [10:55<02:58, 7.43s/it, est. speed input: 9.1encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 108
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "108",
"yield_time_ms": 50000,
"max_tokens": 300
}wait result
Script completed
Wall time 29.0 seconds
Output:
Processed prompts: 14%|▏| 12/84 [11:39<1:09:56, 58.28s/it, est. speed input: 2
Processed prompts: 86%|▊| 72/84 [11:39<01:13, 6.14s/it, est. speed input: 10.encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}exec result
Script running with cell ID 109
Wall time 11.0 seconds
Output: