Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#351
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "183",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 27.9 seconds
Output:
#352
narration

The second v2 full evaluation is underway. This is the key replication: v2 currently stands at 55.56 on one run, while v1 averages 55.14 over two; I’ll retain the method with the stronger replicated mean and then perform final integrity/code checks.

#353
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 184
Wall time 11.0 seconds
Output:
#354
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "184",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 23.9 seconds
Output:

Processed prompts:  12%|▏| 12/96 [03:50<26:51, 19.18s/it, est. speed input: 8.7
#355
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 185
Wall time 11.0 seconds
Output:
#356
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "185",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 32.0 seconds
Output:

Processed prompts:  14%|▏| 12/84 [05:23<32:23, 27.00s/it, est. speed input: 2.8
#357
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 186
Wall time 11.0 seconds
Output:
#358
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "186",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.6 seconds
Output:

Processed prompts:  12%|▏| 12/96 [05:25<37:58, 27.12s/it, est. speed input: 5.4
#359
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 187
Wall time 11.0 seconds
Output:
#360
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "187",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 29.1 seconds
Output:

Processed prompts:  25%|▎| 24/96 [06:43<19:39, 16.38s/it, est. speed input: 9.6
#361
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 188
Wall time 11.0 seconds
Output:
#362
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "188",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 25.0 seconds
Output:
#363
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 189
Wall time 11.0 seconds
Output:
#364
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "189",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 25.3 seconds
Output:

Processed prompts:  29%|▎| 24/84 [08:33<20:24, 20.41s/it, est. speed input: 3.7
Processed prompts:  25%|▎| 24/96 [08:54<25:43, 21.43s/it, est. speed input: 5.9
#365
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 190
Wall time 11.0 seconds
Output:
#366
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "190",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 30.3 seconds
Output:

Processed prompts:  43%|▍| 36/84 [09:23<10:23, 12.98s/it, est. speed input: 4.9
Processed prompts:  38%|▍| 36/96 [10:01<16:26, 16.44s/it, est. speed input: 8.2
#367
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 191
Wall time 11.0 seconds
Output:
#368
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "191",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 32.8 seconds
Output:

Processed prompts:  57%|▌| 48/84 [10:17<05:46,  9.63s/it, est. speed input: 6.5
#369
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 192
Wall time 11.0 seconds
Output:
#370
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "192",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 23.8 seconds
Output:

Processed prompts:  50%|▌| 48/96 [11:12<09:50, 12.31s/it, est. speed input: 9.4
Processed prompts:  71%|▋| 60/84 [11:28<03:18,  8.29s/it, est. speed input: 8.7
Processed prompts:  86%|▊| 72/84 [11:41<01:10,  5.85s/it, est. speed input: 10.
Processed prompts:  62%|▋| 60/96 [12:08<05:43,  9.54s/it, est. speed input: 10.
#371
narration

The repeat evaluation is around halfway overall, with no errors. Once it completes, I’ll compute both variant means and per-problem consistency, then clean only the generated preprocessing caches from this session so the frozen dataset tree is returned to its original state.

#372
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 193
Wall time 11.0 seconds
Output:
#373
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "193",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 32.5 seconds
Output:

Processed prompts:  14%|▏| 12/84 [12:17<1:13:47, 61.50s/it, est. speed input: 1
Processed prompts:  29%|▎| 24/84 [13:16<28:10, 28.17s/it, est. speed input: 3.4
#374
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:51685,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 194
Wall time 11.0 seconds
Output:
#375
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "194",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 34.4 seconds
Output:

Processed prompts:  75%|▊| 72/96 [13:35<03:30,  8.75s/it, est. speed input: 9.9