Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#451
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "238",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 10.8 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 866)
Total output lines: 31

vLLM generation done - elapsed time: 4.89s, prompts: 4, total tokens: 4702, avg le…26 tokens truncated…70, avg length: 1342.5, speed: 1061.0 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1197.1 tok/s
#452
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 239
Wall time 11.0 seconds
Output:
#453
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "239",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 4.7 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1100)
Total output lines: 34

vLLM generation done - elapsed time: 5.04s, prompts: 4, total tokens: 5207, avg l…27 tokens truncated…44, avg length: 1536.0, speed: 1199.5 tok/s
vLLM generation done - elapsed time: 5.12s, prompts: 4, total tokens: 5958, avg length: 1489.5, speed: 1163.3 tok/s
#454
narration

The repeat has reached roughly step 30 with the same stable dynamics. Since the scoring protocol averages five retrains, consistency across these two local full evaluations is more informative than further one-off hyperparameter changes.

#455
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 240
Wall time 11.0 seconds
Output:
#456
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "240",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 12.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 927)
Total output lines: 32

vLLM generation done - elapsed time: 4.87s, prompts: 4, total tokens: 4472, avg le…26 tokens truncated…045, avg length: 1261.2, speed: 998.6 tok/s
vLLM generation done - elapsed time: 5.08s, prompts: 4, total tokens: 5193, avg length: 1298.2, speed: 1022.0 tok/s
#457
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 241
Wall time 11.0 seconds
Output:
#458
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "241",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 11.1 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 865)
Total output lines: 31

vLLM generation done - elapsed time: 4.99s, prompts: 4, total tokens: 4549, avg le…26 tokens truncated…06, avg length: 1351.5, speed: 1059.4 tok/s
vLLM generation done - elapsed time: 5.13s, prompts: 4, total tokens: 6062, avg length: 1515.5, speed: 1182.2 tok/s
#459
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 242
Wall time 11.0 seconds
Output:
#460
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "242",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 983)
Total output lines: 31

vLLM generation done - elapsed time: 4.83s, prompts: 4, total tokens: 4045, avg le…26 tokens truncated… done - elapsed time: 5.08s, prompts: 4, total tokens: 5895, avg length: 1473.8, speed: 1159.9 tok/s

 45%|████▌     | 45/100 [08:52<10:49, 11.81s/it]
#461
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 243
Wall time 11.0 seconds
Output:
#462
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "243",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 13.6 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 931)
Total output lines: 33

vLLM generation done - elapsed time: 4.96s, prompts: 4, total tokens: 4500, avg le…26 tokens truncated…4076097905635834, 'learning_rate': 2.6500000000000005e-06, 'on_policy_loss': -0.0069, 'epoch': 0.42}

 48%|████▊     | 48/100 [09:28<10:17, 11.88s/it]
#463
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 244
Wall time 11.0 seconds
Output:
#464
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "244",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 0.4 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 982)
Total output lines: 35

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4824, avg le…26 tokens truncated…37, avg length: 1409.2, speed: 1106.3 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 5640, avg length: 1410.0, speed: 1097.7 tok/s
#465
narration

At the midpoint, the repeat’s loss and gradient norms closely match the 57.22% run. This is the expected behavior for the same method under nondeterministic on-policy rollouts; the decisive check remains its full AIME score.

#466
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 245
Wall time 11.0 seconds
Output:
#467
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "245",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 5.9 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 2888)
Total output lines: 207

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5686, avg …27 tokens truncated…83, avg length: 1470.8, speed: 1153.0 tok/s
vLLM generation done - elapsed time: 5.11s, prompts: 4, total tokens: 6130, avg length: 1532.5, speed: 1199.8 tok/s
#468
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 246
Wall time 11.0 seconds
Output:
#469
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "246",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 4.9 seconds
Output:
Warning: truncated output (original token count: 107)
Total output lines: 6

Warning: truncated output (original token count: 1174)
Total output lines: 40

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4733, avg l…27 tokens truncated…990, avg length: 1247.5, speed: 984.4 tok/s
vLLM generation done - elapsed time: 5.17s, prompts: 4, total tokens: 6144, avg length: 1536.0, speed: 1189.4 tok/s
#470
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 247
Wall time 11.0 seconds
Output:
#471
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "247",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 2.2 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 868)
Total output lines: 31

vLLM generation done - elapsed time: 5.01s, prompts: 4, total tokens: 4489, avg le…26 tokens truncated…01, avg length: 1500.2, speed: 1168.7 tok/s
vLLM generation done - elapsed time: 5.14s, prompts: 4, total tokens: 6103, avg length: 1525.8, speed: 1187.4 tok/s
#472
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 248
Wall time 11.0 seconds
Output:
#473
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "248",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 7.5 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 863)
Total output lines: 26

vLLM generation done - elapsed time: 5.06s, prompts: 4, total tokens: 5124, avg le…26 tokens truncated…75, avg length: 1393.8, speed: 1096.3 tok/s
vLLM generation done - elapsed time: 5.09s, prompts: 4, total tokens: 5146, avg length: 1286.5, speed: 1010.1 tok/s
#474
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:57840,chars:\"\",yield_time_ms:30000,max_output_tokens:80});\ntext(r.output);\n"
}
exec result
Script running with cell ID 249
Wall time 11.0 seconds
Output:
#475
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "249",
  "yield_time_ms": 30000,
  "max_tokens": 80
}
wait result
Script completed
Wall time 8.9 seconds
Output:
Warning: truncated output (original token count: 106)
Total output lines: 6

Warning: truncated output (original token count: 871)
Total output lines: 31

vLLM generation done - elapsed time: 5.10s, prompts: 4, total tokens: 5757, avg le…26 tokens truncated…71, avg length: 1467.8, speed: 1152.9 tok/s
vLLM generation done - elapsed time: 5.15s, prompts: 4, total tokens: 5974, avg length: 1493.5, speed: 1160.1 tok/s