Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
627
duration
6h 30m
recorded
2026-08-02

run idopsd-aime24-4gpus__RqGcDty

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol684 turns57 narration627 reasoning627 action627 observation1 final
#401
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "214",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 13.0 seconds
Output:

Processed prompts:  71%|███████▏  | 60/84 [07:56<02:35,  6.48s/it, est. speed input: 13.34 toks/s, output: 1566.51 toks/s]
#402
narration

The robust checkpoint is generating normally; one shard is slower because of max-length reasoning, while another has passed 70%. I’m continuing to use only the final merged accuracy for selection.

#403
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 215
Wall time 11.0 seconds
Output:
#404
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "215",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 9.3 seconds
Output:

Processed prompts:  50%|█████     | 48/96 [08:00<07:46,  9.71s/it, est. speed input: 12.16 toks/s, output: 1030.55 toks/s]
Processed prompts:  86%|████████▌ | 72/84 [08:07<00:54,  4.58s/it, est. speed input: 15.61 toks/s, output: 2054.50 toks/s]
#405
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 216
Wall time 11.0 seconds
Output:
#406
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "216",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 13.4 seconds
Output:

Processed prompts:  62%|██████▎   | 60/96 [08:52<04:40,  7.78s/it, est. speed input: 12.56 toks/s, output: 1216.29 toks/s]
Processed prompts:  25%|██▌       | 24/96 [09:11<28:48, 24.00s/it, est. speed input: 5.77 toks/s, output: 527.13 toks/s]
Processed prompts:  75%|███████▌  | 72/96 [09:14<02:17,  5.74s/it, est. speed input: 21.09 toks/s, output: 1703.65 toks/s]
Processed prompts:  29%|██▊       | 24/84 [09:21<21:21, 21.35s/it, est. speed input: 4.08 toks/s, output: 616.40 toks/s]
#407
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 217
Wall time 11.0 seconds
Output:
#408
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "217",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 12.3 seconds
Output:

Processed prompts: 100%|██████████| 84/84 [09:32<00:00,  5.40s/it, est. speed input: 15.26 toks/s, output: 2291.29 toks/s]
Processed prompts: 100%|██████████| 84/84 [09:32<00:00,  5.40s/it, est. speed input: 15.26 toks/s, output: 2291.29 toks/s]
Processed prompts: 100%|██████████| 84/84 [09:32<00:00,  6.81s/it, est. speed input: 15.26 toks/s, output: 2291.29 toks/s]

Processed prompts:  38%|███▊      | 36/96 [09:59<14:52, 14.87s/it, est. speed input: 9.35 toks/s, output: 866.13 toks/s]
Processed prompts:  50%|█████     | 48/96 [09:59<07:12,  9.01s/it, est. speed input: 11.04 toks/s, output: 1238.02 toks/s]
#409
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 218
Wall time 11.0 seconds
Output:
#410
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "218",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 9.8 seconds
Output:

Processed prompts:  43%|████▎     | 36/84 [10:05<10:37, 13.28s/it, est. speed input: 6.16 toks/s, output: 939.19 toks/s]
Processed prompts:  50%|█████     | 48/96 [10:10<07:12,  9.01s/it, est. speed input: 11.04 toks/s, output: 1238.02 toks/s]
Processed prompts:  57%|█████▋    | 48/84 [10:26<05:14,  8.72s/it, est. speed input: 7.87 toks/s, output: 1273.19 toks/s]
Processed prompts:  71%|███████▏  | 60/84 [10:45<02:27,  6.16s/it, est. speed input: 11.73 toks/s, output: 1657.53 toks/s]
#411
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 219
Wall time 11.0 seconds
Output:
#412
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "219",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 10.2 seconds
Output:

Processed prompts:  86%|████████▌ | 72/84 [10:49<00:50,  4.18s/it, est. speed input: 13.48 toks/s, output: 2093.31 toks/s]
Processed prompts: 100%|██████████| 84/84 [11:09<00:00,  3.34s/it, est. speed input: 14.87 toks/s, output: 2374.15 toks/s]
Processed prompts: 100%|██████████| 84/84 [11:09<00:00,  3.34s/it, est. speed input: 14.87 toks/s, output: 2374.15 toks/s]
Processed prompts: 100%|██████████| 84/84 [11:09<00:00,  7.97s/it, est. speed input: 14.87 toks/s, output: 2374.15 toks/s]

Processed prompts:  62%|██████▎   | 60/96 [11:16<04:49,  8.05s/it, est. speed input: 12.00 toks/s, output: 1414.77 toks/s]
Processed prompts:  88%|████████▊ | 84/96 [11:16<01:26,  7.19s/it, est. speed input: 19.42 toks/s, output: 1673.27 toks/s]
Processed prompts: 100%|██████████| 96/96 [11:16<00:00,  7.19s/it, est. speed input: 22.44 toks/s, output: 2152.58 toks/s]
Processed prompts: 100%|██████████| 96/96 [11:16<00:00,  7.05s/it, est. speed input: 22.44 toks/s, output: 2152.58 toks/s]
#413
narration

Two evaluation shards are complete and the others are past halfway. The robust model tends to finish these generations somewhat faster, but accuracy—not speed—will determine whether the new loss is retained.

#414
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 220
Wall time 11.0 seconds
Output:
#415
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "220",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 11.4 seconds
Output:
#416
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 221
Wall time 11.0 seconds
Output:
#417
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "221",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 13.5 seconds
Output:
#418
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 222
Wall time 11.0 seconds
Output:
#419
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "222",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 11.6 seconds
Output:
#420
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 223
Wall time 11.0 seconds
Output:
#421
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "223",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 13.5 seconds
Output:

Processed prompts:  75%|███████▌  | 72/96 [13:28<03:37,  9.05s/it, est. speed input: 11.66 toks/s, output: 1549.59 toks/s]
#422
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:79122,chars:\"\",yield_time_ms:30000,max_output_tokens:300});\ntext(r.output);\n"
}
exec result
Script running with cell ID 224
Wall time 11.0 seconds
Output:
#423
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "224",
  "yield_time_ms": 30000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 1.5 seconds
Output:

Processed prompts:  88%|████████▊ | 84/96 [14:04<01:25,  7.09s/it, est. speed input: 13.00 toks/s, output: 1828.50 toks/s]
Processed prompts: 100%|██████████| 96/96 [14:09<00:00,  4.97s/it, est. speed input: 15.39 toks/s, output: 2210.38 toks/s]
Processed prompts: 100%|██████████| 96/96 [14:09<00:00,  4.97s/it, est. speed input: 15.39 toks/s, output: 2210.38 toks/s]
Processed prompts: 100%|██████████| 96/96 [14:09<00:00,  8.85s/it, est. speed input: 15.39 toks/s, output: 2210.38 toks/s]

FINAL avg@12 accuracy: 50.00% (180/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/robust_kl_2048_full.json
#424
narration

The corrected-KL model scored 50.00% (180/360), decisively worse. The malformed-looking asymmetric estimator is evidently part of the paper recipe’s beneficial update geometry, so theoretical cleanliness did not translate to reasoning accuracy here. I’m reverting that loss change and 2,048-token setting, restoring the validated 57.22% method, then repeating it to estimate robustness.

#425
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch1 = \"*** Begin Patch\\n*** Update File: /workspace/repo/opsd_trainer.py\\n@@\\n-            token_clip:\\n-                If set, bounds the influence of each sequence token after\\n-                summing over the vocabulary. High-divergence tokens retain a\\n-                scaled gradient instead of being hard-clamped to zero.\\n+            token_clip:\\n+                if set, clips per-token divergence values to this maximum before reduction. Prevents style tokens from dominating the gradient signal over math tokens.\\n@@\\n-        # Aggregate vocabulary contributions into an actual divergence for\\n-        # each sequence token.  The released implementation clamps individual\\n-        # signed vocabulary summands before this sum.  That operation is\\n-        # asymmetric (positive terms are clipped, negative terms are not), can\\n-        # make a forward KL negative, and eventually trains away from its own\\n-        # teacher.  We instead use a detached influence cap on the complete\\n-        # token KL.  Easy/math tokens receive their full gradient; unusually\\n-        # large style/context-shift tokens receive a gradient scaled to the cap\\n-        # rather than no gradient at all.\\n-        token_divergence = jsd.sum(dim=-1)\\n+        # Per-token clipping: cap each token's divergence value\\n         if token_clip is not None:\\n-            influence = (token_clip / token_divergence.detach().clamp_min(token_clip)).clamp(max=1.0)\\n-            token_divergence = token_divergence * influence\\n+            jsd = jsd.clamp(max=token_clip)\\n \\n         # Masking\\n         if labels is not None:\\n             mask = labels != -100\\n-            token_divergence = token_divergence[mask]\\n+            jsd = jsd[mask]\\n \\n         # Apply reduction\\n         if reduction == \\\"batchmean\\\":\\n-            return (\\n-                token_divergence.sum() / mask.sum()\\n-                if labels is not None\\n-                else token_divergence.sum() / token_divergence.size(0)\\n-            )\\n+            return jsd.sum() / mask.sum() if labels is not None else jsd.sum() / jsd.size(0)\\n         elif reduction == \\\"sum\\\":\\n-            return token_divergence.sum()\\n+            return jsd.sum()\\n         elif reduction == \\\"mean\\\":\\n-            return token_divergence.mean()\\n+            return jsd.mean()\\n         else:\\n-            return token_divergence\\n+            return jsd\\n*** End Patch\";\nconst patch2 = \"*** Begin Patch\\n*** Update File: /workspace/submission/recipe.env\\n@@\\n-max_completion_length=2048\\n+max_completion_length=1536\\n*** End Patch\";\ntext(await tools.apply_patch(patch1));\ntext(await tools.apply_patch(patch2));\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{}{}