gpt-5.6-sol
on-policy-self-distillation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- AIME24 avg@12 accuracy
- actions
- 596
- duration
- 7h 59m
- recorded
- 2026-08-03
run idopsd-aime24-4gpus__kuc95M7
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗This evaluation repeat is generating longer traces on several fixed shards, another visible consequence of training nondeterminism. The benchmark still uses identical prompts and sampling seed, so the final accuracy remains directly comparable.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 151
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "151",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 41.8 seconds
Output:
{"chunk_id":"2502b2","wall_time_seconds":60.001588063,"session_id":3244,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 152
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "152",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 43.6 seconds
Output:
{"chunk_id":"b26881","wall_time_seconds":60.002524569,"session_id":3244,"original_token_count":33,"output":"\rProcessed prompts: 38%|███▊ | 36/96 [07:31<11:19, 11.33s/it, est. speed input: 10.47 toks/s, output: 693.92 toks/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 153
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "153",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 36.6 seconds
Output:
{"chunk_id":"cc8cb2","wall_time_seconds":60.001474891,"session_id":3244,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 154
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "154",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 41.6 seconds
Output:
{"chunk_id":"f402f7","wall_time_seconds":60.001311744,"session_id":3244,"original_token_count":163,"output":"\rProcessed prompts: 29%|██▊ | 24/84 [10:00<24:12, 24.21s/it, est. speed input: 3.00 toks/s, output: 477.45 toks/s]\rProcessed prompts: 50%|█████ | 48/96 [10:26<10:04, 12.60s/it, est. speed input: 9.01 toks/s, output: 793.74 toks/s] \rProcessed prompts: 14%|█▍ | 12/84 [10:26<1:02:41, 52.25s/it, est. speed input: 1.61 toks/s, output: 282.58 toks/s]\rProcessed prompts: 62%|██████▎ | 60/96 [10:46<05:11, 8.66s/it, est. speed input: 10.97 toks/s, output: 1082.38 toks/s]\rProcessed prompts: 25%|██▌ | 24/96 [11:01<32:06, 26.76s/it, est. speed input: 4.81 toks/s, output: 521.49 toks/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 155
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "155",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 43.6 seconds
Output:
{"chunk_id":"39cd95","wall_time_seconds":60.001383512,"session_id":3244,"original_token_count":68,"output":"\rProcessed prompts: 75%|███████▌ | 72/96 [11:17<02:38, 6.59s/it, est. speed input: 12.03 toks/s, output: 1341.83 toks/s]\rProcessed prompts: 43%|████▎ | 36/84 [12:01<14:13, 17.77s/it, est. speed input: 3.87 toks/s, output: 656.70 toks/s]"}One shard has unusually long traces (only ~14% after ten minutes), while another is 75%. I’m letting the official-style run finish rather than truncating; max tokens and all inference settings remain frozen as required.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 156
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "156",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 38.6 seconds
Output:
Warning: truncated output (original token count: 276)
Total output lines: 1
{"chunk_id":"74247b","wall_time_seconds":60.000855985,"session_id":3244,"original_token_count":246,"output":"\rProcessed prompts: 57%|█████▋ | 48/84 [12:13<06:41, 11.14s/it, est. speed input: 6.02 toks/s, output: 1030.19 toks/s]\rProcessed prompts: 88%|████████▊ | 84/96 [12:28<01:16, 6.38s/it, est. speed input: 17.56 toks/s, output: 1669.78 toks/s]\rProcessed prompts: 71%|███████▏ | 60/84 [12:45<03:13, 8.08s/it, est. speed input: 7.40 toks/s…26 tokens truncated…5.40s/it, est. speed input: 8.85 toks/s, output: 1766.51 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:06<00:00, 5.36s/it, est. speed input: 19.30 toks/s, output: 2109.40 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:06<00:00, 5.36s/it, est. speed input: 19.30 toks/s, output: 2109.40 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [13:06<00:00, 8.19s/it, est. speed input: 19.30 toks/s, output: 2109.40 toks/s]\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 157
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "157",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 43.0 seconds
Output:
Warning: truncated output (original token count: 306)
Total output lines: 1
{"chunk_id":"ca8ad8","wall_time_seconds":60.002537519,"session_id":3244,"original_token_count":306,"output":"Warning: truncated output (original token count: 306)\nTotal output lines: 2\n\n\rProcessed prompts: 100%|██████████| 84/84 [13:25<00:00, 4.66s/it, est. speed input: 10.84 toks/s, output: 2045.79 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [13:25<00:00, 4.66s/it, est. speed input: 10.84 toks/s, output: 2045.79 toks/s]\rProcessed prompts: 10…56 tokens truncated… | 36/84 [13:47<13:43, 17.15s/it, est. speed input: 4.20 toks/s, output: 870.72 toks/s]\rProcessed prompts: 50%|█████ | 48/96 [13:56<10:28, 13.09s/it, est. speed input: 7.92 toks/s, output: 1079.32 toks/s]\rProcessed prompts: 57%|█████▋ | 48/84 [13:58<06:26, 10.73s/it, est. speed input: 5.87 toks/s, output: 1197.87 toks/s]\rProcessed prompts: 71%|███████▏ | 60/84 [14:22<03:02, 7.60s/it, est. speed input: 7.09 toks/s, output: 1494.57 toks/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 158
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "158",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 38.1 seconds
Output:
{"chunk_id":"fe06b5","wall_time_seconds":60.002059484,"session_id":3244,"original_token_count":177,"output":"\rProcessed prompts: 86%|████████▌ | 72/84 [14:49<01:09, 5.78s/it, est. speed input: 8.21 toks/s, output: 1739.12 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:17<00:00, 4.64s/it, est. speed input: 10.84 toks/s, output: 2048.83 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:17<00:00, 4.64s/it, est. speed input: 10.84 toks/s, output: 2048.83 toks/s]\rProcessed prompts: 100%|██████████| 84/84 [15:17<00:00, 10.92s/it, est. speed input: 10.84 toks/s, output: 2048.83 toks/s]\n\rProcessed prompts: 62%|██████▎ | 60/96 [15:21<06:33, 10.93s/it, est. speed input: 8.88 toks/s, output: 1224.45 toks/s]"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3244,chars:\"\",yield_time_ms:60000,max_output_tokens:250});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 159
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "159",
"yield_time_ms": 60000,
"max_tokens": 250
}wait result
Script completed
Wall time 32.8 seconds
Output:
{"chunk_id":"3974fc","wall_time_seconds":60.001815364,"session_id":3244,"original_token_count":213,"output":"\rProcessed prompts: 75%|███████▌ | 72/96 [16:27<03:38, 9.09s/it, est. speed input: 9.79 toks/s, output: 1389.28 toks/s]\rProcessed prompts: 88%|████████▊ | 84/96 [16:48<01:20, 6.70s/it, est. speed input: 10.88 toks/s, output: 1724.54 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:50<00:00, 4.60s/it, est. speed input: 12.95 toks/s, output: 2093.39 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:50<00:00, 4.60s/it, est. speed input: 12.95 toks/s, output: 2093.39 toks/s]\rProcessed prompts: 100%|██████████| 96/96 [16:50<00:00, 10.52s/it, est. speed input: 12.95 toks/s, output: 2093.39 toks/s]\n\nFINAL avg@12 accuracy: 56.11% (202/360 solutions correct on 30 problems)\nResults written to /workspace/eval_results/dev_eval_1785739883.json\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"cat /workspace/.timer/remaining_secs && git -C /workspace/repo diff --stat && git -C /workspace/repo diff -- opsd_trainer.py opsd_train.py && echo RECIPE && sed -n '1,120p' /workspace/submission/recipe.env\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":12000});\ntext(r.output);\n"
}exec result
Script completed
Wall time 0.1 seconds
Output:
12516
README.md | 4 +---
opsd_trainer.py | 24 ++++++++++++++++++------
2 files changed, 19 insertions(+), 9 deletions(-)
diff --git a/opsd_trainer.py b/opsd_trainer.py
index c151bd7..7549be4 100644
--- a/opsd_trainer.py
+++ b/opsd_trainer.py
@@ -463,20 +463,32 @@ class OPSDTrainer(SFTTrainer):
if token_clip is not None:
jsd = jsd.clamp(max=token_clip)
- # Masking
+ # Masking and trajectory credit assignment. Early reasoning tokens are
+ # more likely to encode a consequential branch choice; late tokens are
+ # increasingly determined by the already-long prefix. Give early
+ # positions moderately more credit without changing the total loss
+ # scale. The linear 1.5 -> 0.5 schedule has mean one on every non-empty
+ # trajectory, so learning-rate and clipping behavior remain comparable
+ # to the original uniformly averaged OPSD objective.
if labels is not None:
mask = labels != -100
- jsd = jsd[mask]
+ token_divergence = jsd.sum(dim=-1)
+ positions = mask.long().cumsum(dim=-1) - 1
+ lengths = mask.sum(dim=-1, keepdim=True).clamp_min(1)
+ progress = positions.to(token_divergence.dtype) / (lengths - 1).clamp_min(1)
+ credit = (1.5 - progress).clamp_min(0.5) * mask
# Apply reduction
if reduction == "batchmean":
- return jsd.sum() / mask.sum() if labels is not None else jsd.sum() / jsd.size(0)
+ if labels is not None:
+ return (token_divergence * credit).sum() / mask.sum().clamp_min(1)
+ return jsd.sum() / jsd.size(0)
elif reduction == "sum":
- return jsd.sum()
+ return (token_divergence * credit).sum() if labels is not None else jsd.sum()
elif reduction == "mean":
- return jsd.mean()
+ return (token_divergence * credit).sum() / credit.sum().clamp_min(1) if labels is not None else jsd.mean()
else:
- return jsd
+ return token_divergence * credit if labels is not None else jsd
def _update_ema(self):
"""Update EMA parameters after an optimizer step.
RECIPE
# OPSD method recipe — whitelisted hyper-parameter overrides (KEY=VALUE per line).
# Empty/all-commented == the OPSD baseline recipe on top of your /workspace/repo
# code changes. The scorer re-trains with THIS file + your repo for 100 steps.
# Budget knobs (max_steps, epochs, batch, gradient_accumulation, num_processes,
# model) are FIXED by the scorer and ignored here.
#
# learning_rate=5e-6
# lora_r=64
# lora_alpha=128
# beta=0
# jsd_token_clip=0.05
# temperature=1.1
# lmbda=1
# max_completion_length=1024 # <= 4096
# lr_scheduler_type=constant
# fixed_teacher=true
# use_ema_teacher=false
# reason_first=false
# teacher_thinking=false
The independent repeat also scored exactly 56.11% (202/360). The two-run mean is therefore 56.11, +1.89 over the official baseline—well beyond the scored-noise scale. With that result secured, I’ll spend one experiment tuning the credit slope from 1.5→0.5 to a stronger but still normalized 1.75→0.25; if it does not beat 56.11, I’ll revert immediately.
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/repo/opsd_trainer.py\\n@@\\n- # scale. The linear 1.5 -> 0.5 schedule has mean one on every non-empty\\n+ # scale. The linear 1.75 -> 0.25 schedule has mean one on every non-empty\\n@@\\n- credit = (1.5 - progress).clamp_min(0.5) * mask\\n+ credit = (1.75 - 1.5 * progress).clamp_min(0.25) * mask\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/earlycredit_strong PORT=12954 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":1000,\"max_output_tokens\":500});\ntext(JSON.stringify(r));\n"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{}{"chunk_id":"517f8e","wall_time_seconds":1.001851853,"session_id":53588,"original_token_count":0,"output":""}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:53588,chars:\"\",yield_time_ms:60000,max_output_tokens:220});\ntext(JSON.stringify(r));\n"
}exec result
Script running with cell ID 162
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "162",
"yield_time_ms": 60000,
"max_tokens": 220
}wait result
Script completed
Wall time 43.1 seconds
Output:
Warning: truncated output (original token count: 279)
Total output lines: 1
{"chunk_id":"42395a","wall_time_seconds":60.001255324,"session_id":53588,"original_token_count":4092,"output":"Warning: truncated output (original token count: 4092)\nTotal output lines: 197\n\n<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.\n\n================================================================================\nRUN CONFIGURATION\n===================…59 tokens truncated…, avg length: 735.2, speed: 902.8 tok/s\nvLLM generation done - elapsed time: 3.26s, prompts: 4, total tokens: 2933, avg length: 733.2, speed: 899.8 tok/s\nvLLM generation done - elapsed time: 3.27s, prompts: 4, total tokens: 3232, avg length: 808.0, speed: 987.4 tok/s\nvLLM generation done - elapsed time: 3.33s, prompts: 4, total tokens: 3575, avg length: 893.8, speed: 1073.3 tok/s\n\r 3%|▎ | 3/100 [00:23<12:50, 7.95s/it]"}