Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
636
duration
8h 10m
recorded
2026-08-03

run idopsd-aime24-4gpus__3tyLi4e

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol683 turns47 narration636 reasoning636 action636 observation1 final
#201
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "109",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 34.9 seconds
Output:

Processed prompts:  62%|▋| 60/96 [12:38<06:59, 11.64s/it, est. speed input: 9.5
Processed prompts:  75%|▊| 72/96 [13:08<03:24,  8.52s/it, est. speed input: 10.
#202
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 110
Wall time 11.0 seconds
Output:
#203
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "110",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 33.3 seconds
Output:

Processed prompts:  29%|▎| 24/84 [13:43<30:04, 30.08s/it, est. speed input: 3.1
#204
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 111
Wall time 11.0 seconds
Output:
#205
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "111",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 34.6 seconds
Output:

Processed prompts:  43%|▍| 36/84 [15:01<15:26, 19.31s/it, est. speed input: 4.2
Processed prompts:  88%|▉| 84/96 [15:14<01:50,  9.17s/it, est. speed input: 14.
#206
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 112
Wall time 11.0 seconds
Output:
#207
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "112",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 34.3 seconds
Output:

Processed prompts: 100%|█| 96/96 [15:23<00:00,  6.50s/it, est. speed input: 16.
Processed prompts: 100%|█| 96/96 [15:23<00:00,  6.50s/it, est. speed input: 16.
Processed prompts: 100%|█| 96/96 [15:23<00:00,  9.62s/it, est. speed input: 16.

Processed prompts:  38%|▍| 36/96 [15:51<27:22, 27.37s/it, est. speed input: 4.9
#208
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 113
Wall time 11.0 seconds
Output:
#209
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "113",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 32.0 seconds
Output:

Processed prompts:  57%|▌| 48/84 [16:18<08:31, 14.21s/it, est. speed input: 5.2
Processed prompts: 100%|█| 84/84 [16:57<00:00, 12.79s/it, est. speed input: 8.5
Processed prompts: 100%|█| 84/84 [16:57<00:00, 12.79s/it, est. speed input: 8.5
Processed prompts: 100%|█| 84/84 [16:57<00:00, 12.11s/it, est. speed input: 8.5

Processed prompts:  71%|▋| 60/84 [16:59<04:08, 10.34s/it, est. speed input: 6.0
#210
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 114
Wall time 11.0 seconds
Output:
#211
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "114",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 25.2 seconds
Output:

Processed prompts:  50%|▌| 48/96 [17:34<15:57, 19.94s/it, est. speed input: 6.7
Processed prompts:  62%|▋| 60/96 [18:02<08:09, 13.60s/it, est. speed input: 8.5
#212
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 115
Wall time 11.0 seconds
Output:
#213
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "115",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 33.1 seconds
Output:

Processed prompts:  75%|▊| 72/96 [18:45<04:04, 10.18s/it, est. speed input: 9.4
Processed prompts:  86%|▊| 72/84 [18:57<02:01, 10.16s/it, est. speed input: 7.7
Processed prompts: 100%|█| 84/84 [19:00<00:00,  6.91s/it, est. speed input: 8.7
Processed prompts: 100%|█| 84/84 [19:00<00:00,  6.91s/it, est. speed input: 8.7
Processed prompts: 100%|█| 84/84 [19:00<00:00, 13.57s/it, est. speed input: 8.7
#214
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:50000,max_output_tokens:500});\ntext(r.output);"
}
exec result
Script running with cell ID 116
Wall time 11.0 seconds
Output:
#215
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "116",
  "yield_time_ms": 50000,
  "max_tokens": 500
}
wait result
Script completed
Wall time 35.6 seconds
Output:

Processed prompts:  88%|▉| 84/96 [19:48<01:42,  8.57s/it, est. speed input: 9.7
Processed prompts: 100%|█| 96/96 [20:24<00:00,  6.80s/it, est. speed input: 10.
Processed prompts: 100%|█| 96/96 [20:24<00:00,  6.80s/it, est. speed input: 10.
Processed prompts: 100%|█| 96/96 [20:24<00:00, 12.75s/it, est. speed input: 10.
#216
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:35310,chars:\"\",yield_time_ms:10000,max_output_tokens:1000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.0 seconds
Output:

FINAL avg@12 accuracy: 53.89% (194/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/frontier_v1_rep2_full.json
#217
narration

Repeat two scored 53.89% (194/360), so the two-run mean is 55.14%: still +0.92 over baseline, but with a 2.50-point spread. The candidate is directionally positive, not yet decisive. I’m inspecting whether the low repeat lost reasoning or merely format, then I’ll test the full-response verified term refinement rather than prematurely locking in a noisy result.

#218
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"python - <<'PY'\\nimport json\\nfiles=['/workspace/baseline/logs/official_eval_result.json','/workspace/eval_results/frontier_v1_full.json','/workspace/eval_results/frontier_v1_rep2_full.json']\\nxs=[json.load(open(f)) for f in files]\\nfor f,x in zip(files,xs): print(f,x['average_at_n_pct'],x['format_rate_pct'])\\nRs=[{z['problem_id']:z for z in x['results']} for x in xs]\\nprint('id base r1 r2 avgdelta')\\nfor i in sorted(Rs[0]):\\n a=[r[i]['num_correct'] for r in Rs]\\n print(i,*a,round((a[1]+a[2])/2-a[0],1))\\nPY\\ncat /workspace/.timer/remaining_secs\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":6000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
/workspace/baseline/logs/official_eval_result.json 54.44444444444444 96.94444444444444
/workspace/eval_results/frontier_v1_full.json 56.388888888888886 99.72222222222223
/workspace/eval_results/frontier_v1_rep2_full.json 53.888888888888886 100.0
id base r1 r2 avgdelta
60 12 12 12 0.0
61 10 10 11 0.5
62 0 0 0 0.0
63 0 0 0 0.0
64 5 6 4 0.0
65 6 4 5 -1.5
66 11 12 9 -0.5
67 12 12 12 0.0
68 12 12 10 -1.0
69 12 12 12 0.0
70 5 6 7 1.5
71 12 12 12 0.0
72 12 12 12 0.0
73 0 0 0 0.0
74 10 10 10 0.0
75 2 3 2 0.5
76 10 8 9 -1.5
77 1 3 0 0.5
78 5 9 3 1.0
79 12 12 12 0.0
80 2 3 8 3.5
81 0 0 0 0.0
82 7 7 8 0.5
83 11 10 12 0.0
84 11 12 12 1.0
85 0 0 0 0.0
86 12 10 11 -1.5
87 4 6 1 -0.5
88 0 0 0 0.0
89 0 0 0 0.0
16551
#219
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/repo/opsd_trainer.py\\n@@\\n-                correct_weights = correct_mask.to(token_jsd.dtype) * position_weights.unsqueeze(0)\\n+                # The correctness anchor covers the whole successful trajectory,\\n+                # including late final-answer and boxing tokens.  Unlike the dense\\n+                # teacher KL, it should not discount those late tokens.\\n+                correct_weights = correct_mask.to(token_jsd.dtype)\\n*** End Patch\";\ntext(await tools.apply_patch(patch));\nconst r = await tools.exec_command({cmd:\"python -m py_compile opsd_trainer.py && git diff --check -- opsd_trainer.py data_collator.py opsd_train.py\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.2 seconds
Output:
{}
#220
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"SEED=42 OUT_DIR=/workspace/runs/frontier_v2 PORT=12952 bash /workspace/train_capped.sh\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":600,\"tty\":true});\ntext(JSON.stringify(r));"
}
exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"99a0c3","wall_time_seconds":1.00180569,"session_id":63578,"original_token_count":0,"output":""}
#221
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:300});\ntext(r.output);"
}
exec result
Script running with cell ID 121
Wall time 11.0 seconds
Output:
#222
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "121",
  "yield_time_ms": 50000,
  "max_tokens": 300
}
wait result
Script completed
Wall time 34.8 seconds
Output:
Warning: truncated output (original token count: 327)
Total output lines: 19

Warning: truncated output (original token count: 3394)
Total output lines: 169

<string>:24: FutureWarning: `torch_dtype` is deprecated and will be removed in version 0.27.0, please use `dtype` instead.

================================================================================
RUN CONFIGURATION
================================================================================
WandB Run Name: capped_seed42_lr5e-6_bs32
Output Directory: /workspace/runs/frontier_v2/capped_seed42
================================================================================


=====================…27 tokens truncated…] If this is not desired, please set os.environ['TORCH_CUDA_ARCH_LIST'] to specific architectures.

  0%|          | 0/100 [00:00<?, ?it/s]vLLM generation done - elapsed time: 3.70s, prompts: 4, total tokens: 2785, avg length: 696.2, speed: 753.7 tok/s
vLLM generation done - elapsed time: 3.69s, prompts: 4, total tokens: 3158, avg length: 789.5, speed: 856.4 tok/s
vLLM generation done - elapsed time: 3.74s, prompts: 4, total tokens: 3511, avg length: 877.8, speed: 938.7 tok/s
vLLM generation done - elapsed time: 3.79s, prompts: 4, total tokens: 3617, avg length: 904.2, speed: 955.1 tok/s
#223
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 122
Wall time 11.0 seconds
Output:
#224
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "122",
  "yield_time_ms": 50000,
  "max_tokens": 200
}
wait result
Script completed
Wall time 33.1 seconds
Output:
Warning: truncated output (original token count: 227)
Total output lines: 11

Warning: truncated output (original token count: 3038)
Total output lines: 170

/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None
  warnings.warn(
/usr/local/lib/python3.10/dist-packages/torch/utils/checkpoint.py:85: UserWarning: None of the inputs have requires_grad=True. Gradients will be None
…27 tokens truncated… tokens: 2716, avg length: 679.0, speed: 758.1 tok/s
vLLM generation done - elapsed time: 3.66s, prompts: 4, total tokens: 2993, avg length: 748.2, speed: 817.8 tok/s
vLLM generation done - elapsed time: 3.68s, prompts: 4, total tokens: 3390, avg length: 847.5, speed: 921.2 tok/s
vLLM generation done - elapsed time: 3.71s, prompts: 4, total tokens: 3822, avg length: 955.5, speed: 1029.4 tok/s
#225
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.write_stdin({session_id:63578,chars:\"\",yield_time_ms:50000,max_output_tokens:200});\ntext(r.output);"
}
exec result
Script running with cell ID 123
Wall time 11.0 seconds
Output: