Skip to content
on-policy-self-distillation

gpt-5.6-sol

on-policy-self-distillation — graded by the task's verifier against a measured baseline.

backend
modal
variant
codex azure_ai/gpt-5.6-sol
metric
AIME24 avg@12 accuracy
actions
428
duration
9h 21m
recorded
2026-08-02

run idopsd-aime24-4gpus__eNDoWrD

Everything gpt-5.6-sol did on this task, in order.

download JSON ↗
codexgpt-5.6-sol444 turns16 narration428 reasoning428 action428 observation1 final
#426
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#427
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#428
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#429
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#430
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#431
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#432
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#433
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#434
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#435
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#436
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#437
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script running with cell ID 129
Wall time 61.0 seconds
Output:
#438
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "129",
  "yield_time_ms": 60000,
  "max_tokens": 60
}
wait result
Script completed
Wall time 42.9 seconds
Output:
Warning: truncated output (original token count: 185)
Total output lines: 19

, ?it/s, est. speed input: 0.00 toks/s, output: 0.00 toks/s]


s/it, est. speed input: 18.69 toks/s, output: 916.79 toks…125 tokens truncated…it, est. speed input: 8.88 toks/s, output: 1630.39 toks/s]
s/it, est. speed input: 7.74 toks/s, output: 1101.39 toks/s]
#439
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r=await tools.write_stdin({session_id:63551,chars:\"\",yield_time_ms:60000,max_output_tokens:100});text(r.output);\n"
}
exec result
Script running with cell ID 130
Wall time 11.0 seconds
Output:
#440
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolwait
args
{
  "cell_id": "130",
  "yield_time_ms": 60000,
  "max_tokens": 100
}
wait result
Script completed
Wall time 0.0 seconds
Output:
Warning: truncated output (original token count: 126)
Total output lines: 7

Warning: truncated output (original token count: 177)
Total output lines: 4


Processed prompts:  75%|███████▌  | 72/96 [17:47<03:40,  9.18s/it, est. speed input: 9.06 toks/s, output: …26 tokens truncated…, est. speed input: 12.11 toks/s, output: 2033.85 toks/s]

FINAL avg@12 accuracy: 57.22% (206/360 solutions correct on 30 problems)
Results written to /workspace/eval_results/dev_eval_1785714009.json
#441
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"cat /workspace/.timer/remaining_secs && git diff --check && python -m py_compile /workspace/repo/opsd_train.py /workspace/repo/opsd_trainer.py /workspace/repo/data_collator.py && cat /workspace/submission/recipe.env && git -C /workspace/repo status --short\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
1723
warning: Not a git repository. Use --no-index to compare two paths outside a working tree
usage: git diff --no-index [<options>] <path> <path>

Diff output format options
    -p, --patch           generate patch
    -s, --no-patch        suppress diff output
    -u                    generate patch
    -U, --unified[=<n>]   generate diffs with <n> lines context
    -W, --function-context
                          generate diffs with <n> lines context
    --raw                 generate the diff in raw format
    --patch-with-raw      synonym for '-p --raw'
    --patch-with-stat     synonym for '-p --stat'
    --numstat             machine friendly --stat
    --shortstat           output only the last line of --stat
    -X, --dirstat[=<param1,param2>...]
                          output the distribution of relative amount of changes for each sub-directory
    --cumulative          synonym for --dirstat=cumulative
    --dirstat-by-file[=<param1,param2>...]
                          synonym for --dirstat=files,param1,param2...
    --check               warn if changes introduce conflict markers or whitespace errors
    --summary             condensed summary such as creations, renames and mode changes
    --name-only           show only names of changed files
    --name-status         show only names and status of changed files
    --stat[=<width>[,<name-width>[,<count>]]]
                          generate diffstat
    --stat-width <width>  generate diffstat with a given width
    --stat-name-width <width>
                          generate diffstat with a given name width
    --stat-graph-width <width>
                          generate diffstat with a given graph width
    --stat-count <count>  generate diffstat with limited lines
    --compact-summary     generate compact summary in diffstat
    --binary              output a binary diff that can be applied
    --full-index          show full pre- and post-image object names on the "index" lines
    --color[=<when>]      show colored diff
    --ws-error-highlight <kind>
                          highlight whitespace errors in the 'context', 'old' or 'new' lines in the diff
    -z                    do not munge pathnames and use NULs as output field terminators in --raw or --numstat
    --abbrev[=<n>]        use <n> digits to display object names
    --src-prefix <prefix>
                          show the given source prefix instead of "a/"
    --dst-prefix <prefix>
                          show the given destination prefix instead of "b/"
    --line-prefix <prefix>
                          prepend an additional prefix to every line of output
    --no-prefix           do not show any source or destination prefix
    --inter-hunk-context <n>
                          show context between diff hunks up to the specified number of lines
    --output-indicator-new <char>
                          specify the character to indicate a new line instead of '+'
    --output-indicator-old <char>
                          specify the character to indicate an old line instead of '-'
    --output-indicator-context <char>
                          specify the character to indicate a context instead of ' '

Diff rename options
    -B, --break-rewrites[=<n>[/<m>]]
                          break complete rewrite changes into pairs of delete and create
    -M, --find-renames[=<n>]
                          detect renames
    -D, --irreversible-delete
                          omit the preimage for deletes
    -C, --find-copies[=<n>]
                          detect copies
    --find-copies-harder  use unmodified files as source to find copies
    --no-renames          disable rename detection
    --rename-empty        use empty blobs as rename source
    --follow              continue listing the history of a file beyond renames
    -l <n>                prevent rename/copy detection if the number of rename/copy targets exceeds given limit

Diff algorithm options
    --minimal             produce the smallest possible diff
    -w, --ignore-all-space
                          ignore whitespace when comparing lines
    -b, --ignore-space-change
                          ignore changes in amount of whitespace
    --ignore-space-at-eol
                          ignore changes in whitespace at EOL
    --ignore-cr-at-eol    ignore carrier-return at the end of line
    --ignore-blank-lines  ignore changes whose lines are all blank
    -I, --ignore-matching-lines <regex>
                          ignore changes whose all lines match <regex>
    --indent-heuristic    heuristic to shift diff hunk boundaries for easy reading
    --patience            generate diff using the "patience diff" algorithm
    --histogram           generate diff using the "histogram diff" algorithm
    --diff-algorithm <algorithm>
                          choose a diff algorithm
    --anchored <text>     generate diff using the "anchored diff" algorithm
    --word-diff[=<mode>]  show word diff, using <mode> to delimit changed words
    --word-diff-regex <regex>
                          use <regex> to decide what a word is
    --color-words[=<regex>]
                          equivalent to --word-diff=color --word-diff-regex=<regex>
    --color-moved[=<mode>]
                          moved lines of code are colored differently
    --color-moved-ws <mode>
                          how white spaces are ignored in --color-moved

Other diff options
    --relative[=<prefix>]
                          when run from subdir, exclude changes outside and show relative paths
    -a, --text            treat all files as text
    -R                    swap two inputs, reverse the diff
    --exit-code           exit with 1 if there were differences, 0 otherwise
    --quiet               disable all output of the program
    --ext-diff            allow an external diff helper to be executed
    --textconv            run external text conversion filters when comparing binary files
    --ignore-submodules[=<when>]
                          ignore changes to submodules in the diff generation
    --submodule[=<format>]
                          specify how differences in submodules are shown
    --ita-invisible-in-index
                          hide 'git add -N' entries from the index
    --ita-visible-in-index
                          treat 'git add -N' entries as real in the index
    -S <string>           look for differences that change the number of occurrences of the specified string
    -G <regex>            look for differences that change the number of occurrences of the specified regex
    --pickaxe-all         show all changes in the changeset with -S or -G
    --pickaxe-regex       treat <string> in -S as extended POSIX regular expression
    -O <file>             control the order in which files appear in the output
    --rotate-to <path>    show the change in the specified path first
    --skip-to <path>      skip the output to the specified path
    --find-object <object-id>
                          look for differences that change the number of occurrences of the specified object
    --diff-filter [(A|C|D|M|R|T|U|X|B)...[*]]
                          select files by diff type
    --output <file>       Output to a specific file

#442
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"git diff --check && python -m py_compile opsd_train.py opsd_trainer.py data_collator.py && git status --short && git diff --stat && cat /workspace/submission/recipe.env && python - <<'PY'\\nimport json\\nfor f in ['/workspace/eval_results/dev_eval_1785709356.json','/workspace/eval_results/dev_eval_1785711242.json','/workspace/eval_results/dev_eval_1785714009.json']:\\n d=json.load(open(f));print(f,d['average_at_n_pct'])\\nPY\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":5000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.1 seconds
Output:
README.md:285: trailing whitespace.
+| `--jsd_token_clip` | 0.05 | Clip the JSD loss for each token to a maximum value. This can improve stability. | 
#443
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const r = await tools.exec_command({cmd:\"python -m py_compile opsd_train.py opsd_trainer.py data_collator.py; git status --short; git diff --stat; cat /workspace/submission/recipe.env; python - <<'PY'\\nimport json\\nfor f in ['/workspace/eval_results/dev_eval_1785709356.json','/workspace/eval_results/dev_eval_1785711242.json','/workspace/eval_results/dev_eval_1785714009.json']:\\n d=json.load(open(f));print(f,d['average_at_n_pct'])\\nPY\",\"workdir\":\"/workspace/repo\",\"yield_time_ms\":10000,\"max_output_tokens\":5000});\ntext(r.output);"
}
exec result
Script completed
Wall time 0.2 seconds
Output:
 M README.md
 M opsd_train.py
 README.md     |  4 +---
 opsd_train.py | 13 +++++++++++++
 2 files changed, 14 insertions(+), 3 deletions(-)
# Difficulty-curriculum OPSD with the released optimizer/model settings.
lmbda=0
top_k_loss=0
lora_r=64
lora_alpha=128
max_completion_length=2048
lr_scheduler_type=linear
/workspace/eval_results/dev_eval_1785709356.json 55.55555555555556
/workspace/eval_results/dev_eval_1785711242.json 57.22222222222222
/workspace/eval_results/dev_eval_1785714009.json 57.22222222222222
#444
reasoning (redacted)

encrypted chain-of-thought — recorded, not exposed

toolexec
args
{
  "input": "const p = await tools.update_plan({explanation:\"Implemented and validated a reasoning-effort curriculum. Two independent 1024-token retrains scored 55.56 and 57.22; the final 2048-token interaction test scored 57.22. Removed all rejected experimental loss/anchor branches and retained only the validated dataset filter plus recipe.\",\"plan\":[\n  {step:\"Audit repository, paper, baseline recipe/logs, and exact capped launcher behavior\",status:\"completed\"},\n  {step:\"Design and implement a budget-neutral OPSD method improvement\",status:\"completed\"},\n  {step:\"Run targeted training/evaluation experiments and compare candidates\",status:\"completed\"},\n  {step:\"Finalize code and recipe, then verify scorer-compatible training\",status:\"completed\"}\n]}); text(p);\n"
}
exec result
Script completed
Wall time 0.0 seconds
Output:
{}
final summary

Implemented a reasoning-effort curriculum in opsd_train.py:

  • Trains only on dataset examples with generated_token_count >= 3000—13,565 hard, diverse problems.
  • Uses 2,048-token rollouts to cover later reasoning and answers.
  • Retains the paper’s linear schedule and original OPSD loss.

Full AIME24 dev results:

  • 1,024-token retrains: 55.56%, 57.22%
  • Final 2,048-token method: 57.22%
  • Baseline anchor: 54.22%

Final settings are in recipe.env. Syntax and scorer compatibility were verified.