gpt-5.6-sol
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- codex azure_ai/gpt-5.6-sol
- metric
- held-out multi-domain perplexity
- actions
- 222
- duration
- 57m
- recorded
- 2026-08-04
run iddclm-nano-curation__f5FdhJo
Everything gpt-5.6-sol did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python3 /workspace/work/pack_cache.py /workspace/work/combo_b30.json /workspace/work/train_combo_b30.npy && python3 /workspace/work/train_diag.py --train_npy /workspace/work/train_combo_b30.npy --heldout_npy /workspace/data/multi_dev.npy --out_json /workspace/work/combo_b30_eval.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"3bf88d","wall_time_seconds":1.00216152,"session_id":89729,"original_token_count":11,"output":"packed 12000000 tokens from 16463 documents\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89729,chars:\"\",yield_time_ms:30000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script running with cell ID 129
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "129",
"yield_time_ms": 20000,
"max_tokens": 5000
}wait result
Script completed
Wall time 8.2 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89729,chars:\"\",yield_time_ms:10000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script completed
Wall time 10.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:89729,chars:\"\",yield_time_ms:10000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script completed
Wall time 5.6 seconds
Output:
HELDOUT_PPL 301.856 loss 5.7100 train_s 56 /workspace/work/train_combo_b30.npy
QUARTER 0 loss 6.4874 ppl 656.79
QUARTER 1 loss 5.5847 ppl 266.32
QUARTER 2 loss 5.7614 ppl 317.78
QUARTER 3 loss 5.0033 ppl 148.9
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python3 /workspace/work/pack_cache.py /workspace/work/pen65.json /workspace/work/train_pen65.npy && python3 /workspace/work/train_diag.py --train_npy /workspace/work/train_pen65.npy --heldout_npy /workspace/data/multi_dev.npy --out_json /workspace/work/pen65_eval.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"e096e5","wall_time_seconds":1.001755762,"session_id":39045,"original_token_count":11,"output":"packed 12000000 tokens from 15527 documents\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"cat /workspace/.timer/remaining_secs\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":1000});\ntext(r.output);"
}exec result
Script completed
Wall time 0.1 seconds
Output:
3498
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:39045,chars:\"\",yield_time_ms:30000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script running with cell ID 134
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "134",
"yield_time_ms": 20000,
"max_tokens": 5000
}wait result
Script completed
Wall time 14.5 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:39045,chars:\"\",yield_time_ms:10000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script completed
Wall time 10.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:39045,chars:\"\",yield_time_ms:10000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script completed
Wall time 4.3 seconds
Output:
HELDOUT_PPL 303.234 loss 5.7145 train_s 56 /workspace/work/train_pen65.npy
QUARTER 0 loss 6.5417 ppl 693.47
QUARTER 1 loss 5.6065 ppl 272.19
QUARTER 2 loss 5.7668 ppl 319.53
QUARTER 3 loss 4.94 ppl 139.78
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python3 /workspace/work/pack_cache.py /workspace/work/pen135.json /workspace/work/train_pen135.npy && python3 /workspace/work/train_diag.py --train_npy /workspace/work/train_pen135.npy --heldout_npy /workspace/data/multi_dev.npy --out_json /workspace/work/pen135_eval.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"689afc","wall_time_seconds":1.001471112,"session_id":21543,"original_token_count":11,"output":"packed 12000000 tokens from 17348 documents\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Add File: /workspace/submission/curate.py\\n+#!/usr/bin/env python3\\n+\\\"\\\"\\\"Target-density curation with explicit four-domain token balancing.\\n+\\n+Criterion\\n+---------\\n+Treat the four equal quarters of ``multi_dev.npy`` as disclosed, unlabeled\\n+examples of the desired registers. Estimate smoothed GPT-2 unigram and hashed\\n+bigram density ratios between that target and the whole web pool. A document's\\n+quality score is its mean clipped density ratio minus deterministic crawl-noise\\n+penalties. For each target quarter, rank documents by quality plus 0.85 times\\n+quarter-specific token evidence, then interleave the four rankings until each\\n+has supplied 3M tokens. Append the remaining ids in score order so the output\\n+always provides ample coverage.\\n+\\n+This is intentionally a corpus-level rule: no pool id, URL, or hand label is\\n+embedded here. Running the script recreates ``selection.json`` from the two\\n+provided data files. It uses a temporary 16-bit token cache to keep memory use\\n+reasonable; on the benchmark host it takes roughly ten minutes.\\n+\\\"\\\"\\\"\\n+\\n+import json\\n+import math\\n+import re\\n+import tempfile\\n+import time\\n+from pathlib import Path\\n+\\n+import numpy as np\\n+from transformers import AutoTokenizer\\n+\\n+POOL = Path(\\\"/workspace/data/pool.jsonl\\\")\\n+DEV = Path(\\\"/workspace/data/multi_dev.npy\\\")\\n+OUTPUT = Path(\\\"/workspace/submission/selection.json\\\")\\n+VOCAB = 50_257\\n+HASH_MASK = (1 << 20) - 1\\n+DOMAIN_QUOTA = 3_000_000\\n+\\n+\\n+def bigram_hash(a, b):\\n+ return (a.astype(np.uint64) * 1_000_003 + b.astype(np.uint64)) & HASH_MASK\\n+\\n+\\n+def crawl_penalty(q, lengths):\\n+ words, lines, alpha, printable, eos, questions, code, urls, bars, unique, punct, maxline = q.T\\n+ p = np.zeros(len(lengths), dtype=np.float64)\\n+ p += np.maximum(0, 0.60 - alpha) * 4\\n+ p += np.maximum(0, 0.985 - printable) * 8\\n+ p += np.maximum(0, 0.20 - unique) * 3\\n+ p += np.maximum(0, bars / np.maximum(words, 1) - 0.015) * 8\\n+ p += np.maximum(0, urls / np.maximum(words, 1) - 0.015) * 5\\n+ p += (words < 80) * 0.8 + (words < 30) * 1.5\\n+ p += np.maximum(0, np.log1p(lengths / 6000)) * 0.10\\n+ p += ((lines >= 5) & (punct < 0.12)) * 0.3\\n+ return p\\n+\\n+\\n+def interleave(rankings, lengths):\\n+ \\\"\\\"\\\"Draw from the token-furthest-behind domain, de-duplicating ids.\\\"\\\"\\\"\\n+ pos = [0] * 4\\n+ used = [0] * 4\\n+ out = []\\n+ seen = set()\\n+ while any(used[k] < DOMAIN_QUOTA and pos[k] < len(rankings[k]) for k in range(4)):\\n+ active = [k for k in range(4)\\n+ if used[k] < DOMAIN_QUOTA and pos[k] < len(rankings[k])]\\n+ k = min(active, key=lambda j: used[j] / DOMAIN_QUOTA)\\n+ while pos[k] < len(rankings[k]) and int(rankings[k][pos[k]]) in seen:\\n+ pos[k] += 1\\n+ if pos[k] == len(rankings[k]):\\n+ continue\\n+ doc_id = int(rankings[k][pos[k]])\\n+ pos[k] += 1\\n+ out.append(doc_id)\\n+ seen.add(doc_id)\\n+ used[k] += int(lengths[doc_id]) + 1 # packer adds one EOS per document\\n+ return out, seen, used\\n+\\n+\\n+def tokenize_pool(rows, token_path):\\n+ tokenizer = AutoTokenizer.from_pretrained(\\\"gpt2\\\")\\n+ backend = tokenizer.backend_tokenizer\\n+ n = len(rows)\\n+ offsets = np.zeros(n + 1, dtype=np.int64)\\n+ quality = np.zeros((n, 12), dtype=np.float32)\\n+ total = 0\\n+ started = time.time()\\n+ with token_path.open(\\\"wb\\\") as token_file:\\n+ for base in range(0, n, 256):\\n+ batch = rows[base:base + 256]\\n+ encoded = backend.encode_batch([r[\\\"text\\\"] for r in batch],\\n+ add_special_tokens=False)\\n+ batch_ids = []\\n+ for j, (row, enc) in enumerate(zip(batch, encoded)):\\n+ i = base + j\\n+ s = row[\\\"text\\\"]\\n+ ids = enc.ids\\n+ batch_ids.extend(ids)\\n+ total += len(ids)\\n+ offsets[i + 1] = total\\n+ words = re.findall(r\\\"\\\\b\\\\w+\\\\b\\\", s)\\n+ lines = [x.strip() for x in s.splitlines() if x.strip()]\\n+ alpha = sum(c.isalpha() for c in s)\\n+ printable = sum(c.isprintable() or c in \\\"\\\\n\\\\t\\\" for c in s)\\n+ code = len(re.findall(\\n+ r\\\"</?(?:p|pre|code|blockquote)\\\\b|\\\"\\n+ r\\\"\\\\b(?:def|class|function|SELECT|import)\\\\b|[{};]\\\", s, re.I))\\n+ quality[i] = (\\n+ len(words), len(lines), alpha / max(1, len(s)),\\n+ printable / max(1, len(s)), s.count(\\\"<|endoftext|>\\\"),\\n+ s.count(\\\"?\\\"), code, s.count(\\\"http\\\"), s.count(\\\"|\\\"),\\n+ len(set(w.lower() for w in words)) / max(1, len(words)),\\n+ sum(x[-1:] in \\\".!?\\\\\\\"'\\\" for x in lines) / max(1, len(lines)),\\n+ len(max(lines, key=len, default=\\\"\\\")),\\n+ )\\n+ np.asarray(batch_ids, dtype=np.uint16).tofile(token_file)\\n+ if base and base % 20_480 == 0:\\n+ print(f\\\"tokenized {base}/{n} documents in {time.time()-started:.0f}s\\\",\\n+ flush=True)\\n+ return offsets, quality\\n+\\n+\\n+def main():\\n+ rows = [json.loads(line) for line in POOL.open()]\\n+ if not all(r[\\\"id\\\"] == i for i, r in enumerate(rows)):\\n+ raise ValueError(\\\"pool ids must be the consecutive range 0..N-1\\\")\\n+\\n+ with tempfile.TemporaryDirectory(prefix=\\\"curate-\\\") as temp:\\n+ token_path = Path(temp) / \\\"tokens.uint16\\\"\\n+ offsets, quality = tokenize_pool(rows, token_path)\\n+ del rows\\n+ lengths = np.diff(offsets)\\n+ pool = np.memmap(token_path, dtype=np.uint16, mode=\\\"r\\\")\\n+ dev = np.load(DEV)\\n+ if len(dev) != 1_000_000:\\n+ raise ValueError(\\\"expected the disclosed 1M-token, four-quarter dev set\\\")\\n+\\n+ # Smoothed unigram target-vs-pool density ratio.\\n+ dcount = np.stack([\\n+ np.bincount(dev[k * 250_000:(k + 1) * 250_000], minlength=VOCAB)\\n+ for k in range(4)\\n+ ])\\n+ pcount = np.bincount(pool, minlength=VOCAB)\\n+ alpha = 10.0\\n+ target_lp = np.log((dcount.sum(0) + alpha) /\\n+ (len(dev) + alpha * VOCAB))\\n+ ratio = len(pool) / len(dev)\\n+ pool_lp = np.log((pcount + alpha * ratio) /\\n+ (len(pool) + alpha * ratio * VOCAB))\\n+ unigram_density_weight = np.clip(target_lp - pool_lp, -2.5, 2.5)\\n+ domain_lp = np.log((dcount + alpha) / (250_000 + alpha * VOCAB))\\n+ unigram_domain_weight = domain_lp - domain_lp.mean(0, keepdims=True)\\n+\\n+ # Hashed bigrams add sequence/markup style without a huge 50k^2 table.\\n+ d_hash_count = np.stack([\\n+ np.bincount(\\n+ bigram_hash(dev[k*250_000:(k+1)*250_000-1],\\n+ dev[k*250_000+1:(k+1)*250_000]).astype(np.int64),\\n+ minlength=HASH_MASK + 1)\\n+ for k in range(4)\\n+ ])\\n+ p_hash_count = np.zeros(HASH_MASK + 1, dtype=np.int64)\\n+ for start in range(0, len(pool) - 1, 4_000_000):\\n+ x = np.asarray(pool[start:min(len(pool), start + 4_000_001)])\\n+ p_hash_count += np.bincount(\\n+ bigram_hash(x[:-1], x[1:]).astype(np.int64),\\n+ minlength=HASH_MASK + 1)\\n+ hash_alpha = 1.0\\n+ target_hash_lp = np.log((d_hash_count.sum(0) + hash_alpha) /\\n+ (len(dev) + hash_alpha * (HASH_MASK + 1)))\\n+ pool_hash_lp = np.log((p_hash_count + hash_alpha * ratio) /\\n+ (len(pool) + hash_alpha * ratio * (HASH_MASK + 1)))\\n+ hash_density_weight = np.clip(target_hash_lp - pool_hash_lp, -2.5, 2.5)\\n+ hash_domain_lp = np.log(\\n+ (d_hash_count + hash_alpha) /\\n+ (250_000 + hash_alpha * (HASH_MASK + 1)))\\n+ hash_domain_weight = hash_domain_lp - hash_domain_lp.mean(0, keepdims=True)\\n+\\n+ density = np.zeros(len(lengths), dtype=np.float32)\\n+ domain = np.zeros((len(lengths), 4), dtype=np.float32)\\n+ for i, (lo, hi) in enumerate(zip(offsets[:-1], offsets[1:])):\\n+ x = pool[lo:hi]\\n+ if not len(x):\\n+ continue\\n+ density[i] = unigram_density_weight[x].mean()\\n+ domain[i] = unigram_domain_weight[:, x].mean(1)\\n+ if len(x) > 1:\\n+ h = bigram_hash(x[:-1], x[1:]).astype(np.int64)\\n+ density[i] += 0.45 * hash_density_weight[h].mean()\\n+ domain[i] += 0.45 * hash_domain_weight[:, h].mean(1)\\n+\\n+ penalty = crawl_penalty(quality, lengths)\\n+ base = density.astype(np.float64) - penalty\\n+ valid = (lengths >= 64) & (lengths <= 20_000) & (quality[:, 2] >= 0.45)\\n+ candidates = np.where(valid)[0]\\n+ rankings = []\\n+ for k in range(4):\\n+ score = base + 0.85 * domain[:, k]\\n+ rankings.append(candidates[np.argsort(-score[candidates], kind=\\\"stable\\\")])\\n+\\n+ chosen, seen, used = interleave(rankings, lengths)\\n+ tail_score = base + 0.85 * domain.max(1)\\n+ tail = np.argsort(-tail_score, kind=\\\"stable\\\")\\n+ chosen.extend(int(i) for i in tail if int(i) not in seen)\\n+\\n+ OUTPUT.parent.mkdir(parents=True, exist_ok=True)\\n+ with OUTPUT.open(\\\"w\\\") as f:\\n+ json.dump(chosen, f)\\n+ print(f\\\"wrote {len(chosen)} unique ids to {OUTPUT}; domain tokens {used}\\\")\\n+\\n+\\n+if __name__ == \\\"__main__\\\":\\n+ main()\\n*** End Patch\";\nconst result = await tools.apply_patch(patch);\ntext(result);"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"sed -n '85,175p' /workspace/work/make_candidates.py | nl -ba && python3 -m py_compile /workspace/submission/curate.py\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":12000});\ntext(r.output);"
}exec result
Script completed
Wall time 0.1 seconds
Output:
1 mix_hlp = np.log((dh.sum(0) + ha) / (len(dev) + ha * (HMASK + 1)))
2 hr = len(toks) / len(dev)
3 pool_hlp = np.log((ph + ha * hr) / (len(toks) + ha * hr * (HMASK + 1)))
4 hdensity_w = np.clip(mix_hlp - pool_hlp, -2.5, 2.5)
5 hdlps = np.log((dh + ha) / (250000 + ha * (HMASK + 1)))
6 hdomain_w = hdlps - hdlps.mean(0, keepdims=True)
7 density = np.zeros(len(lengths), dtype=np.float32)
8 domain = np.zeros((len(lengths), 4), dtype=np.float32)
9 bigram_density = np.zeros(len(lengths), dtype=np.float32)
10 bigram_domain = np.zeros((len(lengths), 4), dtype=np.float32)
11 lexical_domain = np.zeros((len(lengths), 4), dtype=np.float32)
12 for i, (a, b) in enumerate(zip(offsets[:-1], offsets[1:])):
13 x = toks[a:b]
14 if len(x):
15 density[i] = density_w[x].mean()
16 domain[i] = domain_w[:, x].mean(1)
17 lx = x[lexical[x]]
18 if len(lx):
19 lexical_domain[i] = domain_w[:, lx].mean(1)
20 if len(x) > 1:
21 h = bhash(x[:-1], x[1:]).astype(np.int64)
22 bigram_density[i] = hdensity_w[h].mean()
23 bigram_domain[i] = hdomain_w[:, h].mean(1)
24 density = density + .45 * bigram_density
25 domain = domain + .45 * bigram_domain
26 pen = penalties(q, lengths)
27 base = density.astype(float) - pen
28 # A soft domain confidence bonus avoids generic/boilerplate text.
29 confident = np.max(domain, axis=1) - np.partition(domain, -2, axis=1)[:, -2]
30 variants = {
31 "density": base,
32 "density_conf": base + .12 * confident,
33 "density_light": density.astype(float) - .35 * pen,
34 }
35 assigned = np.argmax(domain, axis=1)
36 # Very large crawl records are almost always whole-site concatenations; a
37 # single such record would also defeat document diversity under 12M tokens.
38 valid = (lengths >= 64) & (lengths <= 20_000) & (q[:, 2] >= .45)
39 # Crawl snapshots contain the same page under many URLs and with slightly
40 # different leading boilerplate. Deduplicate on a normalized tail fingerprint.
41 unique = np.ones(len(lengths), dtype=bool)
42 seen_fp = set()
43 with open("/workspace/data/pool.jsonl") as f:
44 for line in f:
45 r = json.loads(line)
46 s = re.sub(r"\W+", " ", r["text"][-5000:].lower()).strip()
47 fp = hashlib.blake2b(s.encode(), digest_size=12).digest()
48 if fp in seen_fp:
49 unique[r["id"]] = False
50 else:
51 seen_fp.add(fp)
52 # Keep naturally repeated pages: in this pool their frequency is useful
53 # evidence about the target register. The fingerprint is retained as a
54 # diagnostic but is not made a hard filter.
55 np.savez(W / "scores.npz", density=density, domain=domain, penalty=pen,
56 bigram_density=bigram_density, bigram_domain=bigram_domain,
57 lexical_domain=lexical_domain,
58 lengths=lengths, assigned=assigned, valid=valid)
59 for name, score in variants.items():
60 lists = []
61 for k in range(4):
62 ix = np.where(valid & (assigned == k))[0]
63 lists.append(ix[np.argsort(-score[ix], kind="stable")])
64 ids, used = interleave(lists, lengths)
65 tail = np.argsort(-score, kind="stable")
66 chosen = set(ids)
67 ids.extend(int(i) for i in tail if i not in chosen)
68 json.dump(ids, (W / f"{name}.json").open("w"))
69 print(name, used, len(ids),
70 [(k, int((assigned[np.array(ids[:12000])] == k).sum())) for k in range(4)])
71 for name, score in variants.items():
72 good = np.where(valid)[0]
73 bad = np.where(~valid)[0]
74 ids = np.concatenate((good[np.argsort(-score[good], kind="stable")],
75 bad[np.argsort(-score[bad], kind="stable")])).tolist()
76 json.dump(ids, (W / f"global_{name}.json").open("w"))
77 # Domain-specific rankings use all documents as candidates for every quota;
78 # de-duplication during interleaving assigns a document to its best use.
79 for beta in (.35, .70):
80 lists = []
81 for k in range(4):
82 score = base + beta * domain[:, k]
83 ix = np.where(valid)[0]
84 lists.append(ix[np.argsort(-score[ix], kind="stable")])
85 ids, used = interleave(lists, lengths)
86 score_any = base + beta * np.max(domain, axis=1)
87 tail = np.argsort(-score_any, kind="stable")
88 chosen = set(ids)
89 ids.extend(int(i) for i in tail if i not in chosen)
90 name = f"specialist{int(beta*100)}"
91 json.dump(ids, (W / f"{name}.json").open("w"))
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:21543,chars:\"\",yield_time_ms:30000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script running with cell ID 140
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "140",
"yield_time_ms": 20000,
"max_tokens": 5000
}wait result
Script completed
Wall time 0.0 seconds
Output:
HELDOUT_PPL 305.831 loss 5.7230 train_s 56 /workspace/work/train_pen135.npy
QUARTER 0 loss 6.5544 ppl 702.33
QUARTER 1 loss 5.5837 ppl 266.05
QUARTER 2 loss 5.742 ppl 311.69
QUARTER 3 loss 5.0091 ppl 149.77
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Update File: /workspace/work/make_variants.py\\n@@\\n-for beta in (.55, .70, .85, .95):\\n+for beta in (.55, .70, .80, .85, .88, .95):\\n*** End Patch\";\nconst result = await tools.apply_patch(patch);\ntext(result);"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"PYTHONPATH=/workspace/work python3 /workspace/work/make_variants.py | rg 'mid(80|88)'\",\"workdir\":\"/workspace\",\"yield_time_ms\":10000,\"max_output_tokens\":3000});\ntext(r.output);"
}exec result
Script completed
Wall time 10.1 seconds
Output:
mid80_w30 [3004061, 3003630, 3000026, 3000090]
mid88_w30 [3000172, 3000296, 3000251, 3000353]
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python3 /workspace/work/pack_cache.py /workspace/work/mid80_w30.json /workspace/work/train_mid80.npy && python3 /workspace/work/train_diag.py --train_npy /workspace/work/train_mid80.npy --heldout_npy /workspace/data/multi_dev.npy --out_json /workspace/work/mid80_eval.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"b017a2","wall_time_seconds":1.002179208,"session_id":3240,"original_token_count":11,"output":"packed 12000000 tokens from 16661 documents\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const patch = \"*** Begin Patch\\n*** Add File: /workspace/claim.md\\n+# Hypothesis\\n+\\n+Under a severe token budget, a small model benefits more from clean documents whose local token statistics match the intended evaluation registers than from a random sample of the crawl. A mixture explicitly balanced across encyclopedia, polished web prose, news, and technical Q&A should beat a globally ranked mixture that lets abundant news-like prose dominate.\\n+\\n+# Mechanism\\n+\\n+The selection rule estimates clipped, smoothed GPT-2 unigram and hashed-bigram density ratios between the disclosed target and the full pool. Bigram evidence captures prose/markup style; unigram evidence captures register and vocabulary. It subtracts penalties for short, non-alphabetic, unprintable, repetitive, URL-heavy, table-heavy, or whole-site records. Four domain-conditioned rankings are interleaved by consumed tokens.\\n+\\n+An observable prediction independent of final perplexity is that the consumed prefix will contain approximately 3M tokens from each conditioned ranking and will have substantially higher held-out target-density score and lower crawl-noise penalty than an equal-sized random prefix. It should also retain visibly distinct factual, article/news, and code/Q&A records instead of collapsing onto the most common register.\\n+\\n+# Falsification\\n+\\n+The hypothesis is falsified if the selected prefix does not improve the preregistered density/noise statistics over random data, if removing domain balancing improves all four per-register validation losses, or if gains disappear on a disjoint target sample while the disclosed dev sample improves. A large reversal on the hidden sample would specifically falsify the assumption that smoothed token statistics capture source-level register rather than sample-specific topics.\\n+\\n+# Transfer\\n+\\n+The method transfers to another raw pool when a modest representative calibration corpus and the intended mixture weights are available: retokenize, re-estimate target-versus-pool density ratios, and rerank. The weights and quarter boundaries are not expected to transfer to a differently weighted, multilingual, conversational, or highly specialized target; those require new calibration slices and quotas.\\n*** End Patch\";\nconst result = await tools.apply_patch(patch);\ntext(result);"
}exec result
Script completed
Wall time 0.0 seconds
Output:
{}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3240,chars:\"\",yield_time_ms:30000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script running with cell ID 145
Wall time 11.0 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"cell_id": "145",
"yield_time_ms": 20000,
"max_tokens": 5000
}wait result
Script completed
Wall time 9.5 seconds
Output:
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:3240,chars:\"\",yield_time_ms:10000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script completed
Wall time 7.5 seconds
Output:
HELDOUT_PPL 305.600 loss 5.7223 train_s 56 /workspace/work/train_mid80.npy
QUARTER 0 loss 6.5287 ppl 684.48
QUARTER 1 loss 5.598 ppl 269.88
QUARTER 2 loss 5.7539 ppl 315.42
QUARTER 3 loss 5.0063 ppl 149.36
encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.exec_command({cmd:\"python3 /workspace/work/pack_cache.py /workspace/work/mid88_w30.json /workspace/work/train_mid88.npy && python3 /workspace/work/train_diag.py --train_npy /workspace/work/train_mid88.npy --heldout_npy /workspace/data/multi_dev.npy --out_json /workspace/work/mid88_eval.json\",\"workdir\":\"/workspace\",\"yield_time_ms\":1000,\"max_output_tokens\":5000});\ntext(JSON.stringify(r));"
}exec result
Script completed
Wall time 1.2 seconds
Output:
{"chunk_id":"80cdd9","wall_time_seconds":1.001559795,"session_id":88538,"original_token_count":11,"output":"packed 12000000 tokens from 16344 documents\n"}encrypted chain-of-thought — recorded, not exposed
args
{
"input": "const r = await tools.write_stdin({session_id:88538,chars:\"\",yield_time_ms:30000,max_output_tokens:5000});\ntext(r.output);"
}exec result
Script running with cell ID 148
Wall time 11.0 seconds
Output: