claude-opus-4-8
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-4-8
- metric
- held-out multi-domain perplexity
- actions
- 99
- duration
- 1h 15m
- recorded
- 2026-08-04
run iddclm-nano-curation__wasW5uR
Everything claude-opus-4-8 did on this task, in order.
download JSON ↗Much better — junk unigrams are now e-commerce/spam words (cart, sitemap, clomid, towbars), and target unigrams reflect the dev's registers (cricket, HTML/XML tech, news). Let me verify the actual top-ranked docs are clean prose across registers, and check domain balance.
args
{
"command": "cd /workspace\npython3 -c \"\nimport json\nsel = json.load(open('/workspace/submission/selection.json'))\ntexts={}\nwith open('/workspace/data/pool.jsonl') as f:\n for line in f:\n r=json.loads(line); texts[r['id']]=r['text']\nprint('total selected:', len(sel))\nprint('=== TOP 12 ranked ===')\nfor i in sel[:12]:\n t=texts[i].strip().replace(chr(10),' ')\n print(f'[{i}] len{len(texts[i])}: {t[:150]}')\nprint()\nprint('=== around rank 8000 (near budget cutoff ~11k) ===')\nfor i in sel[8000:8006]:\n t=texts[i].strip().replace(chr(10),' ')\n print(f'[{i}]: {t[:150]}')\n\"",
"description": "Inspect new top-ranked docs"
}Bash result
total selected: 134754
=== TOP 12 ranked ===
[162538] len1423: int Up Module Stdint module Stdint: sig .. end Standard integer types type int8 Signed 8-bit integer type int16 Signed 16-bit integer type int24 Signe
[58452] len1237: <|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India
[28825] len791: WASHINGTON (Reuters) - U.S.-led forces conducted 14 air strikes against Islamic State militants in Syria and nine in Iraq on Friday, the task force co
[41027] len874: Jammu, Sep 30 (IANS) Pakistan continued to violate ceasefire along the Line of Control (LoC) in Jammu and Kashmir’s Poonch district for the seventh co
[37522] len1176: Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's national elections. "I congratulate
[49679] len567: policemen killed in Mosul bombing attack A senior police officer was killed Wednesday in northern Iraq. The first Division Chief of Nineveh Police die
[55624] len926: to local government officials, the incident took place around 9:30 pm local time leaving the district police chief Abdul Manaf dead along with two pol
[68634] len944: UPDATE: Court remands 45-year-old man for raping 15-year-old daughter A 45-year-old man, Durodola Kehinde has been remanded in prison by an Iyaganku C
[79735] len667: <|endoftext|>India’s PM Narendra Modi has tweeted pictures of his mother’s first visit to his official residence in Delhi. Twitter users have praised
[81859] len1424: |Rediff India Abroad Home | All the sections| Bihar: Vigilante justice resurfaces, three people lynched February 18, 2008 17:20 IST Fresh incidents of
[43758] len1702: vel (Maharashtra), June 14 (ANI): A magistrate's court in Panvel, Maharashtra, on Sunday extended suspended senior Nationalist Congress Party (NCP) le
[76107] len903: Major Bhupinder Singh In sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he su
=== around rank 8000 (near budget cutoff ~11k) ===
[89777]: are some exclusive pictures of Kareena Kapoor and Saif Ali Khan shooting for Yash Raj Films action film, Tashan, that will be released next year. Kare
[45375]: Delving into Foucault’s essay can give some clarity and organization to our thoughts on The Avian Gospels and the choices Adam Novy makes, and what he
[49746]: .<|endoftext|>2) The state of MI encompasses a total area of 96,716 sq miles, which makes it the 11th biggest state in the U.S. 4) The tallest point i
[14160]: Jefferson's Demons: Portrait of a Restless Mind Free Press, 2003 - 265 ページ "I have often wondered for what good end the sensations of Grief could be i
[107422]: — Moses Lake River Dogs second baseman Emmitt Tatum shook his head after Saturday’s loss to Mid-County from Nederland, Texas. “We’re not going to lose
[66247]: <|endoftext|>Icelandic weightlifter Hafþór Júlíus Björnsson, who placed third in the World’s Strongest Man competition last month, will play Gregor Cl
[stdout]
total selected: 134754
=== TOP 12 ranked ===
[162538] len1423: int Up Module Stdint module Stdint: sig .. end Standard integer types type int8 Signed 8-bit integer type int16 Signed 16-bit integer type int24 Signe
[58452] len1237: <|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India
[28825] len791: WASHINGTON (Reuters) - U.S.-led forces conducted 14 air strikes against Islamic State militants in Syria and nine in Iraq on Friday, the task force co
[41027] len874: Jammu, Sep 30 (IANS) Pakistan continued to violate ceasefire along the Line of Control (LoC) in Jammu and Kashmir’s Poonch district for the seventh co
[37522] len1176: Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's national elections. "I congratulate
[49679] len567: policemen killed in Mosul bombing attack A senior police officer was killed Wednesday in northern Iraq. The first Division Chief of Nineveh Police die
[55624] len926: to local government officials, the incident took place around 9:30 pm local time leaving the district police chief Abdul Manaf dead along with two pol
[68634] len944: UPDATE: Court remands 45-year-old man for raping 15-year-old daughter A 45-year-old man, Durodola Kehinde has been remanded in prison by an Iyaganku C
[79735] len667: <|endoftext|>India’s PM Narendra Modi has tweeted pictures of his mother’s first visit to his official residence in Delhi. Twitter users have praised
[81859] len1424: |Rediff India Abroad Home | All the sections| Bihar: Vigilante justice resurfaces, three people lynched February 18, 2008 17:20 IST Fresh incidents of
[43758] len1702: vel (Maharashtra), June 14 (ANI): A magistrate's court in Panvel, Maharashtra, on Sunday extended suspended senior Nationalist Congress Party (NCP) le
[76107] len903: Major Bhupinder Singh In sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he su
=== around rank 8000 (near budget cutoff ~11k) ===
[89777]: are some exclusive pictures of Kareena Kapoor and Saif Ali Khan shooting for Yash Raj Films action film, Tashan, that will be released next year. Kare
[45375]: Delving into Foucault’s essay can give some clarity and organization to our thoughts on The Avian Gospels and the choices Adam Novy makes, and what he
[49746]: .<|endoftext|>2) The state of MI encompasses a total area of 96,716 sq miles, which makes it the 11th biggest state in the U.S. 4) The tallest point i
[14160]: Jefferson's Demons: Portrait of a Restless Mind Free Press, 2003 - 265 ページ "I have often wondered for what good end the sensations of Grief could be i
[107422]: — Moses Lake River Dogs second baseman Emmitt Tatum shook his head after Saturday’s loss to Mid-County from Nederland, Texas. “We’re not going to lose
[66247]: <|endoftext|>Icelandic weightlifter Hafþór Júlíus Björnsson, who placed third in the World’s Strongest Man competition last month, will play Gregor Cl
[metadata] {"noOutputExpected": false}Top docs are now clean — but heavily news-skewed (lots of South-Asian news), though the #1 doc is technical (stdint docs) so tech is represented. The target is equal parts 4 registers, so balance matters. Let me launch a real eval on the current selection while I analyze register composition.
args
{
"command": "cd /workspace\nexport TOKENIZERS_PARALLELISM=false\ncp /workspace/submission/selection.json /tmp/sel_v1.json\necho \"launching dev_eval on v1 (classifier selection)...\"\nbash dev_eval.sh /tmp/sel_v1.json 2>&1 | tail -8\necho \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Run dev_eval on v1 selection",
"run_in_background": true
}Bash result
Command running in background with ID: balqi8yfn. Output is being written to: /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/balqi8yfn.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "balqi8yfn", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}Now let me analyze register composition of the budget-filling top slice while the eval trains.
args
{
"command": "cd /workspace\npython3 -c \"\nimport json, re\nsel = json.load(open('/tmp/sel_v1.json'))\ntexts={}\nwith open('/workspace/data/pool.jsonl') as f:\n for line in f:\n r=json.loads(line); texts[r['id']]=r['text']\n\ndef tokest(t): return max(1,len(t)//4) # rough token estimate\n\ndef tag(t):\n tl=t.lower()\n sym = sum(c in '<>{}=;/&|_#' for c in t)/max(1,len(t))\n codey = ('<' in t or '>' in t or '</' in t or '/>' in t or '{' in t and '}' in t\n or re.search(r'\\b(def|function|import|class|return|var|const|public|void|int|string)\\b', tl) and sym>0.01)\n if sym>0.03 or codey: return 'tech/code'\n news = bool(re.search(r'\\((reuters|ians|ani|ap|afp|pti)\\)', tl)) or bool(re.search(r'\\b(said|reported|according to|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b', tl))\n if news: return 'news'\n wiki = (' @.@ ' in t or ' @-@ ' in t) or bool(re.search(r'\\b(is|was) an? [a-z]+ (that|who|which)\\b', tl)) or bool(re.search(r'\\bborn \\d', tl))\n if wiki: return 'wiki/enc'\n return 'web-prose'\n\n# fill to 12M tokens by estimate, tag composition\nfrom collections import Counter\ncomp=Counter(); tok=0; n=0\nfor i in sel:\n comp[tag(texts[i])]+=tokest(texts[i]); tok+=tokest(texts[i]); n+=1\n if tok>=12_000_000: break\nprint(f'budget-fill: {n} docs, ~{tok/1e6:.1f}M est tokens')\nfor k,v in comp.most_common():\n print(f' {k:12s} {100*v/tok:5.1f}%')\n\"",
"description": "Estimate register composition of budget slice"
}Bash result
budget-fill: 11591 docs, ~12.0M est tokens
news 76.8%
web-prose 20.1%
tech/code 1.9%
wiki/enc 1.2%
[stdout]
budget-fill: 11591 docs, ~12.0M est tokens
news 76.8%
web-prose 20.1%
tech/code 1.9%
wiki/enc 1.2%
[metadata] {"noOutputExpected": false}Confirmed the problem: the budget slice is ~77% news, with tech/code and encyclopedic starved (target wants ~25% each). The high-entropy tech/code portion will dominate perplexity if under-covered. I'll build a register-balanced version: split the dev target into 4 registers, train one importance classifier per register, and interleave the selection so the budget fills equally. Let me check the v1 eval, then build it.
args
{
"command": "cat /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/balqi8yfn.output 2>/dev/null | tail -6; echo \"---\"; if ps aux | grep -q \"[t]rain_nano\"; then echo \"TRAINING IN PROGRESS\"; else echo \"not training yet/done\"; fi; echo \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Check v1 eval status"
}Bash result
launching dev_eval on v1 (classifier selection)...
---
TRAINING IN PROGRESS
remaining: 4258s
[stdout]
launching dev_eval on v1 (classifier selection)...
---
TRAINING IN PROGRESS
remaining: 4258s
[metadata] {"noOutputExpected": false}v1 is training. Let me copy v1 and build the register-balanced version while it runs.
args
{
"command": "cp /workspace/submission/curate.py /tmp/curate_v1.py; echo \"backed up v1 curate\"",
"description": "Back up v1 curate"
}Bash result
backed up v1 curate
[stdout]
backed up v1 curate
[metadata] {"noOutputExpected": false}v1 is training. Let me build the register-balanced version — split dev into 4 registers, one importance classifier each, round-robin fill so the budget is balanced by construction.
args
{
"file_path": "/tmp/curate_v2.py",
"content": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation (v2).\n\nSame importance-weighting idea as v1, but the disclosed target is EQUAL PARTS four\nregisters (encyclopedic / general web prose / news / technical Q&A). v1 ranked by a\nsingle global target classifier and the 12M-token budget filled ~77% news, starving\nthe technical and encyclopedic registers -> high perplexity on those quarters.\n\nv2 splits the decoded target sample into four registers, trains one importance\nclassifier per register on a shared common vocabulary, and fills the budget by\nround-robin over the four registers (always extend the register with the fewest\ntokens so far). The 12M budget is thus balanced across registers by construction,\neach register contributing its own best-matching pool documents.\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef quality_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n \"\"\"Assign a decoded dev segment to one of four registers by surface markers.\"\"\"\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n # WikiText-103 formatting: spaces around punctuation, @.@ / @-@\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\", help=\"token allocation weights tech,wiki,news,prose\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)\n # (rewrite cleanly)\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist())\n r = register_of(txt)\n reg_toks[r] += (e - s)\n w = words_of(txt)\n cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register token shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n inv_uni = {k: w for w, k in uni_ix.items()}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n log_t = np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf)\n weight[r] = np.clip(log_t - log_bg, -CLIP, CLIP)\n topw = sorted(range(N_UNI), key=lambda k: -weight[r][k])[:12]\n print(f\" [{r}] top words:\", [inv_uni[k] for k in topw], file=sys.stderr)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not quality_ok(text, w): continue\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n score[r][k] = weight[r][fa].mean()\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} passed gate\", file=sys.stderr)\n\n # dedup: keep the doc with highest max-register score per (near-)dup key\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n order_best = np.argsort(-best)\n keep = np.zeros(N, dtype=bool)\n for k in order_best:\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n # per-register ranked lists over kept docs\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc)\n o = o[keep[o]]\n ranked[r] = o\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000 # emit ~40M est-tokens worth, balanced; pipeline cuts at 12M\n while True:\n # choose register furthest below its allocation target\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n # advance past used\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]):\n continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n\n print(\" emitted register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
File created successfully at: /tmp/curate_v2.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/tmp/curate_v2.py", "content": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation (v2).\n\nSame importance-weighting idea as v1, but the disclosed target is EQUAL PARTS four\nregisters (encyclopedic / general web prose / news / technical Q&A). v1 ranked by a\nsingle global target classifier and the 12M-token budget filled ~77% news, starving\nthe technical and encyclopedic registers -> high perplexity on those quarters.\n\nv2 splits the decoded target sample into four registers, trains one importance\nclassifier per register on a shared common vocabulary, and fills the budget by\nround-robin over the four registers (always extend the register with the fewest\ntokens so far). The 12M budget is thus balanced across registers by construction,\neach register contributing its own best-matching pool documents.\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef quality_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n \"\"\"Assign a decoded dev segment to one of four registers by surface markers.\"\"\"\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n # WikiText-103 formatting: spaces around punctuation, @.@ / @-@\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\", help=\"token allocation weights tech,wiki,news,prose\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)\n # (rewrite cleanly)\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist())\n r = register_of(txt)\n reg_toks[r] += (e - s)\n w = words_of(txt)\n cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register token shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n inv_uni = {k: w for w, k in uni_ix.items()}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n log_t = np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf)\n weight[r] = np.clip(log_t - log_bg, -CLIP, CLIP)\n topw = sorted(range(N_UNI), key=lambda k: -weight[r][k])[:12]\n print(f\" [{r}] top words:\", [inv_uni[k] for k in topw], file=sys.stderr)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not quality_ok(text, w): continue\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n score[r][k] = weight[r][fa].mean()\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} passed gate\", file=sys.stderr)\n\n # dedup: keep the doc with highest max-register score per (near-)dup key\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n order_best = np.argsort(-best)\n keep = np.zeros(N, dtype=bool)\n for k in order_best:\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n # per-register ranked lists over kept docs\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc)\n o = o[keep[o]]\n ranked[r] = o\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000 # emit ~40M est-tokens worth, balanced; pipeline cuts at 12M\n while True:\n # choose register furthest below its allocation target\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n # advance past used\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]):\n continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n\n print(\" emitted register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [], "originalFile": null, "userModified": false}v1 eval finished. Let me read its result and fix a bug in v2 (a garbled bg_counts block).
args
{
"command": "echo \"=== v1 (single global classifier) eval ===\"; cat /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/balqi8yfn.output 2>/dev/null | grep -E \"HELDOUT_PPL|packed\"; echo \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Read v1 eval result"
}Bash result
=== v1 (single global classifier) eval ===
packed 12000000 tokens from 134754 ids -> /tmp/dev_train.npy
HELDOUT_PPL 388.206 loss 5.9615 train_s 56 /tmp/dev_train.npy
remaining: 4138s
[stdout]
=== v1 (single global classifier) eval ===
packed 12000000 tokens from 134754 ids -> /tmp/dev_train.npy
HELDOUT_PPL 388.206 loss 5.9615 train_s 56 /tmp/dev_train.npy
remaining: 4138s
[metadata] {"noOutputExpected": false}v1 gives PPL 388.2. I need the random baseline to calibrate. First let me fix the garbled block in v2.
args
{
"replace_all": false,
"file_path": "/tmp/curate_v2.py",
"old_string": " Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)\n # (rewrite cleanly)\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():",
"new_string": " Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():"
}Edit result
The file /tmp/curate_v2.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/tmp/curate_v2.py", "oldString": " Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)\n # (rewrite cleanly)\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():", "newString": " Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():", "originalFile": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation (v2).\n\nSame importance-weighting idea as v1, but the disclosed target is EQUAL PARTS four\nregisters (encyclopedic / general web prose / news / technical Q&A). v1 ranked by a\nsingle global target classifier and the 12M-token budget filled ~77% news, starving\nthe technical and encyclopedic registers -> high perplexity on those quarters.\n\nv2 splits the decoded target sample into four registers, trains one importance\nclassifier per register on a shared common vocabulary, and fills the budget by\nround-robin over the four registers (always extend the register with the fewest\ntokens so far). The 12M budget is thus balanced across registers by construction,\neach register contributing its own best-matching pool documents.\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef quality_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n \"\"\"Assign a decoded dev segment to one of four registers by surface markers.\"\"\"\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n # WikiText-103 formatting: spaces around punctuation, @.@ / @-@\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\", help=\"token allocation weights tech,wiki,news,prose\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)\n # (rewrite cleanly)\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist())\n r = register_of(txt)\n reg_toks[r] += (e - s)\n w = words_of(txt)\n cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register token shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n inv_uni = {k: w for w, k in uni_ix.items()}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n log_t = np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf)\n weight[r] = np.clip(log_t - log_bg, -CLIP, CLIP)\n topw = sorted(range(N_UNI), key=lambda k: -weight[r][k])[:12]\n print(f\" [{r}] top words:\", [inv_uni[k] for k in topw], file=sys.stderr)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not quality_ok(text, w): continue\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n score[r][k] = weight[r][fa].mean()\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} passed gate\", file=sys.stderr)\n\n # dedup: keep the doc with highest max-register score per (near-)dup key\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n order_best = np.argsort(-best)\n keep = np.zeros(N, dtype=bool)\n for k in order_best:\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n # per-register ranked lists over kept docs\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc)\n o = o[keep[o]]\n ranked[r] = o\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000 # emit ~40M est-tokens worth, balanced; pipeline cuts at 12M\n while True:\n # choose register furthest below its allocation target\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n # advance past used\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]):\n continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n\n print(\" emitted register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [{"oldStart": 84, "oldLines": 10, "newStart": 84, "newLines": 6, "lines": [" Vf = N_UNI + N_BI", " bg_counts = np.zeros(Vf)", " for w, c in uni_bg.items():", "- k = uni_ix.get(w); bg_counts[k] = c if k is not None else bg_counts[0]*0 + (bg_counts[k] if k is not None else 0)", "- # (rewrite cleanly)", "- bg_counts = np.zeros(Vf)", "- for w, c in uni_bg.items():", " k = uni_ix.get(w)", " if k is not None: bg_counts[k] = c", " for b, c in bi_bg.items():"]}], "userModified": false, "replaceAll": false}Now run v2 curation and the random baseline eval concurrently (v2 is CPU, baseline uses the GPU).
args
{
"command": "cd /workspace\nexport TOKENIZERS_PARALLELISM=false\npython3 /tmp/curate_v2.py --out /tmp/sel_v2.json 2>&1 | tail -40 > /tmp/v2_curate.log 2>&1\necho \"v2 curate done\"; tail -20 /tmp/v2_curate.log",
"description": "Run v2 register-balanced curation",
"run_in_background": true
}Bash result
Command running in background with ID: b4ospmk42. Output is being written to: /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/b4ospmk42.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "b4ospmk42", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace\nexport TOKENIZERS_PARALLELISM=false\necho \"launching random baseline eval...\"\nbash dev_eval.sh /tmp/random_sel.json 2>&1 | grep -E \"HELDOUT_PPL|packed\" > /tmp/random_eval.log 2>&1\necho \"random eval done\"; cat /tmp/random_eval.log",
"description": "Run random baseline eval",
"run_in_background": true
}Bash result
Command running in background with ID: bym80be5e. Output is being written to: /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/bym80be5e.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bym80be5e", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}Both launched. Let me wait for them and check the v2 diagnostics (register split + balance).
args
{
"command": "sleep 110\necho \"=== v2 curate progress ===\"; cat /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/b4ospmk42.output 2>/dev/null | tail -20\necho \"=== random eval ===\"; cat /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/bym80be5e.output 2>/dev/null | tail -3\necho \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Check v2 curate and random eval"
}Bash result
=== v2 curate progress ===
=== random eval ===
random eval done
packed 12000000 tokens from 182016 ids -> /tmp/dev_train.npy
HELDOUT_PPL 457.226 loss 6.1252 train_s 56 /tmp/dev_train.npy
remaining: 3987s
[stdout]
=== v2 curate progress ===
=== random eval ===
random eval done
packed 12000000 tokens from 182016 ids -> /tmp/dev_train.npy
HELDOUT_PPL 457.226 loss 6.1252 train_s 56 /tmp/dev_train.npy
remaining: 3987s
[metadata] {"noOutputExpected": false}Calibration established: random = 457.2, v1 (news-heavy classifier) = 388.2 (−15%). Let me wait for v2 curation to finish.
args
{
"command": "while ps aux | grep -q \"[c]urate_v2.py\"; do sleep 5; done\necho \"=== v2 curate diagnostics ===\"; cat /tmp/v2_curate.log 2>/dev/null\necho \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Wait for v2 curate, show diagnostics"
}Bash result
=== v2 curate diagnostics ===
loading pool...
building vocab + background counts...
decoding target + per-register counts...
dev register token shares: {'tech': '26%', 'wiki': '25%', 'news': '34%', 'prose': '15%'}
[tech] top words: ['p', 'code', 'pre', 'href', 'gt', 'lt', 'int', 'rel', 'aspx', 'utf', 'datetime', 'mvc']
[wiki] top words: ["'s", 'monroe', 'jupiter', 'infantry', 'pok', 'wilder', 'airborne', 'regiment', 'caesar', 'eddy', 'lennon', 'artillery']
[news] top words: ['modi', 'bjp', 'narendra', 'macron', 'yadav', 'cbi', 'heineken', 'jed', 'federer', 'pushkar', 'joshi', 'singularity']
[prose] top words: ['boundary', 'silva', 'shah', 'fuller', 'defended', 'stump', 'rowling', 'defends', 'wicket', 'rahman', 'jaime', 'amir']
scoring all docs under 4 registers...
40000/182016
80000/182016
120000/182016
160000/182016
139874/182016 passed gate
emitted register est-token shares: {'tech': '25%', 'wiki': '25%', 'news': '25%', 'prose': '25%'}
wrote 39563 ids -> /tmp/sel_v2.json
remaining: 3867s
[stdout]
=== v2 curate diagnostics ===
loading pool...
building vocab + background counts...
decoding target + per-register counts...
dev register token shares: {'tech': '26%', 'wiki': '25%', 'news': '34%', 'prose': '15%'}
[tech] top words: ['p', 'code', 'pre', 'href', 'gt', 'lt', 'int', 'rel', 'aspx', 'utf', 'datetime', 'mvc']
[wiki] top words: ["'s", 'monroe', 'jupiter', 'infantry', 'pok', 'wilder', 'airborne', 'regiment', 'caesar', 'eddy', 'lennon', 'artillery']
[news] top words: ['modi', 'bjp', 'narendra', 'macron', 'yadav', 'cbi', 'heineken', 'jed', 'federer', 'pushkar', 'joshi', 'singularity']
[prose] top words: ['boundary', 'silva', 'shah', 'fuller', 'defended', 'stump', 'rowling', 'defends', 'wicket', 'rahman', 'jaime', 'amir']
scoring all docs under 4 registers...
40000/182016
80000/182016
120000/182016
160000/182016
139874/182016 passed gate
emitted register est-token shares: {'tech': '25%', 'wiki': '25%', 'news': '25%', 'prose': '25%'}
wrote 39563 ids -> /tmp/sel_v2.json
remaining: 3867s
[metadata] {"noOutputExpected": false}v2 diagnostics are excellent — the 4 register classifiers are coherent (tech=HTML/code, wiki=historical/biographical, news=politics/sports, prose=cricket/entertainment) and the budget is balanced 25% each. Let me eval it.
args
{
"command": "cd /workspace\nexport TOKENIZERS_PARALLELISM=false\nbash dev_eval.sh /tmp/sel_v2.json 2>&1 | grep -E \"HELDOUT_PPL|packed\" > /tmp/v2_eval.log 2>&1\necho \"v2 eval done\"; cat /tmp/v2_eval.log; echo \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Run dev_eval on v2 balanced selection",
"run_in_background": true
}Bash result
Command running in background with ID: bm1j4eiak. Output is being written to: /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/bm1j4eiak.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bm1j4eiak", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}While v2 trains, let me record progress and think about a follow-up allocation variant.
args
{
"taskId": "2",
"status": "completed"
}TaskUpdate result
Updated task #2 status
[metadata] {"success": true, "taskId": "2", "updatedFields": ["status"], "statusChange": {"from": "in_progress", "to": "completed"}}encrypted chain-of-thought — recorded, not exposed
args
{
"taskId": "3",
"status": "in_progress"
}TaskUpdate result
Updated task #3 status
[metadata] {"success": true, "taskId": "3", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 120; echo \"=== v2 balanced eval ===\"; cat /tmp/v2_eval.log 2>/dev/null; if ps aux | grep -q \"[t]rain_nano\"; then echo \"(still training)\"; fi; echo \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Check v2 eval result"
}Bash result · error
Exit code 143
Command timed out after 2m 0s
[error] tool reported failurev2 eval completed. Let me read the result.
args
{
"command": "echo \"=== v2 balanced eval ===\"; cat /tmp/v2_eval.log 2>/dev/null; cat /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/bm1j4eiak.output 2>/dev/null | grep -E \"HELDOUT_PPL|packed\"\necho \"---SUMMARY---\"; echo \"random=457.2 v1(news-heavy)=388.2 v2(balanced)=?\"\necho \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Read v2 balanced eval result"
}Bash result
=== v2 balanced eval ===
packed 12000000 tokens from 39563 ids -> /tmp/dev_train.npy
HELDOUT_PPL 347.523 loss 5.8508 train_s 56 /tmp/dev_train.npy
packed 12000000 tokens from 39563 ids -> /tmp/dev_train.npy
HELDOUT_PPL 347.523 loss 5.8508 train_s 56 /tmp/dev_train.npy
---SUMMARY---
random=457.2 v1(news-heavy)=388.2 v2(balanced)=?
remaining: 3687s
[stdout]
=== v2 balanced eval ===
packed 12000000 tokens from 39563 ids -> /tmp/dev_train.npy
HELDOUT_PPL 347.523 loss 5.8508 train_s 56 /tmp/dev_train.npy
packed 12000000 tokens from 39563 ids -> /tmp/dev_train.npy
HELDOUT_PPL 347.523 loss 5.8508 train_s 56 /tmp/dev_train.npy
---SUMMARY---
random=457.2 v1(news-heavy)=388.2 v2(balanced)=?
remaining: 3687s
[metadata] {"noOutputExpected": false}v2 (balanced) = 347.5 — a clear win: −24% vs random, −10.5% vs v1. Balancing was the key lever. Now let me probe whether the high-entropy tech register wants a larger share (common bottleneck). I'll generate a tech-boosted variant while inspecting v2's per-register doc quality.
args
{
"command": "cd /workspace\nexport TOKENIZERS_PARALLELISM=false\n# tech-boosted allocation: tech 32.5%, wiki 25%, news 21.25%, prose 21.25%\npython3 /tmp/curate_v2.py --out /tmp/sel_v2t.json --alloc \"1.3,1.0,0.85,0.85\" > /tmp/v2t_curate.log 2>&1\necho \"v2-techboost curate done\"; grep -E \"emitted|wrote\" /tmp/v2t_curate.log",
"description": "Generate tech-boosted variant",
"run_in_background": true
}Bash result
Command running in background with ID: b1zodlxkf. Output is being written to: /tmp/claude-0/-workspace/06cb3775-ded4-4ad5-84cf-ed7eaa26a5bd/tasks/b1zodlxkf.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "b1zodlxkf", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace\npython3 -c \"\nimport json\nsel = json.load(open('/tmp/sel_v2.json'))\ntexts={}\nwith open('/workspace/data/pool.jsonl') as f:\n for line in f:\n r=json.loads(line); texts[r['id']]=r['text']\n# round-robin order: first 24 alternate tech,wiki,news,prose\nprint('First 24 (round-robin: tech,wiki,news,prose,...):')\nfor n,i in enumerate(sel[:24]):\n reg=['tech','wiki','news','prose'][n%4]\n t=texts[i].strip().replace(chr(10),' ')\n print(f'{reg:5s}[{i}]: {t[:120]}')\n\"",
"description": "Inspect v2 per-register top docs"
}Bash result
First 24 (round-robin: tech,wiki,news,prose,...):
tech [162538]: int Up Module Stdint module Stdint: sig .. end Standard integer types type int8 Signed 8-bit integer type int16 Signed 1
wiki [76107]: Major Bhupinder Singh In sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani f
news [58452]: <|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, includ
prose[66445]: Derrol tasty snoring residing agro que es six sigma en espanol inflexible. Niccolo que es el algebra de funciones cannib
tech [87261]: for her role as Brittany on the BET comedy-drama series The Game. She appeared in 16 episodes of the series between 2006
wiki [81859]: |Rediff India Abroad Home | All the sections| Bihar: Vigilante justice resurfaces, three people lynched February 18, 200
news [12743]: ZF-5830: Zend_Db_Table_Select doesn't allow use of $select->columns('..') Zend_Db_Table_Select doesn't allow use of $sel
prose[3475]: Anglo-Dutch Wars, also called Dutch Wars, Dutch Engelse Oorlogen, four 17th- and 18th-century naval conflicts between En
tech [180946]: <|endoftext|>bdlymmgs.com - Database Error Discuz! Database Error (1040) notconnect PHP Debug No. File Line Code 1 forum
wiki [138968]: .<|endoftext|>Boletos en para el Excite Tickets Deportes WWE Royal Rumble Boston Bruins Brooklyn Nets Dallas Cowboys Gre
news [27085]: <|endoftext|>News on : Jagan Mohan The Atmakur Civil Judge on Tuesday sent TDP MLA Erra Shekhar to 14-day judicial custo
prose[103973]: pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crow
tech [12364]: The attack was launched at 0730hrs on the 1st July 1916. Along a twenty mile Front 200,000 British and French troops att
wiki [37522]: Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's nation
news [108673]: Here are the list of players included in each of the eight playing team in IPL 2007-08. The tournament will begin on Apr
prose[165621]: Sitemap ↑<|endoftext|>Pkcs11Interop: Net.Pkcs11Interop.HighLevelAPI40.Mechanism Class Reference Pkcs11Interop 4.1.1 Net
tech [39794]: Fabian of Rome The Hieromartyr Fabian II of Rome was the bishop of Rome of the Church of Rome from 236 to 250. The remar
wiki [55624]: to local government officials, the incident took place around 9:30 pm local time leaving the district police chief Abdul
news [36908]: Thomas Robert Gore Browne |Birthplace:||Aylesbury, Buckinghamshire, England| |Death:||Died in London, England| |Managed
prose[91191]: <|endoftext|>BJP-led Goa government has decided to subsidise three more LPG cylinders for the people in addition to the
tech [41186]: Amarnath Yatra 2017 Amarnath Yatra 2017 News Jammu and Kashmir police said three people, alleged conspirators in the 10
wiki [12485]: Disable the server cache In order to disable the cache in the server and let the files be served each time, disable Page
news [178700]: fr © copyright 2011 - 2019 Avalon<|endoftext|>www.adingyuan.com - Database Error Discuz! Database Error (1146) Table 'dy
prose[119828]: Conditions.<|endoftext|>Welcome-香港赛马会主页 Top Work Portfolio Contact 香港赛马会主页 And this is Miniport, a free, fully responsiv
[stdout]
First 24 (round-robin: tech,wiki,news,prose,...):
tech [162538]: int Up Module Stdint module Stdint: sig .. end Standard integer types type int8 Signed 8-bit integer type int16 Signed 1
wiki [76107]: Major Bhupinder Singh In sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani f
news [58452]: <|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, includ
prose[66445]: Derrol tasty snoring residing agro que es six sigma en espanol inflexible. Niccolo que es el algebra de funciones cannib
tech [87261]: for her role as Brittany on the BET comedy-drama series The Game. She appeared in 16 episodes of the series between 2006
wiki [81859]: |Rediff India Abroad Home | All the sections| Bihar: Vigilante justice resurfaces, three people lynched February 18, 200
news [12743]: ZF-5830: Zend_Db_Table_Select doesn't allow use of $select->columns('..') Zend_Db_Table_Select doesn't allow use of $sel
prose[3475]: Anglo-Dutch Wars, also called Dutch Wars, Dutch Engelse Oorlogen, four 17th- and 18th-century naval conflicts between En
tech [180946]: <|endoftext|>bdlymmgs.com - Database Error Discuz! Database Error (1040) notconnect PHP Debug No. File Line Code 1 forum
wiki [138968]: .<|endoftext|>Boletos en para el Excite Tickets Deportes WWE Royal Rumble Boston Bruins Brooklyn Nets Dallas Cowboys Gre
news [27085]: <|endoftext|>News on : Jagan Mohan The Atmakur Civil Judge on Tuesday sent TDP MLA Erra Shekhar to 14-day judicial custo
prose[103973]: pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crow
tech [12364]: The attack was launched at 0730hrs on the 1st July 1916. Along a twenty mile Front 200,000 British and French troops att
wiki [37522]: Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's nation
news [108673]: Here are the list of players included in each of the eight playing team in IPL 2007-08. The tournament will begin on Apr
prose[165621]: Sitemap ↑<|endoftext|>Pkcs11Interop: Net.Pkcs11Interop.HighLevelAPI40.Mechanism Class Reference Pkcs11Interop 4.1.1 Net
tech [39794]: Fabian of Rome The Hieromartyr Fabian II of Rome was the bishop of Rome of the Church of Rome from 236 to 250. The remar
wiki [55624]: to local government officials, the incident took place around 9:30 pm local time leaving the district police chief Abdul
news [36908]: Thomas Robert Gore Browne |Birthplace:||Aylesbury, Buckinghamshire, England| |Death:||Died in London, England| |Managed
prose[91191]: <|endoftext|>BJP-led Goa government has decided to subsidise three more LPG cylinders for the people in addition to the
tech [41186]: Amarnath Yatra 2017 Amarnath Yatra 2017 News Jammu and Kashmir police said three people, alleged conspirators in the 10
wiki [12485]: Disable the server cache In order to disable the cache in the server and let the files be served each time, disable Page
news [178700]: fr © copyright 2011 - 2019 Avalon<|endoftext|>www.adingyuan.com - Database Error Discuz! Database Error (1146) Table 'dy
prose[119828]: Conditions.<|endoftext|>Welcome-香港赛马会主页 Top Work Portfolio Contact 香港赛马会主页 And this is Miniport, a free, fully responsiv
[metadata] {"noOutputExpected": false}I see real problems in v2's picks: keyword-salad SEO spam ("que es six sigma en espanol", "Boletos Excite Tickets [team names]"), foreign/CJK template junk, and error pages are leaking in. The root cause: mean-log-odds structurally favors proper-noun-dense text (matches wiki/news vocab) over natural prose. The fix is a stopword-ratio floor — natural English prose has ~25-40% function words, keyword salad has ~0% — applied to prose registers but not to the code register. Let me build v3.
args
{
"command": "cat /tmp/v2t_curate.log 2>/dev/null | grep -E \"emitted|wrote|error|Error|Traceback\" | head; echo \"(tech-boost variant status)\"; echo \"remaining: $(cat /workspace/.timer/remaining_secs)s\"",
"description": "Check tech-boost curate status"
}Bash result
(tech-boost variant status)
remaining: 3466s
[stdout]
(tech-boost variant status)
remaining: 3466s
[metadata] {"noOutputExpected": false}Let me write v3 with the fluency gates. I'll base it on v2 and add a stopword-ratio floor (prose registers) plus a non-ASCII gate.
args
{
"file_path": "/tmp/curate_v3.py",
"content": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation with fluency gating (v3).\n\nAdds to v2 two robustness gates that remove the junk v2 let through (keyword-salad\nSEO spam, foreign/CJK template pages), without harming the code register:\n\n * non-ASCII gate: drop documents that are largely non-ASCII (foreign / CJK).\n * stopword-ratio floor for the PROSE registers (wiki / news / prose): natural\n English prose contains ~25-40% function words, whereas proper-noun keyword\n salad (\"Boletos Excite Tickets Royal Rumble Boston Bruins ...\") contains almost\n none yet scores high under a mean-log-odds register classifier. Requiring a\n minimum stopword ratio keeps genuine prose and rejects the salad. The code\n register is exempt (code legitimately has few function words).\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nMAX_NONASCII_FRAC = 0.10\nPROSE_STOP_FLOOR = 0.20\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\nSTOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split()\"\"\".split())\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef general_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(1 for c in text if ord(c) > 127) / len(text) > MAX_NONASCII_FRAC: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist()); r = register_of(txt); reg_toks[r] += (e - s)\n w = words_of(txt); cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n weight[r] = np.clip(np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf) - log_bg, -CLIP, CLIP)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not general_ok(text, w): continue\n stop_frac = sum(1 for x in w if x in STOP) / len(w)\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n if r != \"tech\" and stop_frac < PROSE_STOP_FLOOR:\n continue # prose registers require fluent English\n score[r][k] = weight[r][fa].mean()\n if score[\"tech\"][k] > -1e8 or score[\"wiki\"][k] > -1e8 or \\\n score[\"news\"][k] > -1e8 or score[\"prose\"][k] > -1e8:\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} eligible for >=1 register\", file=sys.stderr)\n\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n keep = np.zeros(N, dtype=bool)\n for k in np.argsort(-best):\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc); ranked[r] = o[keep[o] & (sc[o] > -1e8)]\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000\n while True:\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]): continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n print(\" register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n print(\" register doc supply:\", {r: len(ranked[r]) for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
File created successfully at: /tmp/curate_v3.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/tmp/curate_v3.py", "content": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation with fluency gating (v3).\n\nAdds to v2 two robustness gates that remove the junk v2 let through (keyword-salad\nSEO spam, foreign/CJK template pages), without harming the code register:\n\n * non-ASCII gate: drop documents that are largely non-ASCII (foreign / CJK).\n * stopword-ratio floor for the PROSE registers (wiki / news / prose): natural\n English prose contains ~25-40% function words, whereas proper-noun keyword\n salad (\"Boletos Excite Tickets Royal Rumble Boston Bruins ...\") contains almost\n none yet scores high under a mean-log-odds register classifier. Requiring a\n minimum stopword ratio keeps genuine prose and rejects the salad. The code\n register is exempt (code legitimately has few function words).\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nMAX_NONASCII_FRAC = 0.10\nPROSE_STOP_FLOOR = 0.20\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\nSTOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split()\"\"\".split())\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef general_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(1 for c in text if ord(c) > 127) / len(text) > MAX_NONASCII_FRAC: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist()); r = register_of(txt); reg_toks[r] += (e - s)\n w = words_of(txt); cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n weight[r] = np.clip(np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf) - log_bg, -CLIP, CLIP)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not general_ok(text, w): continue\n stop_frac = sum(1 for x in w if x in STOP) / len(w)\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n if r != \"tech\" and stop_frac < PROSE_STOP_FLOOR:\n continue # prose registers require fluent English\n score[r][k] = weight[r][fa].mean()\n if score[\"tech\"][k] > -1e8 or score[\"wiki\"][k] > -1e8 or \\\n score[\"news\"][k] > -1e8 or score[\"prose\"][k] > -1e8:\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} eligible for >=1 register\", file=sys.stderr)\n\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n keep = np.zeros(N, dtype=bool)\n for k in np.argsort(-best):\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc); ranked[r] = o[keep[o] & (sc[o] > -1e8)]\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000\n while True:\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]): continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n print(\" register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n print(\" register doc supply:\", {r: len(ranked[r]) for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [], "originalFile": null, "userModified": false}I have a malformed STOP set literal (stray ".split()) inside the string). Let me fix it.
args
{
"replace_all": false,
"file_path": "/tmp/curate_v3.py",
"old_string": "STOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split())\"\"\".split())",
"new_string": "STOP = set((\n \"the of and to a in is that it for on with as was were be been being by this these those are am \"\n \"at from or an which not no but have has had they you we he she his her their its will would can could \"\n \"i my me your our us them do does did so if then than out up down about into over under after before \"\n \"all any more most some such only own same other new one two first last time year people\"\n).split())"
}Edit result · error
<tool_use_error>String to replace not found in file.
String: STOP = set("""the of and to a in is that it for on with as was were be been being by this these those are am
at from or an which not no but have has had they you we he she his her their its will would can could
i my me your our us them do does did so if then than out up down about into over under after before
all any more most some such only own same other new one two first last time year people".split())""".split())</tool_use_error>
[error] tool reported failureencrypted chain-of-thought — recorded, not exposed
args
{
"command": "sed -n '20,26p' /tmp/curate_v3.py | cat -A | head -20",
"description": "Show exact STOP lines with control chars"
}Bash result
DEV_NPY = "/workspace/data/multi_dev.npy"$
OUT = "/workspace/submission/selection.json"$
$
N_UNI, N_BI = 40000, 40000$
BG_SAMPLE = 60000$
ALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20$
WORD_RE = re.compile(r"[a-z']+")$
[stdout]
DEV_NPY = "/workspace/data/multi_dev.npy"$
OUT = "/workspace/submission/selection.json"$
$
N_UNI, N_BI = 40000, 40000$
BG_SAMPLE = 60000$
ALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20$
WORD_RE = re.compile(r"[a-z']+")$
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "grep -n 'STOP = set' /tmp/curate_v3.py; echo \"---\"; sed -n '/STOP = set/,/).split())/p' /tmp/curate_v3.py | cat -A",
"description": "Locate STOP definition lines"
}Bash result
32:STOP = set("""the of and to a in is that it for on with as was were be been being by this these those are am
---
STOP = set("""the of and to a in is that it for on with as was were be been being by this these those are am$
at from or an which not no but have has had they you we he she his her their its will would can could$
i my me your our us them do does did so if then than out up down about into over under after before$
all any more most some such only own same other new one two first last time year people".split()""".split())$
$
$
def words_of(t): return WORD_RE.findall(t.lower())$
$
$
def general_ok(text, words):$
if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False$
if sum(1 for c in text if ord(c) > 127) / len(text) > MAX_NONASCII_FRAC: return False$
if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False$
lines = [ln for ln in text.split("\n") if ln.strip()]$
if len(lines) >= 8:$
if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False$
if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False$
return True$
$
$
def register_of(txt):$
n = max(1, len(txt))$
sym = sum(c in "<>{}=;/&|_#\\`$" for c in txt)$
if sym / n > 0.02 or "<" in txt or ">" in txt or (txt.count("{") + txt.count("}") >= 4):$
return "tech"$
if " @.@ " in txt or " @-@ " in txt or (txt.count(" .") + txt.count(" ,")) > n / 250:$
return "wiki"$
tl = txt.lower()$
if re.search(r"\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\)", tl) or \$
re.search(r"\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\b", tl):$
return "news"$
return "prose"$
$
$
def main():$
ap = argparse.ArgumentParser()$
ap.add_argument("--out", default=OUT)$
ap.add_argument("--seed", type=int, default=0)$
ap.add_argument("--alloc", default="1,1,1,1")$
args = ap.parse_args()$
rng = np.random.default_rng(args.seed)$
alloc = np.array([float(x) for x in args.alloc.split(",")]); alloc /= alloc.sum()$
$
print("loading pool...", file=sys.stderr)$
ids, texts = [], []$
with open(POOL) as f:$
for line in f:$
r = json.loads(line); ids.append(r["id"]); texts.append(r["text"])$
N = len(ids)$
$
print("building vocab + background counts...", file=sys.stderr)$
bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)$
uni_bg, bi_bg = Counter(), Counter()$
for j in bg_idx:$
w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))$
uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}$
bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}$
Vf = N_UNI + N_BI$
bg_counts = np.zeros(Vf)$
for w, c in uni_bg.items():$
k = uni_ix.get(w)$
if k is not None: bg_counts[k] = c$
for b, c in bi_bg.items():$
k = bi_ix.get(b)$
if k is not None: bg_counts[k] = c$
B = bg_counts.sum()$
log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)$
$
print("decoding target + per-register counts...", file=sys.stderr)$
from transformers import AutoTokenizer$
tok = AutoTokenizer.from_pretrained("gpt2")$
dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256$
cut = np.where(dev == EOS)[0]$
segs, prev = [], 0$
for i in cut:$
if i > prev: segs.append((prev, i))$
prev = i + 1$
if prev < len(dev): segs.append((prev, len(dev)))$
reg_counts = {r: np.zeros(Vf) for r in REGISTERS}$
reg_toks = {r: 0 for r in REGISTERS}$
for s, e in segs:$
txt = tok.decode(dev[s:e].tolist()); r = register_of(txt); reg_toks[r] += (e - s)$
w = words_of(txt); cc = reg_counts[r]$
for x in w:$
k = uni_ix.get(x)$
if k is not None: cc[k] += 1$
for x in zip(w, w[1:]):$
k = bi_ix.get(x)$
if k is not None: cc[k] += 1$
tot = sum(reg_toks.values())$
print(" dev register shares:", {r: f"{100*reg_toks[r]/tot:.0f}%" for r in REGISTERS}, file=sys.stderr)$
$
weight = {}$
for r in REGISTERS:$
T = reg_counts[r].sum()$
weight[r] = np.clip(np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf) - log_bg, -CLIP, CLIP)$
$
print("scoring all docs under 4 registers...", file=sys.stderr)$
score = {r: np.full(N, -1e9) for r in REGISTERS}$
est_tok = np.zeros(N, dtype=np.int64)$
ok = np.zeros(N, dtype=bool)$
for k in range(N):$
text = texts[k]; w = words_of(text)$
est_tok[k] = len(text) // 4$
if not general_ok(text, w): continue$
stop_frac = sum(1 for x in w if x in STOP) / len(w)$
feats = [uni_ix[x] for x in w if x in uni_ix]$
prev = None$
for x in w:$
if prev is not None:$
b = bi_ix.get((prev, x))$
if b is not None: feats.append(b)$
prev = x$
if len(feats) < MIN_FEATS: continue$
fa = np.asarray(feats)$
for r in REGISTERS:$
if r != "tech" and stop_frac < PROSE_STOP_FLOOR:$
continue # prose registers require fluent English$
score[r][k] = weight[r][fa].mean()$
if score["tech"][k] > -1e8 or score["wiki"][k] > -1e8 or \$
score["news"][k] > -1e8 or score["prose"][k] > -1e8:$
ok[k] = True$
if (k + 1) % 40000 == 0: print(f" {k+1}/{N}", file=sys.stderr)$
print(f" {int(ok.sum())}/{N} eligible for >=1 register", file=sys.stderr)$
$
best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)$
ws = re.compile(r"\s+")$
keymap = {}$
keep = np.zeros(N, dtype=bool)$
for k in np.argsort(-best):$
if not ok[k]: continue$
t = texts[k]; full = hash(ws.sub(" ", t).strip())$
core = ws.sub(" ", t)[120:920]; sig = hash(core) if len(core) > 200 else full$
if full in keymap or sig in keymap: continue$
keymap[full] = 1; keymap[sig] = 1; keep[k] = True$
$
ranked = {}$
for r in REGISTERS:$
sc = score[r].copy(); sc[~keep] = -1e9$
o = np.argsort(-sc); ranked[r] = o[keep[o] & (sc[o] > -1e8)]$
ptr = {r: 0 for r in REGISTERS}$
used = np.zeros(N, dtype=bool)$
got = {r: 0 for r in REGISTERS}$
sel = []$
TARGET_TOTAL = 40_000_000$
while True:$
cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]$
if not cand: break$
r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)$
while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:$
ptr[r] += 1$
if ptr[r] >= len(ranked[r]): continue$
k = ranked[r][ptr[r]]; ptr[r] += 1$
used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])$
if sum(got.values()) >= TARGET_TOTAL: break$
print(" register est-token shares:",$
{r: f"{100*got[r]/max(1,sum(got.values())):.0f}%" for r in REGISTERS}, file=sys.stderr)$
print(" register doc supply:", {r: len(ranked[r]) for r in REGISTERS}, file=sys.stderr)$
json.dump(sel, open(args.out, "w"))$
print(f"wrote {len(sel)} ids -> {args.out}", file=sys.stderr)$
$
$
if __name__ == "__main__":$
main()$
[stdout]
32:STOP = set("""the of and to a in is that it for on with as was were be been being by this these those are am
---
STOP = set("""the of and to a in is that it for on with as was were be been being by this these those are am$
at from or an which not no but have has had they you we he she his her their its will would can could$
i my me your our us them do does did so if then than out up down about into over under after before$
all any more most some such only own same other new one two first last time year people".split()""".split())$
$
$
def words_of(t): return WORD_RE.findall(t.lower())$
$
$
def general_ok(text, words):$
if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False$
if sum(1 for c in text if ord(c) > 127) / len(text) > MAX_NONASCII_FRAC: return False$
if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False$
lines = [ln for ln in text.split("\n") if ln.strip()]$
if len(lines) >= 8:$
if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False$
if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False$
return True$
$
$
def register_of(txt):$
n = max(1, len(txt))$
sym = sum(c in "<>{}=;/&|_#\\`$" for c in txt)$
if sym / n > 0.02 or "<" in txt or ">" in txt or (txt.count("{") + txt.count("}") >= 4):$
return "tech"$
if " @.@ " in txt or " @-@ " in txt or (txt.count(" .") + txt.count(" ,")) > n / 250:$
return "wiki"$
tl = txt.lower()$
if re.search(r"\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\)", tl) or \$
re.search(r"\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\b", tl):$
return "news"$
return "prose"$
$
$
def main():$
ap = argparse.ArgumentParser()$
ap.add_argument("--out", default=OUT)$
ap.add_argument("--seed", type=int, default=0)$
ap.add_argument("--alloc", default="1,1,1,1")$
args = ap.parse_args()$
rng = np.random.default_rng(args.seed)$
alloc = np.array([float(x) for x in args.alloc.split(",")]); alloc /= alloc.sum()$
$
print("loading pool...", file=sys.stderr)$
ids, texts = [], []$
with open(POOL) as f:$
for line in f:$
r = json.loads(line); ids.append(r["id"]); texts.append(r["text"])$
N = len(ids)$
$
print("building vocab + background counts...", file=sys.stderr)$
bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)$
uni_bg, bi_bg = Counter(), Counter()$
for j in bg_idx:$
w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))$
uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}$
bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}$
Vf = N_UNI + N_BI$
bg_counts = np.zeros(Vf)$
for w, c in uni_bg.items():$
k = uni_ix.get(w)$
if k is not None: bg_counts[k] = c$
for b, c in bi_bg.items():$
k = bi_ix.get(b)$
if k is not None: bg_counts[k] = c$
B = bg_counts.sum()$
log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)$
$
print("decoding target + per-register counts...", file=sys.stderr)$
from transformers import AutoTokenizer$
tok = AutoTokenizer.from_pretrained("gpt2")$
dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256$
cut = np.where(dev == EOS)[0]$
segs, prev = [], 0$
for i in cut:$
if i > prev: segs.append((prev, i))$
prev = i + 1$
if prev < len(dev): segs.append((prev, len(dev)))$
reg_counts = {r: np.zeros(Vf) for r in REGISTERS}$
reg_toks = {r: 0 for r in REGISTERS}$
for s, e in segs:$
txt = tok.decode(dev[s:e].tolist()); r = register_of(txt); reg_toks[r] += (e - s)$
w = words_of(txt); cc = reg_counts[r]$
for x in w:$
k = uni_ix.get(x)$
if k is not None: cc[k] += 1$
for x in zip(w, w[1:]):$
k = bi_ix.get(x)$
if k is not None: cc[k] += 1$
tot = sum(reg_toks.values())$
print(" dev register shares:", {r: f"{100*reg_toks[r]/tot:.0f}%" for r in REGISTERS}, file=sys.stderr)$
$
weight = {}$
for r in REGISTERS:$
T = reg_counts[r].sum()$
weight[r] = np.clip(np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf) - log_bg, -CLIP, CLIP)$
$
print("scoring all docs under 4 registers...", file=sys.stderr)$
score = {r: np.full(N, -1e9) for r in REGISTERS}$
est_tok = np.zeros(N, dtype=np.int64)$
ok = np.zeros(N, dtype=bool)$
for k in range(N):$
text = texts[k]; w = words_of(text)$
est_tok[k] = len(text) // 4$
if not general_ok(text, w): continue$
stop_frac = sum(1 for x in w if x in STOP) / len(w)$
feats = [uni_ix[x] for x in w if x in uni_ix]$
prev = None$
for x in w:$
if prev is not None:$
b = bi_ix.get((prev, x))$
if b is not None: feats.append(b)$
prev = x$
if len(feats) < MIN_FEATS: continue$
fa = np.asarray(feats)$
for r in REGISTERS:$
if r != "tech" and stop_frac < PROSE_STOP_FLOOR:$
continue # prose registers require fluent English$
score[r][k] = weight[r][fa].mean()$
if score["tech"][k] > -1e8 or score["wiki"][k] > -1e8 or \$
score["news"][k] > -1e8 or score["prose"][k] > -1e8:$
ok[k] = True$
if (k + 1) % 40000 == 0: print(f" {k+1}/{N}", file=sys.stderr)$
print(f" {int(ok.sum())}/{N} eligible for >=1 register", file=sys.stderr)$
$
best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)$
ws = re.compile(r"\s+")$
keymap = {}$
keep = np.zeros(N, dtype=bool)$
for k in np.argsort(-best):$
if not ok[k]: continue$
t = texts[k]; full = hash(ws.sub(" ", t).strip())$
core = ws.sub(" ", t)[120:920]; sig = hash(core) if len(core) > 200 else full$
if full in keymap or sig in keymap: continue$
keymap[full] = 1; keymap[sig] = 1; keep[k] = True$
$
ranked = {}$
for r in REGISTERS:$
sc = score[r].copy(); sc[~keep] = -1e9$
o = np.argsort(-sc); ranked[r] = o[keep[o] & (sc[o] > -1e8)]$
ptr = {r: 0 for r in REGISTERS}$
used = np.zeros(N, dtype=bool)$
got = {r: 0 for r in REGISTERS}$
sel = []$
TARGET_TOTAL = 40_000_000$
while True:$
cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]$
if not cand: break$
r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)$
while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:$
ptr[r] += 1$
if ptr[r] >= len(ranked[r]): continue$
k = ranked[r][ptr[r]]; ptr[r] += 1$
used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])$
if sum(got.values()) >= TARGET_TOTAL: break$
print(" register est-token shares:",$
{r: f"{100*got[r]/max(1,sum(got.values())):.0f}%" for r in REGISTERS}, file=sys.stderr)$
print(" register doc supply:", {r: len(ranked[r]) for r in REGISTERS}, file=sys.stderr)$
json.dump(sel, open(args.out, "w"))$
print(f"wrote {len(sel)} ids -> {args.out}", file=sys.stderr)$
$
$
if __name__ == "__main__":$
main()$
[metadata] {"noOutputExpected": false}The STOP block parses (valid Python) but has a stray token and drops "people". Let me clean it precisely.
args
{
"replace_all": false,
"file_path": "/tmp/curate_v3.py",
"old_string": "STOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split()\"\"\".split())",
"new_string": "STOP = set((\n \"the of and to a in is that it for on with as was were be been being by this these those are am \"\n \"at from or an which not no but have has had they you we he she his her their its will would can could \"\n \"i my me your our us them do does did so if then than out up down about into over under after before \"\n \"all any more most some such only own same other new one two first last time year people\"\n).split())"
}Edit result
The file /tmp/curate_v3.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/tmp/curate_v3.py", "oldString": "STOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split()\"\"\".split())", "newString": "STOP = set((\n \"the of and to a in is that it for on with as was were be been being by this these those are am \"\n \"at from or an which not no but have has had they you we he she his her their its will would can could \"\n \"i my me your our us them do does did so if then than out up down about into over under after before \"\n \"all any more most some such only own same other new one two first last time year people\"\n).split())", "originalFile": "#!/usr/bin/env python3\n\"\"\"Register-balanced DSIR curation with fluency gating (v3).\n\nAdds to v2 two robustness gates that remove the junk v2 let through (keyword-salad\nSEO spam, foreign/CJK template pages), without harming the code register:\n\n * non-ASCII gate: drop documents that are largely non-ASCII (foreign / CJK).\n * stopword-ratio floor for the PROSE registers (wiki / news / prose): natural\n English prose contains ~25-40% function words, whereas proper-noun keyword\n salad (\"Boletos Excite Tickets Royal Rumble Boston Bruins ...\") contains almost\n none yet scores high under a mean-log-odds register classifier. Requiring a\n minimum stopword ratio keeps genuine prose and rejects the salad. The code\n register is exempt (code legitimately has few function words).\n\"\"\"\nimport argparse, json, re, sys, math\nfrom collections import Counter\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV_NPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nN_UNI, N_BI = 40000, 40000\nBG_SAMPLE = 60000\nALPHA, CLIP, MIN_FEATS = 1.0, 4.0, 20\nWORD_RE = re.compile(r\"[a-z']+\")\nMIN_CHARS, MIN_WORDS = 200, 60\nMAX_DIGIT_FRAC, MAX_SHORT_LINE_FRAC, MAX_DUP_LINE_FRAC = 0.15, 0.66, 0.50\nMAX_NONASCII_FRAC = 0.10\nPROSE_STOP_FLOOR = 0.20\nREGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]\nSTOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am\nat from or an which not no but have has had they you we he she his her their its will would can could\ni my me your our us them do does did so if then than out up down about into over under after before\nall any more most some such only own same other new one two first last time year people\".split()\"\"\".split())\n\n\ndef words_of(t): return WORD_RE.findall(t.lower())\n\n\ndef general_ok(text, words):\n if len(text) < MIN_CHARS or len(words) < MIN_WORDS: return False\n if sum(1 for c in text if ord(c) > 127) / len(text) > MAX_NONASCII_FRAC: return False\n if sum(c.isdigit() for c in text) / len(text) > MAX_DIGIT_FRAC: return False\n lines = [ln for ln in text.split(\"\\n\") if ln.strip()]\n if len(lines) >= 8:\n if sum(1 for ln in lines if len(ln.split()) <= 3) / len(lines) > MAX_SHORT_LINE_FRAC: return False\n if 1 - len(set(lines)) / len(lines) > MAX_DUP_LINE_FRAC: return False\n return True\n\n\ndef register_of(txt):\n n = max(1, len(txt))\n sym = sum(c in \"<>{}=;/&|_#\\\\`$\" for c in txt)\n if sym / n > 0.02 or \"<\" in txt or \">\" in txt or (txt.count(\"{\") + txt.count(\"}\") >= 4):\n return \"tech\"\n if \" @.@ \" in txt or \" @-@ \" in txt or (txt.count(\" .\") + txt.count(\" ,\")) > n / 250:\n return \"wiki\"\n tl = txt.lower()\n if re.search(r\"\\((reuters|ians|ani|ap|afp|pti|xinhua|dw)\\)\", tl) or \\\n re.search(r\"\\b(said|told reporters|according to|reported|correspondent|on (monday|tuesday|wednesday|thursday|friday|saturday|sunday))\\b\", tl):\n return \"news\"\n return \"prose\"\n\n\ndef main():\n ap = argparse.ArgumentParser()\n ap.add_argument(\"--out\", default=OUT)\n ap.add_argument(\"--seed\", type=int, default=0)\n ap.add_argument(\"--alloc\", default=\"1,1,1,1\")\n args = ap.parse_args()\n rng = np.random.default_rng(args.seed)\n alloc = np.array([float(x) for x in args.alloc.split(\",\")]); alloc /= alloc.sum()\n\n print(\"loading pool...\", file=sys.stderr)\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line); ids.append(r[\"id\"]); texts.append(r[\"text\"])\n N = len(ids)\n\n print(\"building vocab + background counts...\", file=sys.stderr)\n bg_idx = rng.choice(N, size=min(BG_SAMPLE, N), replace=False)\n uni_bg, bi_bg = Counter(), Counter()\n for j in bg_idx:\n w = words_of(texts[j]); uni_bg.update(w); bi_bg.update(zip(w, w[1:]))\n uni_ix = {w: k for k, (w, _) in enumerate(uni_bg.most_common(N_UNI))}\n bi_ix = {b: N_UNI + k for k, (b, _) in enumerate(bi_bg.most_common(N_BI))}\n Vf = N_UNI + N_BI\n bg_counts = np.zeros(Vf)\n for w, c in uni_bg.items():\n k = uni_ix.get(w)\n if k is not None: bg_counts[k] = c\n for b, c in bi_bg.items():\n k = bi_ix.get(b)\n if k is not None: bg_counts[k] = c\n B = bg_counts.sum()\n log_bg = np.log(bg_counts + ALPHA) - math.log(B + ALPHA * Vf)\n\n print(\"decoding target + per-register counts...\", file=sys.stderr)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n dev = np.load(DEV_NPY).astype(np.int64); EOS = 50256\n cut = np.where(dev == EOS)[0]\n segs, prev = [], 0\n for i in cut:\n if i > prev: segs.append((prev, i))\n prev = i + 1\n if prev < len(dev): segs.append((prev, len(dev)))\n reg_counts = {r: np.zeros(Vf) for r in REGISTERS}\n reg_toks = {r: 0 for r in REGISTERS}\n for s, e in segs:\n txt = tok.decode(dev[s:e].tolist()); r = register_of(txt); reg_toks[r] += (e - s)\n w = words_of(txt); cc = reg_counts[r]\n for x in w:\n k = uni_ix.get(x)\n if k is not None: cc[k] += 1\n for x in zip(w, w[1:]):\n k = bi_ix.get(x)\n if k is not None: cc[k] += 1\n tot = sum(reg_toks.values())\n print(\" dev register shares:\", {r: f\"{100*reg_toks[r]/tot:.0f}%\" for r in REGISTERS}, file=sys.stderr)\n\n weight = {}\n for r in REGISTERS:\n T = reg_counts[r].sum()\n weight[r] = np.clip(np.log(reg_counts[r] + ALPHA) - math.log(T + ALPHA * Vf) - log_bg, -CLIP, CLIP)\n\n print(\"scoring all docs under 4 registers...\", file=sys.stderr)\n score = {r: np.full(N, -1e9) for r in REGISTERS}\n est_tok = np.zeros(N, dtype=np.int64)\n ok = np.zeros(N, dtype=bool)\n for k in range(N):\n text = texts[k]; w = words_of(text)\n est_tok[k] = len(text) // 4\n if not general_ok(text, w): continue\n stop_frac = sum(1 for x in w if x in STOP) / len(w)\n feats = [uni_ix[x] for x in w if x in uni_ix]\n prev = None\n for x in w:\n if prev is not None:\n b = bi_ix.get((prev, x))\n if b is not None: feats.append(b)\n prev = x\n if len(feats) < MIN_FEATS: continue\n fa = np.asarray(feats)\n for r in REGISTERS:\n if r != \"tech\" and stop_frac < PROSE_STOP_FLOOR:\n continue # prose registers require fluent English\n score[r][k] = weight[r][fa].mean()\n if score[\"tech\"][k] > -1e8 or score[\"wiki\"][k] > -1e8 or \\\n score[\"news\"][k] > -1e8 or score[\"prose\"][k] > -1e8:\n ok[k] = True\n if (k + 1) % 40000 == 0: print(f\" {k+1}/{N}\", file=sys.stderr)\n print(f\" {int(ok.sum())}/{N} eligible for >=1 register\", file=sys.stderr)\n\n best = np.max(np.stack([score[r] for r in REGISTERS]), axis=0)\n ws = re.compile(r\"\\s+\")\n keymap = {}\n keep = np.zeros(N, dtype=bool)\n for k in np.argsort(-best):\n if not ok[k]: continue\n t = texts[k]; full = hash(ws.sub(\" \", t).strip())\n core = ws.sub(\" \", t)[120:920]; sig = hash(core) if len(core) > 200 else full\n if full in keymap or sig in keymap: continue\n keymap[full] = 1; keymap[sig] = 1; keep[k] = True\n\n ranked = {}\n for r in REGISTERS:\n sc = score[r].copy(); sc[~keep] = -1e9\n o = np.argsort(-sc); ranked[r] = o[keep[o] & (sc[o] > -1e8)]\n ptr = {r: 0 for r in REGISTERS}\n used = np.zeros(N, dtype=bool)\n got = {r: 0 for r in REGISTERS}\n sel = []\n TARGET_TOTAL = 40_000_000\n while True:\n cand = [r for r in REGISTERS if ptr[r] < len(ranked[r])]\n if not cand: break\n r = min(cand, key=lambda r: got[r] / alloc[REGISTERS.index(r)] if alloc[REGISTERS.index(r)] > 0 else 1e18)\n while ptr[r] < len(ranked[r]) and used[ranked[r][ptr[r]]]:\n ptr[r] += 1\n if ptr[r] >= len(ranked[r]): continue\n k = ranked[r][ptr[r]]; ptr[r] += 1\n used[k] = True; sel.append(int(ids[k])); got[r] += int(est_tok[k])\n if sum(got.values()) >= TARGET_TOTAL: break\n print(\" register est-token shares:\",\n {r: f\"{100*got[r]/max(1,sum(got.values())):.0f}%\" for r in REGISTERS}, file=sys.stderr)\n print(\" register doc supply:\", {r: len(ranked[r]) for r in REGISTERS}, file=sys.stderr)\n json.dump(sel, open(args.out, \"w\"))\n print(f\"wrote {len(sel)} ids -> {args.out}\", file=sys.stderr)\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [{"oldStart": 29, "oldLines": 10, "newStart": 29, "newLines": 12, "lines": [" MAX_NONASCII_FRAC = 0.10", " PROSE_STOP_FLOOR = 0.20", " REGISTERS = [\"tech\", \"wiki\", \"news\", \"prose\"]", "-STOP = set(\"\"\"the of and to a in is that it for on with as was were be been being by this these those are am", "-at from or an which not no but have has had they you we he she his her their its will would can could", "-i my me your our us them do does did so if then than out up down about into over under after before", "-all any more most some such only own same other new one two first last time year people\".split()\"\"\".split())", "+STOP = set((", "+ \"the of and to a in is that it for on with as was were be been being by this these those are am \"", "+ \"at from or an which not no but have has had they you we he she his her their its will would can could \"", "+ \"i my me your our us them do does did so if then than out up down about into over under after before \"", "+ \"all any more most some such only own same other new one two first last time year people\"", "+).split())", " ", " ", " def words_of(t): return WORD_RE.findall(t.lower())"]}], "userModified": false, "replaceAll": false}