claude-opus-5
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- held-out multi-domain perplexity
- actions
- 60
- duration
- 1h 8m
- recorded
- 2026-08-04
run iddclm-nano-curation__uCfhEL6
Everything claude-opus-5 did on this task, in order.
download JSON ↗encrypted chain-of-thought — recorded, not exposed
args
{
"command": "kill -9 763 $(ps -eo pid,ppid | awk '$2==763{print $1}') 2>/dev/null; sleep 1; ps -eo pid,comm|grep -c python"
}Bash result
0
[stdout]
0
[metadata] {"returnCodeInterpretation": "No matches found", "noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "python - <<'EOF'\nimport re\ns=open('submission/curate.py').read()\ns=s.replace(' return i, k, int(len(t) / 4.0) + 1, np.bincount(f, minlength=0), f',\n ' return i, k, int(len(t) / 4.0) + 1, f')\ns=s.replace(' i, k, nt, _, f = out',' i, k, nt, f = out')\ns=s.replace('''def hist(fs, n=NBUCK):\n h = np.zeros(n, dtype=np.float64)\n for f in fs:\n np.add.at(h, f, 1.0)\n return h''',\n'''def hist(fs, n=NBUCK):\n if not len(fs):\n return np.zeros(n)\n return np.bincount(np.concatenate(fs).astype(np.int64), minlength=n).astype(np.float64)''')\ns=s.replace(''' pool_h = hist(fl) + 1.0''',''' lens = np.array([len(f) for f in fl], dtype=np.int64)\n off = np.concatenate([[0], np.cumsum(lens)])\n flat = np.concatenate(fl).astype(np.int64)\n pool_h = np.bincount(flat, minlength=NBUCK).astype(np.float64) + 1.0''')\ns=s.replace(''' lens = np.array([len(f) for f in fl], dtype=np.int64)\n off = np.concatenate([[0], np.cumsum(lens)])\n flat = np.concatenate(fl).astype(np.int64)\n inv_len''',''' inv_len''')\nopen('submission/curate.py','w').write(s)\nEOF\ngrep -n \"return i, k\\|i, k, nt\\|bincount\\|inv_len = \" submission/curate.py"
}Bash result
99: return i, k, int(len(t) / 4.0) + 1, f
105: return np.bincount(np.concatenate(fs).astype(np.int64), minlength=n).astype(np.float64)
140: i, k, nt, f = out
152: pool_h = np.bincount(flat, minlength=NBUCK).astype(np.float64) + 1.0
156: inv_len = 1.0 / np.maximum(lens, 1)
[stdout]
99: return i, k, int(len(t) / 4.0) + 1, f
105: return np.bincount(np.concatenate(fs).astype(np.int64), minlength=n).astype(np.float64)
140: i, k, nt, f = out
152: pool_h = np.bincount(flat, minlength=NBUCK).astype(np.float64) + 1.0
156: inv_len = 1.0 / np.maximum(lens, 1)
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "nohup python -u submission/curate.py > /tmp/cur.log 2>&1 &\nsleep 120; grep -viE \"warning|Token indices\" /tmp/cur.log; cat .timer/remaining_secs",
"timeout": 600000
}Bash result
[1]+ Done nohup python -u submission/curate.py > /tmp/cur.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 148189
wiki score mean/max -0.581 0.394
web score mean/max -0.343 0.274
news score mean/max -0.429 0.439
qa score mean/max -0.823 0.927
selected 14732 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
4307
[stdout]
[1]+ Done nohup python -u submission/curate.py > /tmp/cur.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 148189
wiki score mean/max -0.581 0.394
web score mean/max -0.343 0.274
news score mean/max -0.429 0.439
qa score mean/max -0.823 0.927
selected 14732 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
4307
[metadata] {"noOutputExpected": false}Now evaluating it against the frozen trainer.
args
{
"command": "nohup bash dev_eval.sh > /tmp/eval1.log 2>&1 &\nsleep 30; python -c \"\nimport json\nsel=json.load(open('/workspace/submission/selection.json'))\npool={}\nimport itertools\nwant=set(sel[:8])\nfor l in open('/workspace/data/pool.jsonl'):\n r=json.loads(l)\n if r['id'] in want: pool[r['id']]=r['text']\n if len(pool)==len(want): break\nfor i in sel[:8]:\n print('---',i,repr(pool.get(i,'')[:220]))\n\"",
"timeout": 400000
}Bash result
--- 82269 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army '
--- 66445 'Derrol tasty snoring residing agro que es six sigma en espanol inflexible. Niccolo que es el algebra de funciones cannibalize pruritic, Sheila admit their cooing que choisir juin 2015 formulaire to the outside. lave ling'
--- 37064 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the vandalising of its party headquarters office in the state by the Trina'
--- 169437 "ida Urban Bikes<|endoftext|>'); setTimeout(function() { $body.css('-ms-flex-direction','column'); var headerHeight = $('header').outerHeight(true); var footerHeight = $('footer').outerHeight(true); var bodyHeight = $body"
--- 76107 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he successfully led his forces in cutting of the enemy retreat along the Ga'
--- 18830 'Archaeological dating methods\nFutilities must undress thitherward convened dismissively towerish tomato.\nPossessory tongs being very mechanically fugued calendering.\nBurdensomely tomentum consciously scavenged the differ'
--- 58452 '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India. Justices Sanjay Kishan Kaul, Navin Sinha, Mohan M Shantanagoudar, De'
--- 167171 ", dibs, shotgun\n'; var startBlockContent = '\n'; var endBlockContent = '\n'; var endBlock = '\n'; widgetContent = startBlockContent + widgetContent + endBlockContent; if (widgetTitle && true) { widgetContent = startBlockHea"
[stdout]
--- 82269 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army '
--- 66445 'Derrol tasty snoring residing agro que es six sigma en espanol inflexible. Niccolo que es el algebra de funciones cannibalize pruritic, Sheila admit their cooing que choisir juin 2015 formulaire to the outside. lave ling'
--- 37064 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the vandalising of its party headquarters office in the state by the Trina'
--- 169437 "ida Urban Bikes<|endoftext|>'); setTimeout(function() { $body.css('-ms-flex-direction','column'); var headerHeight = $('header').outerHeight(true); var footerHeight = $('footer').outerHeight(true); var bodyHeight = $body"
--- 76107 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he successfully led his forces in cutting of the enemy retreat along the Ga'
--- 18830 'Archaeological dating methods\nFutilities must undress thitherward convened dismissively towerish tomato.\nPossessory tongs being very mechanically fugued calendering.\nBurdensomely tomentum consciously scavenged the differ'
--- 58452 '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India. Justices Sanjay Kishan Kaul, Navin Sinha, Mohan M Shantanagoudar, De'
--- 167171 ", dibs, shotgun\n'; var startBlockContent = '\n'; var endBlockContent = '\n'; var endBlock = '\n'; widgetContent = startBlockContent + widgetContent + endBlockContent; if (widgetTitle && true) { widgetContent = startBlockHea"
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 150; cat /tmp/eval1.log | grep -E \"PPL|packed\"; cat .timer/remaining_secs",
"timeout": 600000
}Bash result
packed 12000000 tokens from 14732 ids -> /tmp/dev_train.npy
HELDOUT_PPL 348.117 loss 5.8525 train_s 57 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
4106
[stdout]
packed 12000000 tokens from 14732 ids -> /tmp/dev_train.npy
HELDOUT_PPL 348.117 loss 5.8525 train_s 57 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
4106
[metadata] {"noOutputExpected": false}PPL 348 vs random 458. The top-ranked docs revealed junk (SEO word-salad, JS blobs) slipping through — tightening gates and damping the rare-token reward.
args
{
"command": "python - <<'EOF'\ns=open('submission/curate.py').read()\nold_gate = s[s.index(\"def gate(t):\"):s.index(\"def dedup_key\")]\nnew_gate = '''def gate(t):\n \"\"\"Universal junk filters: length, English prose, no mojibake, no markup dumps,\n no SEO word-salad. Returns False to drop, else a code-density flag.\"\"\"\n n = len(t)\n if n < 500 or n > 400_000:\n return False\n samp = t[:20000]\n ws = re.findall(r\"[A-Za-z']+\", samp)\n if len(ws) < 80:\n return False\n lw = [w.lower() for w in ws]\n # (a) English prose: top-20 stopwords are ~25-40% of real English word tokens.\n # SEO word-salad and keyword spam sit far below this.\n if sum(w in STOP for w in lw) / len(lw) < 0.15:\n return False\n # (b) sane word lengths\n mwl = sum(len(w) for w in lw) / len(lw)\n if not (3.0 <= mwl <= 9.0):\n return False\n # (c) mostly letters/space, mostly ASCII\n if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.70:\n return False\n if sum(ord(c) > 127 for c in samp) / len(samp) > 0.08:\n return False\n # (d) not a script/markup dump: density of code punctuation\n if sum(samp.count(c) for c in \"{};=<>$|\") / len(samp) > 0.012:\n return False\n # (e) not a boilerplate list: too many repeated lines\n lines = [l.strip() for l in samp.split(\"\\\\n\") if len(l.strip()) > 15]\n if len(lines) >= 10 and len(set(lines)) / len(lines) < 0.6:\n return False\n # (f) lexical diversity floor/ceiling (ceiling catches random-word spam)\n u = len(set(lw)) / len(lw)\n if u > 0.85 and len(lw) > 200:\n return False\n return True\n\n\n'''\ns = s.replace(old_gate, new_gate)\n# damp the rare-feature reward: clip the per-feature log ratio\ns = s.replace(\" lr = np.log(th / th.sum()) - pool_p # log p_target/p_pool\",\n \" lr = np.log(th / th.sum()) - pool_p # log p_target/p_pool\\n\"\n \" np.clip(lr, -4.0, 2.0, out=lr) # damp: one rare feature can't carry a doc\")\nopen('submission/curate.py','w').write(s)\nEOF\nnohup python -u submission/curate.py > /tmp/cur2.log 2>&1 &\nsleep 100; grep -viE \"warning|Token indices\" /tmp/cur2.log",
"timeout": 600000
}Bash result
[1]+ Done nohup python -u submission/curate.py > /tmp/cur2.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 126299
wiki score mean/max -0.548 0.343
web score mean/max -0.323 0.202
news score mean/max -0.4 0.352
qa score mean/max -0.844 0.678
selected 17040 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
[stdout]
[1]+ Done nohup python -u submission/curate.py > /tmp/cur2.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 126299
wiki score mean/max -0.548 0.343
web score mean/max -0.323 0.202
news score mean/max -0.4 0.352
qa score mean/max -0.844 0.678
selected 17040 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cp submission/selection.json /tmp/sel_v2.json; nohup bash dev_eval.sh /tmp/sel_v2.json > /tmp/eval2.log 2>&1 &\nsleep 5; python - <<'EOF'\nimport json\nsel=json.load(open('/tmp/sel_v2.json'))\nwant={i:r for r,i in enumerate(sel[:40])}\ngot={}\nfor l in open('/workspace/data/pool.jsonl'):\n r=json.loads(l)\n if r['id'] in want: got[want[r['id']]]=r['text']\n if len(got)==len(want): break\nreg=['wiki','web','news','qa']\nfor r in range(12):\n print(reg[r%4], repr(got[r][:150]))\nEOF",
"timeout": 400000
}Bash result
wiki 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill'
web 'ineadóir In-athnuaite/Leabaithe a nascadh\nIt appears JavaScript is disabled. To get the most out of the website we recommend enabling JavaScript in yo'
news 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the '
qa 'YOKE – Jamavar golden fabric with the golden booti. (embroidery will not be provided)\nFLAIR – Shantoon (cotton silk) fabric with gota embellishment st'
wiki 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he su'
web 'Brisbane to London | Cruiseabout NZ\nJump to navigation\nMobile Main Menu\nHome\nDeals\nCruise Lines\nDestinations\nCruise Types\nDiscover Cruising\nFly Cruise'
news '|Rediff India Abroad Home | All the sections|\nBihar: Vigilante justice resurfaces, three people lynched\nFebruary 18, 2008 17:20 IST\nFresh incidents of'
qa 'a href="">\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\nWarning: mysql_result() [function.m'
wiki ' Conditions<|endoftext|>1st Battalion, 26th Infantry Regiment \nSubscribe Now !\nSign In Sign Out\nHome :: Military :: Agencies :: Army :: FORSCOM :: '
web '/.<|endoftext|>Townkomhs witk ‘oshn’ floor Slan locathg nhar train stations\nThe Ivy at Shrewsbury in Monmouth County, N.J., by K. Hovnanian, opened th'
news '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India'
qa 'recated: Methods with the same name as their class will not be constructors in a future version of PHP; acm has a deprecated constructor in /homepages'
[stdout]
wiki 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill'
web 'ineadóir In-athnuaite/Leabaithe a nascadh\nIt appears JavaScript is disabled. To get the most out of the website we recommend enabling JavaScript in yo'
news 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the '
qa 'YOKE – Jamavar golden fabric with the golden booti. (embroidery will not be provided)\nFLAIR – Shantoon (cotton silk) fabric with gota embellishment st'
wiki 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On the 11th Sept. he su'
web 'Brisbane to London | Cruiseabout NZ\nJump to navigation\nMobile Main Menu\nHome\nDeals\nCruise Lines\nDestinations\nCruise Types\nDiscover Cruising\nFly Cruise'
news '|Rediff India Abroad Home | All the sections|\nBihar: Vigilante justice resurfaces, three people lynched\nFebruary 18, 2008 17:20 IST\nFresh incidents of'
qa 'a href="">\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\n교육과정소개\n진행중인 교육과정\n마감된 교육과정\nWarning: mysql_result() [function.m'
wiki ' Conditions<|endoftext|>1st Battalion, 26th Infantry Regiment \nSubscribe Now !\nSign In Sign Out\nHome :: Military :: Agencies :: Army :: FORSCOM :: '
web '/.<|endoftext|>Townkomhs witk ‘oshn’ floor Slan locathg nhar train stations\nThe Ivy at Shrewsbury in Monmouth County, N.J., by K. Hovnanian, opened th'
news '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India'
qa 'recated: Methods with the same name as their class will not be constructors in a future version of PHP; acm has a deprecated constructor in /homepages'
[metadata] {"noOutputExpected": false}The naive-Bayes ratio is picking boilerplate. Switching to a discriminative classifier (GPU logistic regression), which can down-weight pool-typical boilerplate.
args
{
"command": "grep -E \"PPL\" /tmp/eval2.log; cat .timer/remaining_secs"
}Bash result
3936
[stdout]
3936
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"file_path": "/workspace/submission/curate.py",
"content": "\"\"\"Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.\n\nSTATED CRITERION\n----------------\nThe disclosed evaluation target is a BROAD four-register mix: encyclopedic\n(Wikipedia), high-quality general web prose, news, and technical Q&A. The dev\ntarget `multi_dev.npy` is exactly that mix, stored as four contiguous 250k-token\nblocks (verified by decoding it). We select documents by\n\n (1) universal junk gates, then\n (2) a per-register discriminative \"is this text drawn from the target\n register or from the raw pool?\" classifier, then\n (3) a round-robin interleave of the four per-register rankings,\n\nso that the 12M-token budget is spent roughly evenly across the four registers\ninstead of being monopolised by whichever register the pool over-represents.\n\nDetails\n-------\n1. Gates (cheap, universal, register-agnostic). Length; English-prose stopword\n rate (kills SEO word-salad and keyword spam); sane mean word length; mostly\n letters and mostly ASCII (kills mojibake / non-English); repeated-line and\n short-line-ratio limits (kills nav menus and link-farm boilerplate); a\n lexical-diversity ceiling (kills random-word spam); and near-duplicate\n removal via a word-shingle hash.\n\n2. Scoring. Every document is a hashed bag of unigrams + bigrams (2^18\n buckets), value = log(1+count), L2-normalised. For each register r we fit an\n L2-regularised logistic regression separating that register's target\n documents (positives) from a random sample of the gated pool (negatives), and\n score the whole pool with its decision function. A *discriminative* model is\n used rather than a naive-Bayes / DSIR likelihood ratio because the latter is\n dominated by rare features and happily ranks boilerplate and markup dumps at\n the top; logistic regression can assign those pool-typical n-grams negative\n weight.\n\n3. Priority order. Round-robin over the four ranked lists (best first), so\n truncating at the token budget preserves the register balance.\n\nDependencies: numpy, torch (for the LR fit/scoring), transformers (GPT-2\ntokenizer, only to decode the dev target back to text).\n\"\"\"\nimport json, re, hashlib, zlib, numpy as np, torch\nfrom multiprocessing import Pool as MPPool\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nEMIT_TOKENS = 20_000_000 # emit ~1.7x the budget worth of ids\nNBUCK = 1 << 18\nCHARCAP = 8000 # featurise only the first 8k chars of a document\nNEG = 20_000 # pool negatives per register classifier\nSEED = 0\nREGISTERS = [\"wiki\", \"web\", \"news\", \"qa\"]\n\nWORD = re.compile(r\"[a-z0-9']+|[^\\sa-z0-9]\")\nALPHA = re.compile(r\"[A-Za-z']+\")\nSTOP = set(\"the of and to in a is that it for as was on with be by are this an or from at\".split())\n\n_HC = {}\n\n\ndef _wh(w):\n h = _HC.get(w)\n if h is None:\n h = _HC[w] = zlib.crc32(w.encode()) * 2654435761 & 0xFFFFFFFF\n return h\n\n\ndef feats(text):\n \"\"\"Deduplicated hashed unigram+bigram features -> (bucket_ids, log(1+tf)).\"\"\"\n ws = WORD.findall(text[:CHARCAP].lower())\n if len(ws) < 2:\n return np.zeros(0, dtype=np.int32), np.zeros(0, dtype=np.float32)\n h = [_wh(w) for w in ws]\n b = [(h[i] * 1000003 ^ h[i + 1]) & 0xFFFFFFFF for i in range(len(h) - 1)]\n f = (np.array(h + b, dtype=np.int64) % NBUCK)\n u, c = np.unique(f, return_counts=True)\n v = np.log1p(c).astype(np.float32)\n v /= max(np.linalg.norm(v), 1e-6)\n return u.astype(np.int32), v\n\n\ndef gate(t):\n \"\"\"Universal junk filters. True = keep.\"\"\"\n n = len(t)\n if n < 500 or n > 400_000:\n return False\n samp = t[:20000]\n ws = ALPHA.findall(samp)\n if len(ws) < 80:\n return False\n lw = [w.lower() for w in ws]\n # English prose: the top-20 stopwords are ~25-40% of real English word tokens\n if sum(w in STOP for w in lw) / len(lw) < 0.15:\n return False\n mwl = sum(len(w) for w in lw) / len(lw)\n if not (3.0 <= mwl <= 9.0):\n return False\n if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.62:\n return False\n # non-ASCII checked on the head too, so a junk header can't hide in a long doc\n for chunk in (samp, t[:2000]):\n if sum(ord(c) > 127 for c in chunk) / len(chunk) > 0.08:\n return False\n lines = [l.strip() for l in samp.split(\"\\n\") if l.strip()]\n if len(lines) >= 12:\n if len(set(lines)) / len(lines) < 0.6: # repeated boilerplate\n return False\n if sum(len(l) < 40 for l in lines) / len(lines) > 0.75: # nav menus\n return False\n if len(lw) > 200 and len(set(lw)) / len(lw) > 0.85: # random-word spam\n return False\n return True\n\n\ndef dedup_key(t):\n ws = [w.lower() for w in ALPHA.findall(t)][:400]\n if len(ws) < 40:\n return None\n return hashlib.md5(\" \".join(ws[20:40]).encode()).digest()\n\n\ndef work(rec):\n i, t = rec\n if not gate(t):\n return None\n k = dedup_key(t)\n if k is None:\n return None\n u, v = feats(t)\n if len(u) == 0:\n return None\n return i, k, int(len(t) / 4.0) + 1, u, v\n\n\ndef to_csr(rows, dev=\"cuda\"):\n \"\"\"rows: list of (idx array, val array) -> torch sparse CSR on GPU.\"\"\"\n lens = np.array([len(u) for u, _ in rows], dtype=np.int64)\n crow = np.concatenate([[0], np.cumsum(lens)])\n col = np.concatenate([u for u, _ in rows]).astype(np.int64)\n val = np.concatenate([v for _, v in rows]).astype(np.float32)\n return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),\n torch.from_numpy(col).to(dev),\n torch.from_numpy(val).to(dev),\n size=(len(rows), NBUCK))\n\n\ndef fit_lr(Xp, Xn, steps=300, l2=1e-4, lr=0.5):\n \"\"\"Class-balanced logistic regression, full batch, Adam.\"\"\"\n dev = \"cuda\"\n w = torch.zeros(NBUCK, device=dev, requires_grad=True)\n b = torch.zeros(1, device=dev, requires_grad=True)\n opt = torch.optim.Adam([w, b], lr=lr)\n for _ in range(steps):\n sp = torch.sparse.mm(Xp, w.unsqueeze(1)).squeeze(1) + b\n sn = torch.sparse.mm(Xn, w.unsqueeze(1)).squeeze(1) + b\n loss = (torch.nn.functional.softplus(-sp).mean()\n + torch.nn.functional.softplus(sn).mean()) + l2 * (w * w).sum()\n opt.zero_grad(); loss.backward(); opt.step()\n with torch.no_grad():\n acc = ((torch.sparse.mm(Xp, w.unsqueeze(1)).squeeze(1) + b > 0).float().mean().item()\n + (torch.sparse.mm(Xn, w.unsqueeze(1)).squeeze(1) + b < 0).float().mean().item()) / 2\n return w.detach(), b.detach(), acc\n\n\ndef main():\n rng = np.random.default_rng(SEED)\n\n # -------------------------------------------------------- target documents\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n EOS = tok.eos_token_id\n dev_arr = np.load(DEV).astype(np.int64)\n Q = len(dev_arr) // 4\n pos = {r: [] for r in REGISTERS}\n for qi, name in enumerate(REGISTERS):\n blk = dev_arr[qi * Q:(qi + 1) * Q]\n prev = 0\n for c in list(np.flatnonzero(blk == EOS)) + [len(blk)]:\n seg = blk[prev:c]; prev = c + 1\n if len(seg) < 64:\n continue\n u, v = feats(tok.decode(seg))\n if len(u):\n pos[name].append((u, v))\n print(\"target docs:\", {r: len(v) for r, v in pos.items()}, flush=True)\n\n # ----------------------------------------------------------- pool + gating\n recs = []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n recs.append((r[\"id\"], r[\"text\"]))\n print(\"pool docs:\", len(recs), flush=True)\n\n ids, ntok, rows = [], [], []\n seen = set()\n with MPPool(16) as mp:\n for out in mp.imap(work, recs, chunksize=256):\n if out is None:\n continue\n i, k, nt, u, v = out\n if k in seen:\n continue\n seen.add(k)\n ids.append(i); ntok.append(nt); rows.append((u, v))\n del recs\n ids = np.array(ids); ntok = np.array(ntok)\n print(\"kept after gates+dedup:\", len(ids), flush=True)\n\n # ------------------------------------------------- classifiers + scoring\n Xall = to_csr(rows)\n neg_idx = rng.choice(len(rows), size=min(NEG, len(rows)), replace=False)\n Xn = to_csr([rows[i] for i in neg_idx])\n\n scores = {}\n for name in REGISTERS:\n Xp = to_csr(pos[name])\n w, b, acc = fit_lr(Xp, Xn)\n with torch.no_grad():\n s = (torch.sparse.mm(Xall, w.unsqueeze(1)).squeeze(1) + b).float().cpu().numpy()\n scores[name] = s\n print(f\"{name}: balanced train acc {acc:.3f} score mean {s.mean():.2f} \"\n f\"max {s.max():.2f}\", flush=True)\n del Xp\n\n # --------------------------------------------- per-register rank, round robin\n order = {n: np.argsort(-scores[n]) for n in REGISTERS}\n ptr = {n: 0 for n in REGISTERS}\n taken, sel, tot = set(), [], 0\n while tot < EMIT_TOKENS:\n moved = False\n for n in REGISTERS:\n o = order[n]\n while ptr[n] < len(o) and o[ptr[n]] in taken:\n ptr[n] += 1\n if ptr[n] >= len(o):\n continue\n j = int(o[ptr[n]]); ptr[n] += 1\n taken.add(j); sel.append(int(ids[j])); tot += int(ntok[j])\n moved = True\n if not moved:\n break\n\n assert len(sel) == len(set(sel))\n json.dump(sel, open(OUT, \"w\"))\n np.save(\"/tmp/curate_scores.npy\",\n np.stack([ids] + [scores[n] for n in REGISTERS]))\n print(f\"selected {len(sel)} ids, ~{tot/1e6:.1f}M est tokens \"\n f\"(budget {BUDGET/1e6:.0f}M) -> {OUT}\", flush=True)\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result · error
<tool_use_error>File has been modified since read, either by the user or by a linter. Read it again before attempting to write it.</tool_use_error>
[error] tool reported failureencrypted chain-of-thought — recorded, not exposed
args
{
"command": "grep -E \"PPL\" /tmp/eval2.log; cp submission/curate.py /tmp/v2_curate.py; echo ok"
}Bash result
ok
[stdout]
ok
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"file_path": "/workspace/submission/curate.py",
"limit": 5
}Read result
1 """Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.
2
3 STATED CRITERION
4 ----------------
5 The disclosed evaluation target is a BROAD four-register mix: encyclopedic
[metadata] {"type": "text", "file": {"filePath": "/workspace/submission/curate.py", "content": "\"\"\"Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.\n\nSTATED CRITERION\n----------------\nThe disclosed evaluation target is a BROAD four-register mix: encyclopedic", "numLines": 5, "startLine": 1, "totalLines": 216}}encrypted chain-of-thought — recorded, not exposed
args
{
"file_path": "/workspace/submission/curate.py",
"content": "\"\"\"Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.\n\nSTATED CRITERION\n----------------\nThe disclosed evaluation target is a BROAD four-register mix: encyclopedic\n(Wikipedia), high-quality general web prose, news, and technical Q&A. The dev\ntarget `multi_dev.npy` is exactly that mix, stored as four contiguous 250k-token\nblocks (verified by decoding it). We select documents by\n\n (1) universal junk gates, then\n (2) a per-register discriminative \"is this text from the target register or\n from the raw pool?\" classifier, then\n (3) a round-robin interleave of the four per-register rankings,\n\nso the 12M-token budget is spent evenly across the four registers instead of\nbeing monopolised by whichever register the pool over-represents.\n\nDetails\n-------\n1. Gates (cheap, universal, register-agnostic). Length; English-prose stopword\n rate (kills SEO word-salad and keyword spam); sane mean word length; mostly\n letters and mostly ASCII (kills mojibake / non-English); repeated-line and\n short-line-ratio limits (kill nav menus and link-farm boilerplate); a\n lexical-diversity ceiling (kills random-word spam); plus near-duplicate\n removal via a word-shingle hash.\n\n2. Scoring. Every document is a hashed bag of unigrams + bigrams (2^18\n buckets), value = log(1+count), L2-normalised. For each register r we fit an\n L2-regularised logistic regression separating that register's target\n documents (positives) from a random sample of the gated pool (negatives), and\n score the whole pool with its decision function. A *discriminative* model is\n used rather than a naive-Bayes / DSIR likelihood ratio because the latter is\n dominated by rare features and ranks boilerplate and markup dumps at the top;\n logistic regression can give those pool-typical n-grams negative weight.\n\n3. Priority order. Round-robin over the four ranked lists (best first), so\n truncating at the token budget preserves the register balance.\n\nDependencies: numpy, torch (LR fit/scoring), transformers (GPT-2 tokenizer, used\nonly to decode the dev target back into text).\n\"\"\"\nimport json, re, hashlib, zlib, numpy as np, torch\nfrom multiprocessing import Pool as MPPool\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nEMIT_TOKENS = 20_000_000 # emit ~1.7x the budget worth of ids\nNBUCK = 1 << 18\nCHARCAP = 8000 # featurise only the first 8k chars of a document\nNEG = 20_000 # pool negatives per register classifier\nSEED = 0\nREGISTERS = [\"wiki\", \"web\", \"news\", \"qa\"]\n\nWORD = re.compile(r\"[a-z0-9']+|[^\\sa-z0-9]\")\nALPHA = re.compile(r\"[A-Za-z']+\")\nSTOP = set(\"the of and to in a is that it for as was on with be by are this an or from at\".split())\n\n_HC = {}\n\n\ndef _wh(w):\n h = _HC.get(w)\n if h is None:\n h = _HC[w] = zlib.crc32(w.encode()) * 2654435761 & 0xFFFFFFFF\n return h\n\n\ndef feats(text):\n \"\"\"Deduplicated hashed unigram+bigram features -> (bucket_ids, log(1+tf)).\"\"\"\n ws = WORD.findall(text[:CHARCAP].lower())\n if len(ws) < 2:\n return np.zeros(0, dtype=np.int32), np.zeros(0, dtype=np.float32)\n h = [_wh(w) for w in ws]\n b = [(h[i] * 1000003 ^ h[i + 1]) & 0xFFFFFFFF for i in range(len(h) - 1)]\n f = np.array(h + b, dtype=np.int64) % NBUCK\n u, c = np.unique(f, return_counts=True)\n v = np.log1p(c).astype(np.float32)\n v /= max(float(np.linalg.norm(v)), 1e-6)\n return u.astype(np.int32), v\n\n\ndef gate(t):\n \"\"\"Universal junk filters. True = keep.\"\"\"\n n = len(t)\n if n < 500 or n > 400_000:\n return False\n samp = t[:20000]\n ws = ALPHA.findall(samp)\n if len(ws) < 80:\n return False\n lw = [w.lower() for w in ws]\n # English prose: the top-20 stopwords are ~25-40% of real English word tokens\n if sum(w in STOP for w in lw) / len(lw) < 0.15:\n return False\n mwl = sum(len(w) for w in lw) / len(lw)\n if not (3.0 <= mwl <= 9.0):\n return False\n if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.62:\n return False\n # non-ASCII checked on the head too, so a junk header can't hide in a long doc\n for chunk in (samp, t[:2000]):\n if sum(ord(c) > 127 for c in chunk) / len(chunk) > 0.08:\n return False\n lines = [l.strip() for l in samp.split(\"\\n\") if l.strip()]\n if len(lines) >= 12:\n if len(set(lines)) / len(lines) < 0.6: # repeated boilerplate\n return False\n if sum(len(l) < 40 for l in lines) / len(lines) > 0.75: # nav menus\n return False\n if len(lw) > 200 and len(set(lw)) / len(lw) > 0.85: # random-word spam\n return False\n return True\n\n\ndef dedup_key(t):\n ws = [w.lower() for w in ALPHA.findall(t)][:400]\n if len(ws) < 40:\n return None\n return hashlib.md5(\" \".join(ws[20:40]).encode()).digest()\n\n\ndef work(rec):\n i, t = rec\n if not gate(t):\n return None\n k = dedup_key(t)\n if k is None:\n return None\n u, v = feats(t)\n if len(u) == 0:\n return None\n return i, k, int(len(t) / 4.0) + 1, u, v\n\n\ndef to_csr(rows, dev=\"cuda\"):\n \"\"\"rows: list of (idx array, val array) -> torch sparse CSR matrix on GPU.\"\"\"\n lens = np.array([len(u) for u, _ in rows], dtype=np.int64)\n crow = np.concatenate([[0], np.cumsum(lens)])\n col = np.concatenate([u for u, _ in rows]).astype(np.int64)\n val = np.concatenate([v for _, v in rows]).astype(np.float32)\n return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),\n torch.from_numpy(col).to(dev),\n torch.from_numpy(val).to(dev),\n size=(len(rows), NBUCK))\n\n\ndef _mv(X, w, b):\n return torch.sparse.mm(X, w.unsqueeze(1)).squeeze(1) + b\n\n\ndef fit_lr(Xp, Xn, steps=300, l2=1e-4, lr=0.5):\n \"\"\"Class-balanced L2-regularised logistic regression, full batch, Adam.\"\"\"\n w = torch.zeros(NBUCK, device=\"cuda\", requires_grad=True)\n b = torch.zeros(1, device=\"cuda\", requires_grad=True)\n opt = torch.optim.Adam([w, b], lr=lr)\n F = torch.nn.functional\n for _ in range(steps):\n loss = (F.softplus(-_mv(Xp, w, b)).mean() + F.softplus(_mv(Xn, w, b)).mean()\n + l2 * (w * w).sum())\n opt.zero_grad(); loss.backward(); opt.step()\n with torch.no_grad():\n acc = 0.5 * ((_mv(Xp, w, b) > 0).float().mean().item()\n + (_mv(Xn, w, b) < 0).float().mean().item())\n return w.detach(), b.detach(), acc\n\n\ndef main():\n rng = np.random.default_rng(SEED)\n\n # -------------------------------------------------------- target documents\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n EOS = tok.eos_token_id\n dev_arr = np.load(DEV).astype(np.int64)\n Q = len(dev_arr) // 4\n pos = {r: [] for r in REGISTERS}\n for qi, name in enumerate(REGISTERS):\n blk = dev_arr[qi * Q:(qi + 1) * Q]\n prev = 0\n for c in list(np.flatnonzero(blk == EOS)) + [len(blk)]:\n seg = blk[prev:c]; prev = c + 1\n if len(seg) < 64:\n continue\n u, v = feats(tok.decode(seg))\n if len(u):\n pos[name].append((u, v))\n print(\"target docs:\", {r: len(v) for r, v in pos.items()}, flush=True)\n\n # ----------------------------------------------------------- pool + gating\n recs = []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n recs.append((r[\"id\"], r[\"text\"]))\n print(\"pool docs:\", len(recs), flush=True)\n\n ids, ntok, rows = [], [], []\n seen = set()\n with MPPool(16) as mp:\n for out in mp.imap(work, recs, chunksize=256):\n if out is None:\n continue\n i, k, nt, u, v = out\n if k in seen:\n continue\n seen.add(k)\n ids.append(i); ntok.append(nt); rows.append((u, v))\n del recs\n ids = np.array(ids); ntok = np.array(ntok)\n print(\"kept after gates+dedup:\", len(ids), flush=True)\n\n # --------------------------------------------------- classifiers + scoring\n Xall = to_csr(rows)\n neg_idx = rng.choice(len(rows), size=min(NEG, len(rows)), replace=False)\n Xn = to_csr([rows[i] for i in neg_idx])\n\n scores = {}\n for name in REGISTERS:\n Xp = to_csr(pos[name])\n w, b, acc = fit_lr(Xp, Xn)\n with torch.no_grad():\n scores[name] = _mv(Xall, w, b).float().cpu().numpy()\n print(f\"{name}: balanced train acc {acc:.3f} pool score mean \"\n f\"{scores[name].mean():.2f} max {scores[name].max():.2f}\", flush=True)\n del Xp\n\n # ------------------------------------------ per-register rank, round robin\n order = {n: np.argsort(-scores[n]) for n in REGISTERS}\n ptr = {n: 0 for n in REGISTERS}\n taken, sel, tot = set(), [], 0\n while tot < EMIT_TOKENS:\n moved = False\n for n in REGISTERS:\n o = order[n]\n while ptr[n] < len(o) and o[ptr[n]] in taken:\n ptr[n] += 1\n if ptr[n] >= len(o):\n continue\n j = int(o[ptr[n]]); ptr[n] += 1\n taken.add(j); sel.append(int(ids[j])); tot += int(ntok[j])\n moved = True\n if not moved:\n break\n\n assert len(sel) == len(set(sel))\n json.dump(sel, open(OUT, \"w\"))\n np.save(\"/tmp/curate_scores.npy\", np.stack([ids] + [scores[n] for n in REGISTERS]))\n print(f\"selected {len(sel)} ids, ~{tot/1e6:.1f}M est tokens \"\n f\"(budget {BUDGET/1e6:.0f}M) -> {OUT}\", flush=True)\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"type": "update", "filePath": "/workspace/submission/curate.py", "content": "\"\"\"Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.\n\nSTATED CRITERION\n----------------\nThe disclosed evaluation target is a BROAD four-register mix: encyclopedic\n(Wikipedia), high-quality general web prose, news, and technical Q&A. The dev\ntarget `multi_dev.npy` is exactly that mix, stored as four contiguous 250k-token\nblocks (verified by decoding it). We select documents by\n\n (1) universal junk gates, then\n (2) a per-register discriminative \"is this text from the target register or\n from the raw pool?\" classifier, then\n (3) a round-robin interleave of the four per-register rankings,\n\nso the 12M-token budget is spent evenly across the four registers instead of\nbeing monopolised by whichever register the pool over-represents.\n\nDetails\n-------\n1. Gates (cheap, universal, register-agnostic). Length; English-prose stopword\n rate (kills SEO word-salad and keyword spam); sane mean word length; mostly\n letters and mostly ASCII (kills mojibake / non-English); repeated-line and\n short-line-ratio limits (kill nav menus and link-farm boilerplate); a\n lexical-diversity ceiling (kills random-word spam); plus near-duplicate\n removal via a word-shingle hash.\n\n2. Scoring. Every document is a hashed bag of unigrams + bigrams (2^18\n buckets), value = log(1+count), L2-normalised. For each register r we fit an\n L2-regularised logistic regression separating that register's target\n documents (positives) from a random sample of the gated pool (negatives), and\n score the whole pool with its decision function. A *discriminative* model is\n used rather than a naive-Bayes / DSIR likelihood ratio because the latter is\n dominated by rare features and ranks boilerplate and markup dumps at the top;\n logistic regression can give those pool-typical n-grams negative weight.\n\n3. Priority order. Round-robin over the four ranked lists (best first), so\n truncating at the token budget preserves the register balance.\n\nDependencies: numpy, torch (LR fit/scoring), transformers (GPT-2 tokenizer, used\nonly to decode the dev target back into text).\n\"\"\"\nimport json, re, hashlib, zlib, numpy as np, torch\nfrom multiprocessing import Pool as MPPool\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nEMIT_TOKENS = 20_000_000 # emit ~1.7x the budget worth of ids\nNBUCK = 1 << 18\nCHARCAP = 8000 # featurise only the first 8k chars of a document\nNEG = 20_000 # pool negatives per register classifier\nSEED = 0\nREGISTERS = [\"wiki\", \"web\", \"news\", \"qa\"]\n\nWORD = re.compile(r\"[a-z0-9']+|[^\\sa-z0-9]\")\nALPHA = re.compile(r\"[A-Za-z']+\")\nSTOP = set(\"the of and to in a is that it for as was on with be by are this an or from at\".split())\n\n_HC = {}\n\n\ndef _wh(w):\n h = _HC.get(w)\n if h is None:\n h = _HC[w] = zlib.crc32(w.encode()) * 2654435761 & 0xFFFFFFFF\n return h\n\n\ndef feats(text):\n \"\"\"Deduplicated hashed unigram+bigram features -> (bucket_ids, log(1+tf)).\"\"\"\n ws = WORD.findall(text[:CHARCAP].lower())\n if len(ws) < 2:\n return np.zeros(0, dtype=np.int32), np.zeros(0, dtype=np.float32)\n h = [_wh(w) for w in ws]\n b = [(h[i] * 1000003 ^ h[i + 1]) & 0xFFFFFFFF for i in range(len(h) - 1)]\n f = np.array(h + b, dtype=np.int64) % NBUCK\n u, c = np.unique(f, return_counts=True)\n v = np.log1p(c).astype(np.float32)\n v /= max(float(np.linalg.norm(v)), 1e-6)\n return u.astype(np.int32), v\n\n\ndef gate(t):\n \"\"\"Universal junk filters. True = keep.\"\"\"\n n = len(t)\n if n < 500 or n > 400_000:\n return False\n samp = t[:20000]\n ws = ALPHA.findall(samp)\n if len(ws) < 80:\n return False\n lw = [w.lower() for w in ws]\n # English prose: the top-20 stopwords are ~25-40% of real English word tokens\n if sum(w in STOP for w in lw) / len(lw) < 0.15:\n return False\n mwl = sum(len(w) for w in lw) / len(lw)\n if not (3.0 <= mwl <= 9.0):\n return False\n if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.62:\n return False\n # non-ASCII checked on the head too, so a junk header can't hide in a long doc\n for chunk in (samp, t[:2000]):\n if sum(ord(c) > 127 for c in chunk) / len(chunk) > 0.08:\n return False\n lines = [l.strip() for l in samp.split(\"\\n\") if l.strip()]\n if len(lines) >= 12:\n if len(set(lines)) / len(lines) < 0.6: # repeated boilerplate\n return False\n if sum(len(l) < 40 for l in lines) / len(lines) > 0.75: # nav menus\n return False\n if len(lw) > 200 and len(set(lw)) / len(lw) > 0.85: # random-word spam\n return False\n return True\n\n\ndef dedup_key(t):\n ws = [w.lower() for w in ALPHA.findall(t)][:400]\n if len(ws) < 40:\n return None\n return hashlib.md5(\" \".join(ws[20:40]).encode()).digest()\n\n\ndef work(rec):\n i, t = rec\n if not gate(t):\n return None\n k = dedup_key(t)\n if k is None:\n return None\n u, v = feats(t)\n if len(u) == 0:\n return None\n return i, k, int(len(t) / 4.0) + 1, u, v\n\n\ndef to_csr(rows, dev=\"cuda\"):\n \"\"\"rows: list of (idx array, val array) -> torch sparse CSR matrix on GPU.\"\"\"\n lens = np.array([len(u) for u, _ in rows], dtype=np.int64)\n crow = np.concatenate([[0], np.cumsum(lens)])\n col = np.concatenate([u for u, _ in rows]).astype(np.int64)\n val = np.concatenate([v for _, v in rows]).astype(np.float32)\n return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),\n torch.from_numpy(col).to(dev),\n torch.from_numpy(val).to(dev),\n size=(len(rows), NBUCK))\n\n\ndef _mv(X, w, b):\n return torch.sparse.mm(X, w.unsqueeze(1)).squeeze(1) + b\n\n\ndef fit_lr(Xp, Xn, steps=300, l2=1e-4, lr=0.5):\n \"\"\"Class-balanced L2-regularised logistic regression, full batch, Adam.\"\"\"\n w = torch.zeros(NBUCK, device=\"cuda\", requires_grad=True)\n b = torch.zeros(1, device=\"cuda\", requires_grad=True)\n opt = torch.optim.Adam([w, b], lr=lr)\n F = torch.nn.functional\n for _ in range(steps):\n loss = (F.softplus(-_mv(Xp, w, b)).mean() + F.softplus(_mv(Xn, w, b)).mean()\n + l2 * (w * w).sum())\n opt.zero_grad(); loss.backward(); opt.step()\n with torch.no_grad():\n acc = 0.5 * ((_mv(Xp, w, b) > 0).float().mean().item()\n + (_mv(Xn, w, b) < 0).float().mean().item())\n return w.detach(), b.detach(), acc\n\n\ndef main():\n rng = np.random.default_rng(SEED)\n\n # -------------------------------------------------------- target documents\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n EOS = tok.eos_token_id\n dev_arr = np.load(DEV).astype(np.int64)\n Q = len(dev_arr) // 4\n pos = {r: [] for r in REGISTERS}\n for qi, name in enumerate(REGISTERS):\n blk = dev_arr[qi * Q:(qi + 1) * Q]\n prev = 0\n for c in list(np.flatnonzero(blk == EOS)) + [len(blk)]:\n seg = blk[prev:c]; prev = c + 1\n if len(seg) < 64:\n continue\n u, v = feats(tok.decode(seg))\n if len(u):\n pos[name].append((u, v))\n print(\"target docs:\", {r: len(v) for r, v in pos.items()}, flush=True)\n\n # ----------------------------------------------------------- pool + gating\n recs = []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n recs.append((r[\"id\"], r[\"text\"]))\n print(\"pool docs:\", len(recs), flush=True)\n\n ids, ntok, rows = [], [], []\n seen = set()\n with MPPool(16) as mp:\n for out in mp.imap(work, recs, chunksize=256):\n if out is None:\n continue\n i, k, nt, u, v = out\n if k in seen:\n continue\n seen.add(k)\n ids.append(i); ntok.append(nt); rows.append((u, v))\n del recs\n ids = np.array(ids); ntok = np.array(ntok)\n print(\"kept after gates+dedup:\", len(ids), flush=True)\n\n # --------------------------------------------------- classifiers + scoring\n Xall = to_csr(rows)\n neg_idx = rng.choice(len(rows), size=min(NEG, len(rows)), replace=False)\n Xn = to_csr([rows[i] for i in neg_idx])\n\n scores = {}\n for name in REGISTERS:\n Xp = to_csr(pos[name])\n w, b, acc = fit_lr(Xp, Xn)\n with torch.no_grad():\n scores[name] = _mv(Xall, w, b).float().cpu().numpy()\n print(f\"{name}: balanced train acc {acc:.3f} pool score mean \"\n f\"{scores[name].mean():.2f} max {scores[name].max():.2f}\", flush=True)\n del Xp\n\n # ------------------------------------------ per-register rank, round robin\n order = {n: np.argsort(-scores[n]) for n in REGISTERS}\n ptr = {n: 0 for n in REGISTERS}\n taken, sel, tot = set(), [], 0\n while tot < EMIT_TOKENS:\n moved = False\n for n in REGISTERS:\n o = order[n]\n while ptr[n] < len(o) and o[ptr[n]] in taken:\n ptr[n] += 1\n if ptr[n] >= len(o):\n continue\n j = int(o[ptr[n]]); ptr[n] += 1\n taken.add(j); sel.append(int(ids[j])); tot += int(ntok[j])\n moved = True\n if not moved:\n break\n\n assert len(sel) == len(set(sel))\n json.dump(sel, open(OUT, \"w\"))\n np.save(\"/tmp/curate_scores.npy\", np.stack([ids] + [scores[n] for n in REGISTERS]))\n print(f\"selected {len(sel)} ids, ~{tot/1e6:.1f}M est tokens \"\n f\"(budget {BUDGET/1e6:.0f}M) -> {OUT}\", flush=True)\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [{"oldStart": 5, "oldLines": 27, "newStart": 5, "newLines": 41, "lines": [" The disclosed evaluation target is a BROAD four-register mix: encyclopedic", " (Wikipedia), high-quality general web prose, news, and technical Q&A. The dev", " target `multi_dev.npy` is exactly that mix, stored as four contiguous 250k-token", "-blocks (verified by decoding). We select pool documents by *importance", "-resampling against that target distribution* (DSIR, Xie et al. 2023), done", "-per-register so the budget is spent evenly across all four registers:", "+blocks (verified by decoding it). We select documents by", " ", "- 1. Decode `multi_dev.npy`, split on <|endoftext|>, label every target document", "- by which quarter (register) it came from.", "- 2. Universal quality gates on the pool: length, English-prose stopword rate,", "- alphabetic/non-ASCII ratios, and near-duplicate removal via a word-shingle", "- hash. These remove junk that no register wants.", "- 3. Represent every document as a hashed bag of unigrams+bigrams (2^18 buckets).", "- Fit a unigram categorical distribution p_r over features for each target", "- register r, and p_pool over the gated pool.", "- Score(d, r) = (1/|d|) * sum_f c_f(d) * log( p_r[f] / p_pool[f] )", "- i.e. the length-normalised log importance weight of d under register r.", "- 4. Rank the pool by Score(., r) for each r and emit a round-robin interleave of", "- the four ranked lists. Priority order is therefore best-first *and*", "- register-balanced, so truncating at the 12M-token budget keeps the mix.", "+ (1) universal junk gates, then", "+ (2) a per-register discriminative \"is this text from the target register or", "+ from the raw pool?\" classifier, then", "+ (3) a round-robin interleave of the four per-register rankings,", " ", "-Only numpy + the GPT-2 tokenizer are required.", "+so the 12M-token budget is spent evenly across the four registers instead of", "+being monopolised by whichever register the pool over-represents.", "+", "+Details", "+-------", "+1. Gates (cheap, universal, register-agnostic). Length; English-prose stopword", "+ rate (kills SEO word-salad and keyword spam); sane mean word length; mostly", "+ letters and mostly ASCII (kills mojibake / non-English); repeated-line and", "+ short-line-ratio limits (kill nav menus and link-farm boilerplate); a", "+ lexical-diversity ceiling (kills random-word spam); plus near-duplicate", "+ removal via a word-shingle hash.", "+", "+2. Scoring. Every document is a hashed bag of unigrams + bigrams (2^18", "+ buckets), value = log(1+count), L2-normalised. For each register r we fit an", "+ L2-regularised logistic regression separating that register's target", "+ documents (positives) from a random sample of the gated pool (negatives), and", "+ score the whole pool with its decision function. A *discriminative* model is", "+ used rather than a naive-Bayes / DSIR likelihood ratio because the latter is", "+ dominated by rare features and ranks boilerplate and markup dumps at the top;", "+ logistic regression can give those pool-typical n-grams negative weight.", "+", "+3. Priority order. Round-robin over the four ranked lists (best first), so", "+ truncating at the token budget preserves the register balance.", "+", "+Dependencies: numpy, torch (LR fit/scoring), transformers (GPT-2 tokenizer, used", "+only to decode the dev target back into text).", " \"\"\"", "-import json, re, hashlib, zlib, numpy as np", "+import json, re, hashlib, zlib, numpy as np, torch", " from multiprocessing import Pool as MPPool", " from transformers import AutoTokenizer", " "]}, {"oldStart": 33, "oldLines": 20, "newStart": 47, "newLines": 21, "lines": [" DEV = \"/workspace/data/multi_dev.npy\"", " OUT = \"/workspace/submission/selection.json\"", " BUDGET = 12_000_000", "-EMIT_TOKENS = 20_000_000 # emit ~1.7x budget worth of ids", "+EMIT_TOKENS = 20_000_000 # emit ~1.7x the budget worth of ids", " NBUCK = 1 << 18", "-CHARCAP = 8000 # only the first 8k chars of a doc are featurised", "+CHARCAP = 8000 # featurise only the first 8k chars of a document", "+NEG = 20_000 # pool negatives per register classifier", "+SEED = 0", " REGISTERS = [\"wiki\", \"web\", \"news\", \"qa\"]", " ", " WORD = re.compile(r\"[a-z0-9']+|[^\\sa-z0-9]\")", "+ALPHA = re.compile(r\"[A-Za-z']+\")", " STOP = set(\"the of and to in a is that it for as was on with be by are this an or from at\".split())", " ", "-", " _HC = {}", " ", " ", " def _wh(w):", "- \"\"\"Deterministic (run-independent) word hash, memoised.\"\"\"", " h = _HC.get(w)", " if h is None:", " h = _HC[w] = zlib.crc32(w.encode()) * 2654435761 & 0xFFFFFFFF"]}, {"oldStart": 54, "oldLines": 97, "newStart": 69, "newLines": 126, "lines": [" ", " ", " def feats(text):", "- \"\"\"Hashed unigram+bigram bucket ids for a document.\"\"\"", "+ \"\"\"Deduplicated hashed unigram+bigram features -> (bucket_ids, log(1+tf)).\"\"\"", " ws = WORD.findall(text[:CHARCAP].lower())", "- if not ws:", "- return np.zeros(0, dtype=np.int32)", "+ if len(ws) < 2:", "+ return np.zeros(0, dtype=np.int32), np.zeros(0, dtype=np.float32)", " h = [_wh(w) for w in ws]", " b = [(h[i] * 1000003 ^ h[i + 1]) & 0xFFFFFFFF for i in range(len(h) - 1)]", "- return (np.array(h + b, dtype=np.int64) % NBUCK).astype(np.int32)", "+ f = np.array(h + b, dtype=np.int64) % NBUCK", "+ u, c = np.unique(f, return_counts=True)", "+ v = np.log1p(c).astype(np.float32)", "+ v /= max(float(np.linalg.norm(v)), 1e-6)", "+ return u.astype(np.int32), v", " ", " ", " def gate(t):", "- \"\"\"Universal junk filters: length, English prose, no mojibake, no markup dumps,", "- no SEO word-salad. Returns False to drop, else a code-density flag.\"\"\"", "+ \"\"\"Universal junk filters. True = keep.\"\"\"", " n = len(t)", " if n < 500 or n > 400_000:", " return False", " samp = t[:20000]", "- ws = re.findall(r\"[A-Za-z']+\", samp)", "+ ws = ALPHA.findall(samp)", " if len(ws) < 80:", " return False", " lw = [w.lower() for w in ws]", "- # (a) English prose: top-20 stopwords are ~25-40% of real English word tokens.", "- # SEO word-salad and keyword spam sit far below this.", "+ # English prose: the top-20 stopwords are ~25-40% of real English word tokens", " if sum(w in STOP for w in lw) / len(lw) < 0.15:", " return False", "- # (b) sane word lengths", " mwl = sum(len(w) for w in lw) / len(lw)", " if not (3.0 <= mwl <= 9.0):", " return False", "- # (c) mostly letters/space, mostly ASCII", "- if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.70:", "+ if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.62:", " return False", "- if sum(ord(c) > 127 for c in samp) / len(samp) > 0.08:", "+ # non-ASCII checked on the head too, so a junk header can't hide in a long doc", "+ for chunk in (samp, t[:2000]):", "+ if sum(ord(c) > 127 for c in chunk) / len(chunk) > 0.08:", "+ return False", "+ lines = [l.strip() for l in samp.split(\"\\n\") if l.strip()]", "+ if len(lines) >= 12:", "+ if len(set(lines)) / len(lines) < 0.6: # repeated boilerplate", "+ return False", "+ if sum(len(l) < 40 for l in lines) / len(lines) > 0.75: # nav menus", "+ return False", "+ if len(lw) > 200 and len(set(lw)) / len(lw) > 0.85: # random-word spam", " return False", "- # (d) not a script/markup dump: density of code punctuation", "- if sum(samp.count(c) for c in \"{};=<>$|\") / len(samp) > 0.012:", "- return False", "- # (e) not a boilerplate list: too many repeated lines", "- lines = [l.strip() for l in samp.split(\"\\n\") if len(l.strip()) > 15]", "- if len(lines) >= 10 and len(set(lines)) / len(lines) < 0.6:", "- return False", "- # (f) lexical diversity floor/ceiling (ceiling catches random-word spam)", "- u = len(set(lw)) / len(lw)", "- if u > 0.85 and len(lw) > 200:", "- return False", " return True", " ", " ", " def dedup_key(t):", "- ws = [w.lower() for w in re.findall(r\"[A-Za-z']+\", t)][:400]", "+ ws = [w.lower() for w in ALPHA.findall(t)][:400]", " if len(ws) < 40:", " return None", " return hashlib.md5(\" \".join(ws[20:40]).encode()).digest()", " ", " ", " def work(rec):", "- \"\"\"Per-document worker: gate, dedup key, token estimate, hashed features.\"\"\"", " i, t = rec", " if not gate(t):", " return None", " k = dedup_key(t)", " if k is None:", " return None", "- f = feats(t)", "- return i, k, int(len(t) / 4.0) + 1, f", "+ u, v = feats(t)", "+ if len(u) == 0:", "+ return None", "+ return i, k, int(len(t) / 4.0) + 1, u, v", " ", " ", "-def hist(fs, n=NBUCK):", "- if not len(fs):", "- return np.zeros(n)", "- return np.bincount(np.concatenate(fs).astype(np.int64), minlength=n).astype(np.float64)", "+def to_csr(rows, dev=\"cuda\"):", "+ \"\"\"rows: list of (idx array, val array) -> torch sparse CSR matrix on GPU.\"\"\"", "+ lens = np.array([len(u) for u, _ in rows], dtype=np.int64)", "+ crow = np.concatenate([[0], np.cumsum(lens)])", "+ col = np.concatenate([u for u, _ in rows]).astype(np.int64)", "+ val = np.concatenate([v for _, v in rows]).astype(np.float32)", "+ return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),", "+ torch.from_numpy(col).to(dev),", "+ torch.from_numpy(val).to(dev),", "+ size=(len(rows), NBUCK))", " ", " ", "+def _mv(X, w, b):", "+ return torch.sparse.mm(X, w.unsqueeze(1)).squeeze(1) + b", "+", "+", "+def fit_lr(Xp, Xn, steps=300, l2=1e-4, lr=0.5):", "+ \"\"\"Class-balanced L2-regularised logistic regression, full batch, Adam.\"\"\"", "+ w = torch.zeros(NBUCK, device=\"cuda\", requires_grad=True)", "+ b = torch.zeros(1, device=\"cuda\", requires_grad=True)", "+ opt = torch.optim.Adam([w, b], lr=lr)", "+ F = torch.nn.functional", "+ for _ in range(steps):", "+ loss = (F.softplus(-_mv(Xp, w, b)).mean() + F.softplus(_mv(Xn, w, b)).mean()", "+ + l2 * (w * w).sum())", "+ opt.zero_grad(); loss.backward(); opt.step()", "+ with torch.no_grad():", "+ acc = 0.5 * ((_mv(Xp, w, b) > 0).float().mean().item()", "+ + (_mv(Xn, w, b) < 0).float().mean().item())", "+ return w.detach(), b.detach(), acc", "+", "+", " def main():", "- # ------------------------------------------------------------- target text", "+ rng = np.random.default_rng(SEED)", "+", "+ # -------------------------------------------------------- target documents", " tok = AutoTokenizer.from_pretrained(\"gpt2\")", " EOS = tok.eos_token_id", "- dev = np.load(DEV).astype(np.int64)", "- Q = len(dev) // 4", "- pos_feats = {r: [] for r in REGISTERS}", "+ dev_arr = np.load(DEV).astype(np.int64)", "+ Q = len(dev_arr) // 4", "+ pos = {r: [] for r in REGISTERS}", " for qi, name in enumerate(REGISTERS):", "- blk = dev[qi * Q:(qi + 1) * Q]", "- cuts = list(np.flatnonzero(blk == EOS)) + [len(blk)]", "+ blk = dev_arr[qi * Q:(qi + 1) * Q]", " prev = 0", "- for c in cuts:", "+ for c in list(np.flatnonzero(blk == EOS)) + [len(blk)]:", " seg = blk[prev:c]; prev = c + 1", " if len(seg) < 64:", " continue", "- pos_feats[name].append(feats(tok.decode(seg)))", "- print(\"target docs:\", {r: len(v) for r, v in pos_feats.items()}, flush=True)", "+ u, v = feats(tok.decode(seg))", "+ if len(u):", "+ pos[name].append((u, v))", "+ print(\"target docs:\", {r: len(v) for r, v in pos.items()}, flush=True)", " ", "- # ------------------------------------------------------- pool load + gate", "+ # ----------------------------------------------------------- pool + gating", " recs = []", " with open(POOL) as f:", " for line in f:"]}, {"oldStart": 152, "oldLines": 42, "newStart": 196, "newLines": 37, "lines": [" recs.append((r[\"id\"], r[\"text\"]))", " print(\"pool docs:\", len(recs), flush=True)", " ", "- ids, ntok, fl = [], [], []", "+ ids, ntok, rows = [], [], []", " seen = set()", " with MPPool(16) as mp:", " for out in mp.imap(work, recs, chunksize=256):", " if out is None:", " continue", "- i, k, nt, f = out", "+ i, k, nt, u, v = out", " if k in seen:", " continue", " seen.add(k)", "- ids.append(i); ntok.append(nt); fl.append(f)", "+ ids.append(i); ntok.append(nt); rows.append((u, v))", "+ del recs", " ids = np.array(ids); ntok = np.array(ntok)", " print(\"kept after gates+dedup:\", len(ids), flush=True)", " ", "- # ----------------------------------------------------- feature histograms", "- lens = np.array([len(f) for f in fl], dtype=np.int64)", "- off = np.concatenate([[0], np.cumsum(lens)])", "- flat = np.concatenate(fl).astype(np.int64)", "- pool_h = np.bincount(flat, minlength=NBUCK).astype(np.float64) + 1.0", "- pool_p = np.log(pool_h / pool_h.sum())", "+ # --------------------------------------------------- classifiers + scoring", "+ Xall = to_csr(rows)", "+ neg_idx = rng.choice(len(rows), size=min(NEG, len(rows)), replace=False)", "+ Xn = to_csr([rows[i] for i in neg_idx])", " ", "- # pack pool features into one flat array with offsets for fast scoring", "- inv_len = 1.0 / np.maximum(lens, 1)", "-", " scores = {}", " for name in REGISTERS:", "- th = hist(pos_feats[name]) + 1.0", "- lr = np.log(th / th.sum()) - pool_p # log p_target/p_pool", "- np.clip(lr, -4.0, 2.0, out=lr) # damp: one rare feature can't carry a doc", "- per_tok = lr[flat]", "- s = np.add.reduceat(per_tok, off[:-1]) * inv_len", "- s[lens == 0] = -1e9", "- scores[name] = s", "- print(name, \"score mean/max\", round(s.mean(), 3), round(s.max(), 3), flush=True)", "+ Xp = to_csr(pos[name])", "+ w, b, acc = fit_lr(Xp, Xn)", "+ with torch.no_grad():", "+ scores[name] = _mv(Xall, w, b).float().cpu().numpy()", "+ print(f\"{name}: balanced train acc {acc:.3f} pool score mean \"", "+ f\"{scores[name].mean():.2f} max {scores[name].max():.2f}\", flush=True)", "+ del Xp", " ", "- # ------------------------------------------- per-register rank + round robin", "+ # ------------------------------------------ per-register rank, round robin", " order = {n: np.argsort(-scores[n]) for n in REGISTERS}", " ptr = {n: 0 for n in REGISTERS}", " taken, sel, tot = set(), [], 0"]}, {"oldStart": 207, "oldLines": 6, "newStart": 246, "newLines": 7, "lines": [" ", " assert len(sel) == len(set(sel))", " json.dump(sel, open(OUT, \"w\"))", "+ np.save(\"/tmp/curate_scores.npy\", np.stack([ids] + [scores[n] for n in REGISTERS]))", " print(f\"selected {len(sel)} ids, ~{tot/1e6:.1f}M est tokens \"", " f\"(budget {BUDGET/1e6:.0f}M) -> {OUT}\", flush=True)", " "]}], "originalFile": "\"\"\"Curate /workspace/data/pool.jsonl for a fixed 12M-token pretraining budget.\n\nSTATED CRITERION\n----------------\nThe disclosed evaluation target is a BROAD four-register mix: encyclopedic\n(Wikipedia), high-quality general web prose, news, and technical Q&A. The dev\ntarget `multi_dev.npy` is exactly that mix, stored as four contiguous 250k-token\nblocks (verified by decoding). We select pool documents by *importance\nresampling against that target distribution* (DSIR, Xie et al. 2023), done\nper-register so the budget is spent evenly across all four registers:\n\n 1. Decode `multi_dev.npy`, split on <|endoftext|>, label every target document\n by which quarter (register) it came from.\n 2. Universal quality gates on the pool: length, English-prose stopword rate,\n alphabetic/non-ASCII ratios, and near-duplicate removal via a word-shingle\n hash. These remove junk that no register wants.\n 3. Represent every document as a hashed bag of unigrams+bigrams (2^18 buckets).\n Fit a unigram categorical distribution p_r over features for each target\n register r, and p_pool over the gated pool.\n Score(d, r) = (1/|d|) * sum_f c_f(d) * log( p_r[f] / p_pool[f] )\n i.e. the length-normalised log importance weight of d under register r.\n 4. Rank the pool by Score(., r) for each r and emit a round-robin interleave of\n the four ranked lists. Priority order is therefore best-first *and*\n register-balanced, so truncating at the 12M-token budget keeps the mix.\n\nOnly numpy + the GPT-2 tokenizer are required.\n\"\"\"\nimport json, re, hashlib, zlib, numpy as np\nfrom multiprocessing import Pool as MPPool\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nEMIT_TOKENS = 20_000_000 # emit ~1.7x budget worth of ids\nNBUCK = 1 << 18\nCHARCAP = 8000 # only the first 8k chars of a doc are featurised\nREGISTERS = [\"wiki\", \"web\", \"news\", \"qa\"]\n\nWORD = re.compile(r\"[a-z0-9']+|[^\\sa-z0-9]\")\nSTOP = set(\"the of and to in a is that it for as was on with be by are this an or from at\".split())\n\n\n_HC = {}\n\n\ndef _wh(w):\n \"\"\"Deterministic (run-independent) word hash, memoised.\"\"\"\n h = _HC.get(w)\n if h is None:\n h = _HC[w] = zlib.crc32(w.encode()) * 2654435761 & 0xFFFFFFFF\n return h\n\n\ndef feats(text):\n \"\"\"Hashed unigram+bigram bucket ids for a document.\"\"\"\n ws = WORD.findall(text[:CHARCAP].lower())\n if not ws:\n return np.zeros(0, dtype=np.int32)\n h = [_wh(w) for w in ws]\n b = [(h[i] * 1000003 ^ h[i + 1]) & 0xFFFFFFFF for i in range(len(h) - 1)]\n return (np.array(h + b, dtype=np.int64) % NBUCK).astype(np.int32)\n\n\ndef gate(t):\n \"\"\"Universal junk filters: length, English prose, no mojibake, no markup dumps,\n no SEO word-salad. Returns False to drop, else a code-density flag.\"\"\"\n n = len(t)\n if n < 500 or n > 400_000:\n return False\n samp = t[:20000]\n ws = re.findall(r\"[A-Za-z']+\", samp)\n if len(ws) < 80:\n return False\n lw = [w.lower() for w in ws]\n # (a) English prose: top-20 stopwords are ~25-40% of real English word tokens.\n # SEO word-salad and keyword spam sit far below this.\n if sum(w in STOP for w in lw) / len(lw) < 0.15:\n return False\n # (b) sane word lengths\n mwl = sum(len(w) for w in lw) / len(lw)\n if not (3.0 <= mwl <= 9.0):\n return False\n # (c) mostly letters/space, mostly ASCII\n if sum(c.isalpha() or c.isspace() for c in samp) / len(samp) < 0.70:\n return False\n if sum(ord(c) > 127 for c in samp) / len(samp) > 0.08:\n return False\n # (d) not a script/markup dump: density of code punctuation\n if sum(samp.count(c) for c in \"{};=<>$|\") / len(samp) > 0.012:\n return False\n # (e) not a boilerplate list: too many repeated lines\n lines = [l.strip() for l in samp.split(\"\\n\") if len(l.strip()) > 15]\n if len(lines) >= 10 and len(set(lines)) / len(lines) < 0.6:\n return False\n # (f) lexical diversity floor/ceiling (ceiling catches random-word spam)\n u = len(set(lw)) / len(lw)\n if u > 0.85 and len(lw) > 200:\n return False\n return True\n\n\ndef dedup_key(t):\n ws = [w.lower() for w in re.findall(r\"[A-Za-z']+\", t)][:400]\n if len(ws) < 40:\n return None\n return hashlib.md5(\" \".join(ws[20:40]).encode()).digest()\n\n\ndef work(rec):\n \"\"\"Per-document worker: gate, dedup key, token estimate, hashed features.\"\"\"\n i, t = rec\n if not gate(t):\n return None\n k = dedup_key(t)\n if k is None:\n return None\n f = feats(t)\n return i, k, int(len(t) / 4.0) + 1, f\n\n\ndef hist(fs, n=NBUCK):\n if not len(fs):\n return np.zeros(n)\n return np.bincount(np.concatenate(fs).astype(np.int64), minlength=n).astype(np.float64)\n\n\ndef main():\n # ------------------------------------------------------------- target text\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n EOS = tok.eos_token_id\n dev = np.load(DEV).astype(np.int64)\n Q = len(dev) // 4\n pos_feats = {r: [] for r in REGISTERS}\n for qi, name in enumerate(REGISTERS):\n blk = dev[qi * Q:(qi + 1) * Q]\n cuts = list(np.flatnonzero(blk == EOS)) + [len(blk)]\n prev = 0\n for c in cuts:\n seg = blk[prev:c]; prev = c + 1\n if len(seg) < 64:\n continue\n pos_feats[name].append(feats(tok.decode(seg)))\n print(\"target docs:\", {r: len(v) for r, v in pos_feats.items()}, flush=True)\n\n # ------------------------------------------------------- pool load + gate\n recs = []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n recs.append((r[\"id\"], r[\"text\"]))\n print(\"pool docs:\", len(recs), flush=True)\n\n ids, ntok, fl = [], [], []\n seen = set()\n with MPPool(16) as mp:\n for out in mp.imap(work, recs, chunksize=256):\n if out is None:\n continue\n i, k, nt, f = out\n if k in seen:\n continue\n seen.add(k)\n ids.append(i); ntok.append(nt); fl.append(f)\n ids = np.array(ids); ntok = np.array(ntok)\n print(\"kept after gates+dedup:\", len(ids), flush=True)\n\n # ----------------------------------------------------- feature histograms\n lens = np.array([len(f) for f in fl], dtype=np.int64)\n off = np.concatenate([[0], np.cumsum(lens)])\n flat = np.concatenate(fl).astype(np.int64)\n pool_h = np.bincount(flat, minlength=NBUCK).astype(np.float64) + 1.0\n pool_p = np.log(pool_h / pool_h.sum())\n\n # pack pool features into one flat array with offsets for fast scoring\n inv_len = 1.0 / np.maximum(lens, 1)\n\n scores = {}\n for name in REGISTERS:\n th = hist(pos_feats[name]) + 1.0\n lr = np.log(th / th.sum()) - pool_p # log p_target/p_pool\n np.clip(lr, -4.0, 2.0, out=lr) # damp: one rare feature can't carry a doc\n per_tok = lr[flat]\n s = np.add.reduceat(per_tok, off[:-1]) * inv_len\n s[lens == 0] = -1e9\n scores[name] = s\n print(name, \"score mean/max\", round(s.mean(), 3), round(s.max(), 3), flush=True)\n\n # ------------------------------------------- per-register rank + round robin\n order = {n: np.argsort(-scores[n]) for n in REGISTERS}\n ptr = {n: 0 for n in REGISTERS}\n taken, sel, tot = set(), [], 0\n while tot < EMIT_TOKENS:\n moved = False\n for n in REGISTERS:\n o = order[n]\n while ptr[n] < len(o) and o[ptr[n]] in taken:\n ptr[n] += 1\n if ptr[n] >= len(o):\n continue\n j = int(o[ptr[n]]); ptr[n] += 1\n taken.add(j); sel.append(int(ids[j])); tot += int(ntok[j])\n moved = True\n if not moved:\n break\n\n assert len(sel) == len(set(sel))\n json.dump(sel, open(OUT, \"w\"))\n print(f\"selected {len(sel)} ids, ~{tot/1e6:.1f}M est tokens \"\n f\"(budget {BUDGET/1e6:.0f}M) -> {OUT}\", flush=True)\n\n\nif __name__ == \"__main__\":\n main()\n", "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "grep -E \"PPL\" /tmp/eval2.log; cat .timer/remaining_secs; nvidia-smi --query-gpu=memory.used --format=csv"
}Bash result
HELDOUT_PPL 344.849 loss 5.8431 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
3846
memory.used [MiB]
4 MiB
[stdout]
HELDOUT_PPL 344.849 loss 5.8431 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
3846
memory.used [MiB]
4 MiB
[metadata] {"noOutputExpected": false}v2 (tighter gates): 344.9. Now running the classifier version.
args
{
"command": "nohup python -u submission/curate.py > /tmp/cur3.log 2>&1 &\nsleep 200; grep -viE \"warning|Token indices\" /tmp/cur3.log; cat .timer/remaining_secs",
"timeout": 600000
}Bash result
[1]+ Done nohup python -u submission/curate.py > /tmp/cur3.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 114705
return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),
wiki: balanced train acc 0.995 pool score mean -3.95 max 2.53
web: balanced train acc 0.985 pool score mean -2.55 max 2.08
news: balanced train acc 0.986 pool score mean -2.89 max 3.63
qa: balanced train acc 1.000 pool score mean -4.77 max 1.71
selected 20232 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
3635
[stdout]
[1]+ Done nohup python -u submission/curate.py > /tmp/cur3.log 2>&1
target docs: {'wiki': 1466, 'web': 203, 'news': 423, 'qa': 359}
pool docs: 182016
kept after gates+dedup: 114705
return torch.sparse_csr_tensor(torch.from_numpy(crow).to(dev),
wiki: balanced train acc 0.995 pool score mean -3.95 max 2.53
web: balanced train acc 0.985 pool score mean -2.55 max 2.08
news: balanced train acc 0.986 pool score mean -2.89 max 3.63
qa: balanced train acc 1.000 pool score mean -4.77 max 1.71
selected 20232 ids, ~20.0M est tokens (budget 12M) -> /workspace/submission/selection.json
3635
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cp submission/selection.json /tmp/sel_v3.json; nohup bash dev_eval.sh /tmp/sel_v3.json > /tmp/eval3.log 2>&1 &\nsleep 10; python - <<'EOF'\nimport json\nsel=json.load(open('/tmp/sel_v3.json'))\nwant={i:r for r,i in enumerate(sel[:24])}\ngot={}\nfor l in open('/workspace/data/pool.jsonl'):\n r=json.loads(l)\n if r['id'] in want: got[want[r['id']]]=r['text']\n if len(got)==len(want): break\nreg=['wiki','web','news','qa']\nfor r in range(16):\n print(reg[r%4], repr(got[r][:130]))\nEOF",
"timeout": 400000
}Bash result
wiki 'Jerry Dupont played for the Toronto Marlboros of the Ontario Hockey League at the age of 16. He was drafted in the first round, 15'
web 'TOKYO (AP) — In an apparent contradiction of his national security adviser, President Donald Trump on Sunday downplayed recent Nor'
news 'Mumbai: A Maharashtra court Thursday dropped all charges against two women who wrote comments on Facebook questioning the shutdown'
qa "I'm not sure if I worded my topic title properly, which is probably why I haven't been able to search for the answer to my problem"
wiki 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On '
web 'Lawmakers emerged from a classified administration briefing expressing concern about administration policy on Syria and the legal '
news "NEW DELHI (AP) — India's prime minister demanded Wednesday that Italy return two marines accused of killing a pair of fishermen la"
qa 'XForms/Read and write with get and put\nSometimes all you need to do is to put a nice user friendly form that edits a single static'
wiki 'Mellie Dunham playing his fiddle, Norway, ca. 1927Item 79936 info\nNorway Historical Society\nAlanson Mellen "Mellie" Dunham started'
web 'by Monica Showalter\nThe news is out that then-FBI director James Comey did indeed have some credible prosecutors for Hillary Clint'
news 'LUCKNOW, India (Reuters) - Thousands of youngsters in India have burned down empty train coaches and blocked rail traffic this wee'
qa "Well the panel jam is over. It ended.\nThe serial kind of just faded out.\nOh well... at least you can read everybody else's comics."
wiki "Larrifin, from Savage Empire\nWhen the Avatar first met Lariffin, he was reluctant to give his name. It was only with Chafblum's pr"
web 'President Trump’s ouster of national security adviser Michael Flynn, and the circumstances leading up to it, have quickly become a'
news 'Uttar Pradesh Special Task Force (STF) on Saturday arrested a wanted animal smuggler, carrying a reward of Rs 50,000, who is alleg'
qa 'Scaling the Windows Stack George Beech @GABeech PICC ‘12.\nout of 23\nPost on 27-Dec-2015\nEmbed Size (px)\n<p>PowerPoint Presentation'
[stdout]
wiki 'Jerry Dupont played for the Toronto Marlboros of the Ontario Hockey League at the age of 16. He was drafted in the first round, 15'
web 'TOKYO (AP) — In an apparent contradiction of his national security adviser, President Donald Trump on Sunday downplayed recent Nor'
news 'Mumbai: A Maharashtra court Thursday dropped all charges against two women who wrote comments on Facebook questioning the shutdown'
qa "I'm not sure if I worded my topic title properly, which is probably why I haven't been able to search for the answer to my problem"
wiki 'Major Bhupinder Singh\nIn sept. 1965, Major Bhupinder Singh commanded the "B" squadron of the 4 Horse against Pakistani forces. On '
web 'Lawmakers emerged from a classified administration briefing expressing concern about administration policy on Syria and the legal '
news "NEW DELHI (AP) — India's prime minister demanded Wednesday that Italy return two marines accused of killing a pair of fishermen la"
qa 'XForms/Read and write with get and put\nSometimes all you need to do is to put a nice user friendly form that edits a single static'
wiki 'Mellie Dunham playing his fiddle, Norway, ca. 1927Item 79936 info\nNorway Historical Society\nAlanson Mellen "Mellie" Dunham started'
web 'by Monica Showalter\nThe news is out that then-FBI director James Comey did indeed have some credible prosecutors for Hillary Clint'
news 'LUCKNOW, India (Reuters) - Thousands of youngsters in India have burned down empty train coaches and blocked rail traffic this wee'
qa "Well the panel jam is over. It ended.\nThe serial kind of just faded out.\nOh well... at least you can read everybody else's comics."
wiki "Larrifin, from Savage Empire\nWhen the Avatar first met Lariffin, he was reluctant to give his name. It was only with Chafblum's pr"
web 'President Trump’s ouster of national security adviser Michael Flynn, and the circumstances leading up to it, have quickly become a'
news 'Uttar Pradesh Special Task Force (STF) on Saturday arrested a wanted animal smuggler, carrying a reward of Rs 50,000, who is alleg'
qa 'Scaling the Windows Stack George Beech @GABeech PICC ‘12.\nout of 23\nPost on 27-Dec-2015\nEmbed Size (px)\n<p>PowerPoint Presentation'
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 230; grep -E \"PPL\" /tmp/eval3.log; cat .timer/remaining_secs",
"timeout": 600000
}Bash result
HELDOUT_PPL 340.524 loss 5.8305 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
3375
[stdout]
HELDOUT_PPL 340.524 loss 5.8305 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
3375
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "grep -c \" @-@ \" data/pool.jsonl; grep -c \"s wikipedia\" data/pool.jsonl; python - <<'EOF'\nimport json,re\nn=0;tot=0\nfor l in open('data/pool.jsonl'):\n r=json.loads(l); t=r['text']; tot+=1\n if ' , ' in t[:3000] and ' . ' in t[:3000]: n+=1\nprint('spaced-punct docs',n,'of',tot)\nEOF",
"timeout": 300000
}Bash result
0
7
spaced-punct docs 536 of 182016
[stdout]
0
7
spaced-punct docs 536 of 182016
[metadata] {"noOutputExpected": false}No wikitext artifacts in the pool to exploit. Testing two variants: wiki-weighted round-robin, and a single global ranking (the falsification test from claim.md).
args
{
"command": "python - <<'EOF'\nimport json,numpy as np\nS=np.load('/tmp/curate_scores.npy')\nids=S[0].astype(int); sc={n:S[i+1] for i,n in enumerate(['wiki','web','news','qa'])}\nntok={}\nfor l in open('data/pool.jsonl'):\n r=json.loads(l); ntok[r['id']]=int(len(r['text'])/4.)+1\nnt=np.array([ntok[i] for i in ids])\ndef rr(weights, out, emit=20_000_000):\n order={n:np.argsort(-sc[n]) for n in sc}; ptr={n:0 for n in sc}\n taken=set(); sel=[]; tot=0\n while tot<emit:\n moved=False\n for n,w in weights.items():\n for _ in range(w):\n o=order[n]\n while ptr[n]<len(o) and o[ptr[n]] in taken: ptr[n]+=1\n if ptr[n]>=len(o): continue\n j=int(o[ptr[n]]); ptr[n]+=1; taken.add(j); sel.append(int(ids[j])); tot+=int(nt[j]); moved=True\n if not moved: break\n json.dump(sel,open(out,'w')); print(out,len(sel),tot/1e6)\nrr({'wiki':2,'web':1,'news':1,'qa':1}, '/tmp/sel_B.json')\n# D: single global ranking by max register score (no balancing)\ng=np.max(np.stack([sc[n] for n in sc]),0)\no=np.argsort(-g); sel=[];tot=0\nfor j in o:\n sel.append(int(ids[j])); tot+=int(nt[j])\n if tot>=20_000_000: break\njson.dump(sel,open('/tmp/sel_D.json','w')); print('D',len(sel),tot/1e6)\nEOF",
"timeout": 400000
}Bash result
/tmp/sel_B.json 22345 20.000288
D 15863 20.006813
[stdout]
/tmp/sel_B.json 22345 20.000288
D 15863 20.006813
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "export TOKENIZERS_PARALLELISM=false && (python pack_selection.py /tmp/sel_B.json /tmp/tB.npy > /tmp/pB.log 2>&1 &) && python pack_selection.py /tmp/sel_D.json /tmp/tD.npy 2>&1 | tail -1; sleep 20; tail -1 /tmp/pB.log",
"timeout": 900000
}Bash result
packed 12000000 tokens from 15863 ids -> /tmp/tD.npy
packed 12000000 tokens from 22345 ids -> /tmp/tB.npy
[stdout]
packed 12000000 tokens from 15863 ids -> /tmp/tD.npy
packed 12000000 tokens from 22345 ids -> /tmp/tB.npy
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "for v in B D; do python train_nano.py --train_npy /tmp/t$v.npy --heldout_npy data/multi_dev.npy --out_json /tmp/$v.json 2>&1 | grep HELDOUT | sed \"s/^/$v /\"; done; cat .timer/remaining_secs",
"timeout": 900000
}Bash result
B HELDOUT_PPL 332.557 loss 5.8068 train_s 56 /tmp/tB.npy
D HELDOUT_PPL 441.816 loss 6.0909 train_s 56 /tmp/tD.npy
3104
[stdout]
B HELDOUT_PPL 332.557 loss 5.8068 train_s 56 /tmp/tB.npy
D HELDOUT_PPL 441.816 loss 6.0909 train_s 56 /tmp/tD.npy
3104
[metadata] {"noOutputExpected": false}Strong result: global ranking (D) = 442 ≈ random; balanced = 340; wiki-weighted = 332.6. Pushing the wiki share further, with explicit token shares.
args
{
"command": "python - <<'EOF'\nimport json,numpy as np\nS=np.load('/tmp/curate_scores.npy')\nids=S[0].astype(int); names=['wiki','web','news','qa']; sc={n:S[i+1] for i,n in enumerate(names)}\nntok={}\nfor l in open('data/pool.jsonl'):\n r=json.loads(l); ntok[r['id']]=int(len(r['text'])/4.)+1\nnt=np.array([ntok[i] for i in ids])\nnp.save('/tmp/nt.npy',nt)\ndef build(share,out,emit=20_000_000):\n order={n:np.argsort(-sc[n]) for n in names}; ptr={n:0 for n in names}\n used={n:0 for n in names}; taken=set(); sel=[]; tot=0\n while tot<emit:\n # pick the register furthest below its token quota\n n=min(names,key=lambda n: used[n]/share[n])\n o=order[n]\n while ptr[n]<len(o) and o[ptr[n]] in taken: ptr[n]+=1\n if ptr[n]>=len(o):\n share.pop(n); names.remove(n)\n if not names: break\n continue\n j=int(o[ptr[n]]); ptr[n]+=1; taken.add(j); sel.append(int(ids[j]))\n used[n]+=int(nt[j]); tot+=int(nt[j])\n json.dump(sel,open(out,'w')); print(out,len(sel),round(tot/1e6,1),{k:round(v/tot,2) for k,v in used.items()})\nbuild({'wiki':.40,'web':.20,'news':.20,'qa':.20},'/tmp/sel_W40.json')\nimport importlib\nEOF",
"timeout": 400000
}Bash result
/tmp/sel_W40.json 27339 20.0 {'wiki': 0.4, 'web': 0.2, 'news': 0.2, 'qa': 0.2}
[stdout]
/tmp/sel_W40.json 27339 20.0 {'wiki': 0.4, 'web': 0.2, 'news': 0.2, 'qa': 0.2}
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cat > /tmp/mkvar.py <<'EOF'\nimport json,sys,numpy as np\nS=np.load('/tmp/curate_scores.npy'); nt=np.load('/tmp/nt.npy')\nids=S[0].astype(int); names=['wiki','web','news','qa']; sc={n:S[i+1] for i,n in enumerate(names)}\nshare=dict(zip(names,[float(x) for x in sys.argv[1].split(',')])); out=sys.argv[2]\norder={n:np.argsort(-sc[n]) for n in names}; ptr={n:0 for n in names}\nused={n:0 for n in names}; taken=set(); sel=[]; tot=0; live=list(names)\nwhile tot<20_000_000 and live:\n n=min(live,key=lambda n: used[n]/share[n])\n o=order[n]\n while ptr[n]<len(o) and o[ptr[n]] in taken: ptr[n]+=1\n if ptr[n]>=len(o): live.remove(n); continue\n j=int(o[ptr[n]]); ptr[n]+=1; taken.add(j); sel.append(int(ids[j])); used[n]+=int(nt[j]); tot+=int(nt[j])\njson.dump(sel,open(out,'w')); print(out,len(sel),round(tot/1e6,1),{k:round(v/tot,2) for k,v in used.items()})\nEOF\npython /tmp/mkvar.py 0.60,0.1333,0.1333,0.1333 /tmp/sel_W60.json\nexport TOKENIZERS_PARALLELISM=false\n(python pack_selection.py /tmp/sel_W40.json /tmp/tW40.npy > /tmp/p40.log 2>&1 &); python pack_selection.py /tmp/sel_W60.json /tmp/tW60.npy 2>&1|tail -1; sleep 15; tail -1 /tmp/p40.log; cat .timer/remaining_secs",
"timeout": 900000
}Bash result
/tmp/sel_W60.json 29913 20.0 {'wiki': 0.6, 'web': 0.13, 'news': 0.13, 'qa': 0.13}
packed 12000000 tokens from 29913 ids -> /tmp/tW60.npy
packed 12000000 tokens from 27339 ids -> /tmp/tW40.npy
2994
[stdout]
/tmp/sel_W60.json 29913 20.0 {'wiki': 0.6, 'web': 0.13, 'news': 0.13, 'qa': 0.13}
packed 12000000 tokens from 29913 ids -> /tmp/tW60.npy
packed 12000000 tokens from 27339 ids -> /tmp/tW40.npy
2994
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "for v in W40 W60; do python train_nano.py --train_npy /tmp/t$v.npy --heldout_npy data/multi_dev.npy --out_json /tmp/$v.json 2>&1|grep HELDOUT|sed \"s/^/$v /\"; done; cat .timer/remaining_secs",
"timeout": 900000
}Bash result
W40 HELDOUT_PPL 326.537 loss 5.7885 train_s 56 /tmp/tW40.npy
W60 HELDOUT_PPL 331.116 loss 5.8025 train_s 56 /tmp/tW60.npy
2853
[stdout]
W40 HELDOUT_PPL 326.537 loss 5.7885 train_s 56 /tmp/tW40.npy
W60 HELDOUT_PPL 331.116 loss 5.8025 train_s 56 /tmp/tW60.npy
2853
[metadata] {"noOutputExpected": false}W40 (40% wiki) is best at 326.5. Refining around that.
args
{
"command": "python /tmp/mkvar.py 0.33,0.2233,0.2233,0.2233 /tmp/sel_W33.json; python /tmp/mkvar.py 0.48,0.1733,0.1733,0.1733 /tmp/sel_W48.json\nexport TOKENIZERS_PARALLELISM=false\n(python pack_selection.py /tmp/sel_W33.json /tmp/tW33.npy > /tmp/p33.log 2>&1 &); python pack_selection.py /tmp/sel_W48.json /tmp/tW48.npy 2>&1|tail -1; sleep 15; tail -1 /tmp/p33.log",
"timeout": 900000
}Bash result
/tmp/sel_W33.json 26362 20.0 {'wiki': 0.33, 'web': 0.22, 'news': 0.22, 'qa': 0.22}
/tmp/sel_W48.json 28554 20.0 {'wiki': 0.48, 'web': 0.17, 'news': 0.17, 'qa': 0.17}
packed 12000000 tokens from 28554 ids -> /tmp/tW48.npy
packed 12000000 tokens from 26362 ids -> /tmp/tW33.npy
[stdout]
/tmp/sel_W33.json 26362 20.0 {'wiki': 0.33, 'web': 0.22, 'news': 0.22, 'qa': 0.22}
/tmp/sel_W48.json 28554 20.0 {'wiki': 0.48, 'web': 0.17, 'news': 0.17, 'qa': 0.17}
packed 12000000 tokens from 28554 ids -> /tmp/tW48.npy
packed 12000000 tokens from 26362 ids -> /tmp/tW33.npy
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "for v in W33 W48; do python train_nano.py --train_npy /tmp/t$v.npy --heldout_npy data/multi_dev.npy --out_json /tmp/$v.json 2>&1|grep HELDOUT|sed \"s/^/$v /\"; done; cat .timer/remaining_secs",
"timeout": 900000
}Bash result
W33 HELDOUT_PPL 321.606 loss 5.7733 train_s 57 /tmp/tW33.npy
W48 HELDOUT_PPL 319.198 loss 5.7658 train_s 56 /tmp/tW48.npy
2643
[stdout]
W33 HELDOUT_PPL 321.606 loss 5.7733 train_s 57 /tmp/tW33.npy
W48 HELDOUT_PPL 319.198 loss 5.7658 train_s 56 /tmp/tW48.npy
2643
[metadata] {"noOutputExpected": false}