claude-opus-5
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-5
- metric
- held-out multi-domain perplexity
- actions
- 80
- duration
- 1h 6m
- recorded
- 2026-08-03
run iddclm-nano-curation__xNCncBo
Everything claude-opus-5 did on this task, in order.
download JSON ↗I'll start by getting oriented — checking the time budget and understanding the environment.
args
{
"command": "cat /workspace/.timer/remaining_secs 2>/dev/null; echo \"---\"; ls -la /workspace/; echo \"---\"; ls -la /workspace/data/; echo \"---\"; nvidia-smi 2>/dev/null | head -20",
"description": "Check timer, list workspace and data"
}Bash result
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 3 23:25 .
drwxr-xr-x 1 root root 160 Aug 3 23:25 ..
drwxr-xr-x 1 root root 60 Aug 3 23:25 .live
drwxr-xr-x 1 root root 100 Aug 3 23:25 .metrics
drwxr-xr-x 1 root root 140 Aug 3 23:25 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 3 23:25 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---
Mon Aug 3 23:25:58 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 80GB HBM3 On | 00000000:85:00.0 Off | 0 |
| N/A 31C P0 74W / 700W | 4MiB / 81559MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
[stdout]
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 3 23:25 .
drwxr-xr-x 1 root root 160 Aug 3 23:25 ..
drwxr-xr-x 1 root root 60 Aug 3 23:25 .live
drwxr-xr-x 1 root root 100 Aug 3 23:25 .metrics
drwxr-xr-x 1 root root 140 Aug 3 23:25 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 3 23:25 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---
Mon Aug 3 23:25:58 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 80GB HBM3 On | 00000000:85:00.0 Off | 0 |
| N/A 31C P0 74W / 700W | 4MiB / 81559MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
[metadata] {"noOutputExpected": false}args
{
"file_path": "/workspace/train_nano.py"
}Read result
1 """From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
2 Frozen recipe — the ONLY thing that varies between runs is the training token set."""
3 import argparse, math, time, numpy as np, torch
4 from model import GPT, GPTConfig
5
6 ap = argparse.ArgumentParser()
7 ap.add_argument("--train_npy", required=True)
8 ap.add_argument("--heldout_npy", required=True)
9 ap.add_argument("--out_json", required=True)
10 ap.add_argument("--max_iters", type=int, default=3000)
11 ap.add_argument("--warmup", type=int, default=150)
12 ap.add_argument("--batch", type=int, default=32)
13 ap.add_argument("--block", type=int, default=256)
14 ap.add_argument("--lr", type=float, default=6e-4)
15 ap.add_argument("--seed", type=int, default=1337)
16 a = ap.parse_args()
17
18 torch.manual_seed(a.seed); np.random.seed(a.seed)
19 dev = "cuda"
20 tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
21 tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
22 rng = np.random.default_rng(a.seed)
23
24 model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
25 n_embd=384, dropout=0.0, bias=False)).to(dev)
26 opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
27
28 def lr_at(it):
29 if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
30 r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
31 return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
32
33 def get_batch(src):
34 ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
35 x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
36 y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
37 return x, y
38
39 t0 = time.time()
40 model.train()
41 for it in range(a.max_iters):
42 for g in opt.param_groups: g["lr"] = lr_at(it)
43 x, y = get_batch(tr)
44 with torch.autocast("cuda", dtype=torch.bfloat16):
45 _, loss = model(x, y)
46 opt.zero_grad(set_to_none=True); loss.backward()
47 torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
48 train_s = time.time() - t0
49
50 # held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
51 import json
52 model.eval()
53 nwin = (len(ho) - 1) // a.block
54 starts = [j * a.block for j in range(nwin)]
55 sum_loss, nseen = 0.0, 0
56 with torch.no_grad():
57 for i in range(0, nwin, a.batch):
58 bs = starts[i:i+a.batch]
59 x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
60 y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
61 with torch.autocast("cuda", dtype=torch.bfloat16):
62 _, l = model(x, y)
63 sum_loss += l.item() * len(bs); nseen += len(bs)
64 mean_loss = sum_loss / max(1, nseen)
65 ppl = math.exp(mean_loss)
66 out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
67 "max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
68 json.dump(out, open(a.out_json, "w"))
69 print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
70
[metadata] {"type": "text", "file": {"filePath": "/workspace/train_nano.py", "content": "\"\"\"From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.\nFrozen recipe — the ONLY thing that varies between runs is the training token set.\"\"\"\nimport argparse, math, time, numpy as np, torch\nfrom model import GPT, GPTConfig\n\nap = argparse.ArgumentParser()\nap.add_argument(\"--train_npy\", required=True)\nap.add_argument(\"--heldout_npy\", required=True)\nap.add_argument(\"--out_json\", required=True)\nap.add_argument(\"--max_iters\", type=int, default=3000)\nap.add_argument(\"--warmup\", type=int, default=150)\nap.add_argument(\"--batch\", type=int, default=32)\nap.add_argument(\"--block\", type=int, default=256)\nap.add_argument(\"--lr\", type=float, default=6e-4)\nap.add_argument(\"--seed\", type=int, default=1337)\na = ap.parse_args()\n\ntorch.manual_seed(a.seed); np.random.seed(a.seed)\ndev = \"cuda\"\ntr = np.load(a.train_npy); ho = np.load(a.heldout_npy)\ntr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))\nrng = np.random.default_rng(a.seed)\n\nmodel = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,\n n_embd=384, dropout=0.0, bias=False)).to(dev)\nopt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), \"cuda\")\n\ndef lr_at(it):\n if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)\n r = (it - a.warmup) / max(1, a.max_iters - a.warmup)\n return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)\n\ndef get_batch(src):\n ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)\n x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)\n y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)\n return x, y\n\nt0 = time.time()\nmodel.train()\nfor it in range(a.max_iters):\n for g in opt.param_groups: g[\"lr\"] = lr_at(it)\n x, y = get_batch(tr)\n with torch.autocast(\"cuda\", dtype=torch.bfloat16):\n _, loss = model(x, y)\n opt.zero_grad(set_to_none=True); loss.backward()\n torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()\ntrain_s = time.time() - t0\n\n# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)\nimport json\nmodel.eval()\nnwin = (len(ho) - 1) // a.block\nstarts = [j * a.block for j in range(nwin)]\nsum_loss, nseen = 0.0, 0\nwith torch.no_grad():\n for i in range(0, nwin, a.batch):\n bs = starts[i:i+a.batch]\n x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)\n y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)\n with torch.autocast(\"cuda\", dtype=torch.bfloat16):\n _, l = model(x, y)\n sum_loss += l.item() * len(bs); nseen += len(bs)\nmean_loss = sum_loss / max(1, nseen)\nppl = math.exp(mean_loss)\nout = {\"train_npy\": a.train_npy, \"heldout_ppl\": ppl, \"mean_loss\": mean_loss,\n \"max_iters\": a.max_iters, \"seed\": a.seed, \"train_s\": round(train_s, 1)}\njson.dump(out, open(a.out_json, \"w\"))\nprint(f\"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}\")\n", "numLines": 70, "startLine": 1, "totalLines": 70}}90 minutes of wall clock. Let me move fast and check the key details.
args
{
"command": "cat pack_selection.py; echo \"=== dev_eval.sh ===\"; cat dev_eval.sh; echo \"=== model.py head ===\"; head -30 model.py; echo \"=== python pkgs ===\"; python -c \"import tiktoken; print('tiktoken', tiktoken.__version__)\" 2>&1 | tail -1; python -c \"import transformers; print('tf', transformers.__version__)\" 2>&1|tail -1; python -c \"import sklearn; print('sk', sklearn.__version__)\" 2>&1|tail -1; python -c \"import numpy,torch; print(numpy.__version__, torch.__version__)\"; nproc"
}Bash result
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
=== model.py head ===
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
=== python pkgs ===
ModuleNotFoundError: No module named 'tiktoken'
tf 4.57.1
ModuleNotFoundError: No module named 'sklearn'
1.26.4 2.8.0+cu128
16
[stdout]
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
=== model.py head ===
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
=== python pkgs ===
ModuleNotFoundError: No module named 'tiktoken'
tf 4.57.1
ModuleNotFoundError: No module named 'sklearn'
1.26.4 2.8.0+cu128
16
[metadata] {"noOutputExpected": false}args
{
"command": "head -c 3000 data/pool.jsonl; echo; echo \"=== dev npy ===\"; python -c \"\nimport numpy as np\nd=np.load('data/multi_dev.npy'); print(d.shape, d.dtype, d.min(), d.max())\n\"",
"description": "Inspect pool sample and dev target"
}Bash result
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is not indicative of someone who made an innocent mistake. Its weird.\nSo what with this frantic running? Is the inability to say \"I was wrong\" a pathological feature of Creationists? Or are you hiding something? Or is it both? Or is it more?\nAnd now we get Casey weighing in on the issue, according to cre8id at AboveTopSecret.com-- PBS/NOVA online - Intelligent Design on trial:\n...to my knowledge, Discovery Institute has neither authorized nor received nor is making use of any presentation that used that animation. We have had nothing to do with creating or selling a DVD of that animation, nor do we have anything to do with placing that presentation on Google Video.I dont know what he is talking about with that last part, but the first part sounds similar to DIs claims post-Dover (\"WE HAD NOTHING TO DO WITH DOVER!\"). Maybe Luskin is telling the truth. Maybe this was a magic non-science Creation-friendly narration with convenient edits that AiG or ICR would have killed for... but onl
=== dev npy ===
(1000000,) uint16 0 50256
[stdout]
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is not indicative of someone who made an innocent mistake. Its weird.\nSo what with this frantic running? Is the inability to say \"I was wrong\" a pathological feature of Creationists? Or are you hiding something? Or is it both? Or is it more?\nAnd now we get Casey weighing in on the issue, according to cre8id at AboveTopSecret.com-- PBS/NOVA online - Intelligent Design on trial:\n...to my knowledge, Discovery Institute has neither authorized nor received nor is making use of any presentation that used that animation. We have had nothing to do with creating or selling a DVD of that animation, nor do we have anything to do with placing that presentation on Google Video.I dont know what he is talking about with that last part, but the first part sounds similar to DIs claims post-Dover (\"WE HAD NOTHING TO DO WITH DOVER!\"). Maybe Luskin is telling the truth. Maybe this was a magic non-science Creation-friendly narration with convenient edits that AiG or ICR would have killed for... but onl
=== dev npy ===
(1000000,) uint16 0 50256
[metadata] {"noOutputExpected": false}Raw web pool + a 1M-token high-quality target I can decode. Plan: pre-tokenize the whole pool once (gives exact token counts + instant packing), then build a quality classifier against the decoded dev target. Launching tokenization now in the background.
args
{
"file_path": "/workspace/work/tokenize_pool.py",
"content": "\"\"\"Pre-tokenize the entire pool once with GPT-2 BPE.\n\nWrites /workspace/work/tok/shard_XX.npy (uint16 token stream, EOS-terminated per doc)\nand shard_XX_meta.npy (int64 [id, ntokens] per doc, in file order).\nThis makes exact token accounting and instant packing possible for any selection.\n\"\"\"\nimport os, json, numpy as np\nfrom multiprocessing import Pool\n\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\nPOOL = \"/workspace/data/pool.jsonl\"\nOUT = \"/workspace/work/tok\"\nNSHARD = 14\nos.makedirs(OUT, exist_ok=True)\n\n\ndef byte_offsets(path, n):\n \"\"\"Split the file into n chunks that start on line boundaries.\"\"\"\n size = os.path.getsize(path)\n bounds = [0]\n with open(path, \"rb\") as f:\n for k in range(1, n):\n f.seek(size * k // n)\n f.readline() # discard partial line\n bounds.append(f.tell())\n bounds.append(size)\n return list(zip(bounds[:-1], bounds[1:]))\n\n\ndef run(job):\n k, (start, end) = job\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n toks, meta = [], []\n with open(POOL, \"rb\") as f:\n f.seek(start)\n buf_txt, buf_id = [], []\n\n def flush():\n if not buf_txt:\n return\n enc = tok(buf_txt, add_special_tokens=False)[\"input_ids\"]\n for i, ids in zip(buf_id, enc):\n toks.append(np.array(ids + [eos], dtype=np.uint16))\n meta.append((i, len(ids) + 1))\n buf_txt.clear(); buf_id.clear()\n\n while f.tell() < end:\n line = f.readline()\n if not line:\n break\n r = json.loads(line)\n buf_id.append(r[\"id\"]); buf_txt.append(r[\"text\"])\n if len(buf_txt) >= 256:\n flush()\n flush()\n np.save(f\"{OUT}/shard_{k:02d}.npy\", np.concatenate(toks))\n np.save(f\"{OUT}/shard_{k:02d}_meta.npy\", np.array(meta, dtype=np.int64))\n return k, len(meta)\n\n\nif __name__ == \"__main__\":\n jobs = list(enumerate(byte_offsets(POOL, NSHARD)))\n with Pool(NSHARD) as p:\n for k, n in p.imap_unordered(run, jobs):\n print(f\"shard {k}: {n} docs\", flush=True)\n print(\"DONE\")\n"
}Write result
File created successfully at: /workspace/work/tokenize_pool.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/work/tokenize_pool.py", "content": "\"\"\"Pre-tokenize the entire pool once with GPT-2 BPE.\n\nWrites /workspace/work/tok/shard_XX.npy (uint16 token stream, EOS-terminated per doc)\nand shard_XX_meta.npy (int64 [id, ntokens] per doc, in file order).\nThis makes exact token accounting and instant packing possible for any selection.\n\"\"\"\nimport os, json, numpy as np\nfrom multiprocessing import Pool\n\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\nPOOL = \"/workspace/data/pool.jsonl\"\nOUT = \"/workspace/work/tok\"\nNSHARD = 14\nos.makedirs(OUT, exist_ok=True)\n\n\ndef byte_offsets(path, n):\n \"\"\"Split the file into n chunks that start on line boundaries.\"\"\"\n size = os.path.getsize(path)\n bounds = [0]\n with open(path, \"rb\") as f:\n for k in range(1, n):\n f.seek(size * k // n)\n f.readline() # discard partial line\n bounds.append(f.tell())\n bounds.append(size)\n return list(zip(bounds[:-1], bounds[1:]))\n\n\ndef run(job):\n k, (start, end) = job\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n toks, meta = [], []\n with open(POOL, \"rb\") as f:\n f.seek(start)\n buf_txt, buf_id = [], []\n\n def flush():\n if not buf_txt:\n return\n enc = tok(buf_txt, add_special_tokens=False)[\"input_ids\"]\n for i, ids in zip(buf_id, enc):\n toks.append(np.array(ids + [eos], dtype=np.uint16))\n meta.append((i, len(ids) + 1))\n buf_txt.clear(); buf_id.clear()\n\n while f.tell() < end:\n line = f.readline()\n if not line:\n break\n r = json.loads(line)\n buf_id.append(r[\"id\"]); buf_txt.append(r[\"text\"])\n if len(buf_txt) >= 256:\n flush()\n flush()\n np.save(f\"{OUT}/shard_{k:02d}.npy\", np.concatenate(toks))\n np.save(f\"{OUT}/shard_{k:02d}_meta.npy\", np.array(meta, dtype=np.int64))\n return k, len(meta)\n\n\nif __name__ == \"__main__\":\n jobs = list(enumerate(byte_offsets(POOL, NSHARD)))\n with Pool(NSHARD) as p:\n for k, n in p.imap_unordered(run, jobs):\n print(f\"shard {k}: {n} docs\", flush=True)\n print(\"DONE\")\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "nohup python work/tokenize_pool.py > work/tok.log 2>&1 &\necho \"launched $!\"",
"description": "Launch pool tokenization in background"
}Bash result
launched 368
[stdout]
launched 368
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "python -c \"\nimport numpy as np, os\nos.environ['TOKENIZERS_PARALLELISM']='false'\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained('gpt2')\nd=np.load('data/multi_dev.npy').astype(np.int64)\neos=tok.eos_token_id\nprint('n eos:', (d==eos).sum())\n# split on EOS to get documents\nidx=np.where(d==eos)[0]\nprint('first 20 eos idx:', idx[:20])\nsegs=[]\nprev=0\nfor i in idx:\n segs.append(d[prev:i]); prev=i+1\nsegs.append(d[prev:])\nprint('nsegs', len(segs), 'lens sample', [len(s) for s in segs[:20]])\ntxts=[tok.decode(s) for s in segs]\nimport json\njson.dump(txts, open('work/dev_docs.json','w'))\nfor t in txts[:6]:\n print('-----', len(t))\n print(t[:600].replace(chr(10),' | '))\n\"\n",
"description": "Decode dev target into documents"
}Bash result
n eos: 2704
first 20 eos idx: [ 206 332 556 652 854 986 1010 1080 1231 1555 1789 2010 2179 2412
2473 2516 2702 2875 3017 3128]
nsegs 2705 lens sample [206, 125, 223, 95, 201, 131, 23, 69, 150, 323, 233, 220, 168, 232, 60, 42, 185, 172, 141, 110]
----- 1043
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
----- 603
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
----- 1037
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
----- 435
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused . |
----- 889
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
----- 603
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
[stdout]
n eos: 2704
first 20 eos idx: [ 206 332 556 652 854 986 1010 1080 1231 1555 1789 2010 2179 2412
2473 2516 2702 2875 3017 3128]
nsegs 2705 lens sample [206, 125, 223, 95, 201, 131, 23, 69, 150, 323, 233, 220, 168, 232, 60, 42, 185, 172, 141, 110]
----- 1043
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
----- 603
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
----- 1037
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
----- 435
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused . |
----- 889
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
----- 603
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
[metadata] {"noOutputExpected": false}Important: the encyclopedic portion is WikiText-style ( @-@ , spaces around punctuation). Let me survey all four registers before building any classifier.
args
{
"command": "python -c \"\nimport json\nt=json.load(open('work/dev_docs.json'))\nn=len(t)\nimport textwrap\nfor frac in [0.05,0.2,0.3,0.35,0.45,0.5,0.6,0.7,0.75,0.85,0.95]:\n i=int(frac*n)\n print('==== idx',i,'frac',frac,'len',len(t[i]))\n print(t[i][:450].replace(chr(10),' | '))\n\"\n",
"description": "Sample dev docs across the file"
}Bash result
==== idx 135 frac 0.05 len 893
According to a 2011 PricewaterhouseCoopers report , India 's GDP at purchasing power parity could overtake that of the United States by 2045 . During the next four decades , Indian GDP is expected to grow at an annualised average of 8 % , making it potentially the world 's fastest @-@ growing major economy until 2050 . The report highlights key growth factors : a young and rapidly growing working @-@ age population ; growth in the manufacturing
==== idx 541 frac 0.2 len 1114
The two battalions of the 27th Infantry went into a perimeter defense just north of the village of Soi @-@ ri . The 1st Battalion , on the left of the road , took a position with C Company on high ground in front , and with A Company on a ridge behind it . On their right , B Company was placed parallel to A Company , and carried the line across the stream and the narrow valley to the road . There the 2nd Battalion took up the defense line with E
==== idx 811 frac 0.3 len 1235
In Traveling Shoes , Angelou was able to recognize similarities between African and African @-@ American culture ; as Lupton put it , the " blue songs , shouts , and gospels " she has grown up with in America " echo the rhythms of West Africa " . Marcia Ann Gillespie and her colleagues , writing in A Glorious Celebration , the book published in 2008 for Angelou 's 80th birthday , agreed , stating that Angelou recognized the connections between A
==== idx 946 frac 0.35 len 670
The sixth match was for the WWE Championship between John " Bradshaw " Layfield ( JBL ) and Booker T. The match saw both men take the advantage over one another . Orlando Jordan interfered several times by attacking Booker T. During the match , the referee was knocked out by JBL . Booker T managed to perform the scissors kick on JBL , as a new referee came down to the ring to officiate the match . Jordan , however , removed JBL from the ring and
==== idx 1217 frac 0.45 len 442
Rich and productive , Pittsburgh was also the " Smoky City , " with smog sometimes so thick that streetlights burned during the day as well as rivers that resembled open sewers . Civic leaders , notably Mayor David L. Lawrence , elected in 1945 , Richard K. Mellon , chairman of Mellon Bank and John P. Robin began smoke control and urban revitalization , also known as Urban Renewal projects that transformed the city in unforeseen ways . |
==== idx 1352 frac 0.5 len 493
Mining subsidence coupled with structural and political changes to the mining industry began the decline in Astley 's industrial activities during the mid @-@ 20th century ; its cotton mill closed in 1955 , and the last coal was brought to the surface in 1970 . However , Astley has grown as part of a commuter belt , supported by its proximity to Manchester city centre and inter @-@ city transport links . Astley Green Colliery Museum houses colle
==== idx 1623 frac 0.6 len 745
Commercial concentrating solar power ( CSP ) plants , also called " solar thermal power stations " , were first developed in the 1980s . The 377 MW Ivanpah Solar Power Facility , located in California 's Mojave Desert , is the world ’ s largest solar thermal power plant project . Other large CSP plants include the Solnova Solar Power Station ( 150 MW ) , the Andasol solar power station ( 150 MW ) , and Extresol Solar Power Station ( 150 MW ) , a
==== idx 1893 frac 0.7 len 1967
The year 1845 was a time of unimaginable deprivations: No smart phones, no Twitter, no Words With Friends, plus there was a lot of cholera, and those Irish street gangs, and also slavery somewhere. The Gowanus Canal was not just a repulsive sewage channel and heartbreaking symbol of environmental devastation, but a primary means of public transport through Brooklyn — construction on the great bridge to Manhattan wouldn’t be begin for another quar
==== idx 2028 frac 0.75 len 3804
When Congress president Rahul Gandhi enters the imposing corridors of the 1300-year-old Sharada Peeth on Wednesday, historians will remember the time his grandmother Indira Gandhi visited the spot 40 years ago in 1978.Like the Congress of today, the party was going through a slump in 1978. The excesses of Emergency had cost her the chair of the Prime Minister, she was not even a Member of Parliament.Her family was going through enormous personal
==== idx 2299 frac 0.85 len 2103
The benchmark BSE Sensex rose 147 points and Nifty reclaimed 10,600 mark in early trade on Wednesday on sustained buying by domestic institutional investors amid firm Asian cues.Encouraging earnings by blue-chip companies also added to the uptrend.The 30-share index rose 147.25 points, or 0.42 percent, to 34,598.02. The gauge had gained 35.19 points in the previous session. The NSE Nifty was up 23.90 points, or 0.22 percent, to 10,608.60Major gai
==== idx 2569 frac 0.95 len 1439
<p>I am trying to make this <a href="http://www.webdesignbeach.com/beachbar/ajax-fancy-captcha-jquery-plugin" rel="nofollow noreferrer">captcha</a> jquery plugin to work. The a certain line of code is executed, the error pops up.</p> | | <p>This is the line of code that causes the error :</p> | | <pre><code>$(".ajax-fc-" + rand).draggable({ containment: '#ajax-fc-content' }); | </code></pre> | | <p>What I am assuming is that there is some kind of conflict w
[stdout]
==== idx 135 frac 0.05 len 893
According to a 2011 PricewaterhouseCoopers report , India 's GDP at purchasing power parity could overtake that of the United States by 2045 . During the next four decades , Indian GDP is expected to grow at an annualised average of 8 % , making it potentially the world 's fastest @-@ growing major economy until 2050 . The report highlights key growth factors : a young and rapidly growing working @-@ age population ; growth in the manufacturing
==== idx 541 frac 0.2 len 1114
The two battalions of the 27th Infantry went into a perimeter defense just north of the village of Soi @-@ ri . The 1st Battalion , on the left of the road , took a position with C Company on high ground in front , and with A Company on a ridge behind it . On their right , B Company was placed parallel to A Company , and carried the line across the stream and the narrow valley to the road . There the 2nd Battalion took up the defense line with E
==== idx 811 frac 0.3 len 1235
In Traveling Shoes , Angelou was able to recognize similarities between African and African @-@ American culture ; as Lupton put it , the " blue songs , shouts , and gospels " she has grown up with in America " echo the rhythms of West Africa " . Marcia Ann Gillespie and her colleagues , writing in A Glorious Celebration , the book published in 2008 for Angelou 's 80th birthday , agreed , stating that Angelou recognized the connections between A
==== idx 946 frac 0.35 len 670
The sixth match was for the WWE Championship between John " Bradshaw " Layfield ( JBL ) and Booker T. The match saw both men take the advantage over one another . Orlando Jordan interfered several times by attacking Booker T. During the match , the referee was knocked out by JBL . Booker T managed to perform the scissors kick on JBL , as a new referee came down to the ring to officiate the match . Jordan , however , removed JBL from the ring and
==== idx 1217 frac 0.45 len 442
Rich and productive , Pittsburgh was also the " Smoky City , " with smog sometimes so thick that streetlights burned during the day as well as rivers that resembled open sewers . Civic leaders , notably Mayor David L. Lawrence , elected in 1945 , Richard K. Mellon , chairman of Mellon Bank and John P. Robin began smoke control and urban revitalization , also known as Urban Renewal projects that transformed the city in unforeseen ways . |
==== idx 1352 frac 0.5 len 493
Mining subsidence coupled with structural and political changes to the mining industry began the decline in Astley 's industrial activities during the mid @-@ 20th century ; its cotton mill closed in 1955 , and the last coal was brought to the surface in 1970 . However , Astley has grown as part of a commuter belt , supported by its proximity to Manchester city centre and inter @-@ city transport links . Astley Green Colliery Museum houses colle
==== idx 1623 frac 0.6 len 745
Commercial concentrating solar power ( CSP ) plants , also called " solar thermal power stations " , were first developed in the 1980s . The 377 MW Ivanpah Solar Power Facility , located in California 's Mojave Desert , is the world ’ s largest solar thermal power plant project . Other large CSP plants include the Solnova Solar Power Station ( 150 MW ) , the Andasol solar power station ( 150 MW ) , and Extresol Solar Power Station ( 150 MW ) , a
==== idx 1893 frac 0.7 len 1967
The year 1845 was a time of unimaginable deprivations: No smart phones, no Twitter, no Words With Friends, plus there was a lot of cholera, and those Irish street gangs, and also slavery somewhere. The Gowanus Canal was not just a repulsive sewage channel and heartbreaking symbol of environmental devastation, but a primary means of public transport through Brooklyn — construction on the great bridge to Manhattan wouldn’t be begin for another quar
==== idx 2028 frac 0.75 len 3804
When Congress president Rahul Gandhi enters the imposing corridors of the 1300-year-old Sharada Peeth on Wednesday, historians will remember the time his grandmother Indira Gandhi visited the spot 40 years ago in 1978.Like the Congress of today, the party was going through a slump in 1978. The excesses of Emergency had cost her the chair of the Prime Minister, she was not even a Member of Parliament.Her family was going through enormous personal
==== idx 2299 frac 0.85 len 2103
The benchmark BSE Sensex rose 147 points and Nifty reclaimed 10,600 mark in early trade on Wednesday on sustained buying by domestic institutional investors amid firm Asian cues.Encouraging earnings by blue-chip companies also added to the uptrend.The 30-share index rose 147.25 points, or 0.42 percent, to 34,598.02. The gauge had gained 35.19 points in the previous session. The NSE Nifty was up 23.90 points, or 0.22 percent, to 10,608.60Major gai
==== idx 2569 frac 0.95 len 1439
<p>I am trying to make this <a href="http://www.webdesignbeach.com/beachbar/ajax-fancy-captcha-jquery-plugin" rel="nofollow noreferrer">captcha</a> jquery plugin to work. The a certain line of code is executed, the error pops up.</p> | | <p>This is the line of code that causes the error :</p> | | <pre><code>$(".ajax-fc-" + rand).draggable({ containment: '#ajax-fc-content' }); | </code></pre> | | <p>What I am assuming is that there is some kind of conflict w
[metadata] {"noOutputExpected": false}args
{
"command": "cat /workspace/work/tok.log; echo \"---\"; ls -la /workspace/work/tok/ 2>/dev/null | head; cat /workspace/.timer/remaining_secs",
"description": "Check tokenization progress"
}Bash result
Token indices sequence length is longer than the specified maximum sequence length for this model (1098 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1545 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1381 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (4013 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (2977 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (5281 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (7144 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1897 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (7364 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1508 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1313 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (2736 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1285 > 1024). Running this sequence through the model will result in indexing errors
---
total 0
drwxr-xr-x 1 root root 40 Aug 3 23:26 .
drwxr-xr-x 1 root root 140 Aug 3 23:27 ..
5310
[stdout]
Token indices sequence length is longer than the specified maximum sequence length for this model (1098 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1545 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1381 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (4013 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (2977 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (5281 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (7144 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1897 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (7364 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1508 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1313 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (2736 > 1024). Running this sequence through the model will result in indexing errors
Token indices sequence length is longer than the specified maximum sequence length for this model (1285 > 1024). Running this sequence through the model will result in indexing errors
---
total 0
drwxr-xr-x 1 root root 40 Aug 3 23:26 .
drwxr-xr-x 1 root root 140 Aug 3 23:27 ..
5310
[metadata] {"noOutputExpected": false}Four registers confirmed, in contiguous blocks. Let me pin the boundaries and check what the pool contains.
args
{
"command": "python -c \"\nimport json,re\nt=json.load(open('work/dev_docs.json'))\ndef feat(s):\n return ('@-@' in s or '@,@' in s or re.search(r\\\" 's \\\", s) is not None, s.strip().startswith('<p>') or '<pre><code>' in s or '</a>' in s)\n# scan for transitions\nlabs=[]\nfor i,s in enumerate(t):\n wiki = ('@-@' in s) or ('@,@' in s) or bool(re.search(r\\\"\\s'(s|t|ll|ve|re|m|d)\\s\\\", s)) or bool(re.search(r'\\s[,.;:]\\s', s))\n html = bool(re.search(r'</?(p|pre|code|a|blockquote|em|strong|ul|li)[ >]', s))\n labs.append(('W' if wiki else '.')+('H' if html else '.'))\nprint(''.join(l[0] for l in labs))\nprint()\nprint(''.join(l[1] for l in labs))\n\"\n",
"description": "Locate register boundaries in dev"
}Bash result
WWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWW.WWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWW.......W......W...W...............W...W.........W...................W................W.W.W..............W.......W....W.W..........W..........................................WW.............W.......W....................................W...W.......W.......W.....................W................................................................................W.................W............................W.........................................................................W.....................W........................................................................................W....................W...........................W....W.......W..WWW...W...WW..W..W.............W..................W........W..WW............WW..................W.............W...W....W....W....W........W.........W..........WWW.......W.............WW............W......W..............WW..........W.............W.............W..W..........WW......W........W.........WW..W..............W.....WW......W..W.
..........................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
[stdout]
WWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWW.WWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWWW.......W......W...W...............W...W.........W...................W................W.W.W..............W.......W....W.W..........W..........................................WW.............W.......W....................................W...W.......W.......W.....................W................................................................................W.................W............................W.........................................................................W.....................W........................................................................................W....................W...........................W....W.......W..WWW...W...WW..W..W.............W..................W........W..WW............WW..................W.............W...W....W....W....W........W.........W..........WWW.......W.............WW............W......W..............WW..........W.............W.............W..W..........WW......W........W.........WW..W..............W.....WW......W..W.
..........................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................................HHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHHH
[metadata] {"noOutputExpected": false}args
{
"command": "head -20000 data/pool.jsonl | python -c \"\nimport sys,json,re\nnh=0;n=0;samples=[]\nfor line in sys.stdin:\n r=json.loads(line); s=r['text']; n+=1\n if re.search(r'</?(p|pre|code|blockquote)[ >]', s):\n nh+=1\n if len(samples)<3: samples.append((r['id'],s[:300]))\nprint('html-ish docs:',nh,'/',n)\nfor i,s in samples: print('--',i,s.replace(chr(10),' | '))\n\"",
"description": "Check pool for HTML/code markup docs"
}Bash result
html-ish docs: 10 / 20000
-- 361 Twenty Eleven is the default theme for WordPress in 2011. There are tons of really cool improvements and features packed into it — more Post Formats, beautiful new header images, alternate color schemes and layouts, even custom link colors. All those features add up to the best part of Twenty Eleven
-- 2033 WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Related posts:Win a Trip to the Streets at Seven SpringsBurton Shakes Up Brand Lineu
-- 4313 Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology | | The Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on the web is defined, making it possible for the web to understand and satisfy the
[stdout]
html-ish docs: 10 / 20000
-- 361 Twenty Eleven is the default theme for WordPress in 2011. There are tons of really cool improvements and features packed into it — more Post Formats, beautiful new header images, alternate color schemes and layouts, even custom link colors. All those features add up to the best part of Twenty Eleven
-- 2033 WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Related posts:Win a Trip to the Streets at Seven SpringsBurton Shakes Up Brand Lineu
-- 4313 Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology | | The Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on the web is defined, making it possible for the web to understand and satisfy the
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "python -c \"\nimport numpy as np, json, re\nd=np.load('data/multi_dev.npy').astype(np.int64)\nEOS=50256\nidx=np.where(d==EOS)[0]\nt=json.load(open('work/dev_docs.json'))\n# doc token lengths incl eos\nbounds=np.concatenate([[-1],idx])\nlens=np.diff(np.concatenate([[-1],idx,[len(d)-1]]))\n# classify each doc\ndef kind(s):\n if re.search(r'</?(p|pre|code|blockquote|em|strong|ul|li|a)[ >]', s): return 'SE'\n if ('@-@' in s) or ('@,@' in s) or re.search(r\\\"\\s'(s|t|ll|ve|re|m|d)\\s\\\", s) or re.search(r'\\s[,.;:]\\s',s): return 'WIKI'\n return 'OTHER'\nks=[kind(x) for x in t]\n# find contiguous runs\nprint('counts', {k:ks.count(k) for k in set(ks)})\n# block boundaries: last WIKI index, first SE index\nwi=[i for i,k in enumerate(ks) if k=='WIKI']; se=[i for i,k in enumerate(ks) if k=='SE']\nprint('wiki run end approx', max(i for i in wi if i<1900), 'SE start', min(se))\n# token totals for blocks\nL=lens[:len(t)]\nb1=1691; b2=min(se)\nprint('tokens: wiki[0:%d]='%b1, L[:b1].sum(), ' mid[%d:%d]='%(b1,b2), L[b1:b2].sum(), ' SE[%d:]='%b2, L[b2:].sum(), 'total', L.sum())\n# mid block: split news vs web -- print a few around middle\nfor i in range(b1, b2, 90):\n print(i, repr(t[i][:110]))\n\" 2>&1 | tail -30",
"description": "Compute register token shares"
}Bash result
counts {'OTHER': 602, 'WIKI': 1744, 'SE': 359}
wiki run end approx 1888 SE start 2346
tokens: wiki[0:1691]= 247513 mid[1691:2346]= 502249 SE[2346:]= 250238 total 1000000
1691 ' On Friday , June 17 , a fourth person , later identified as disc jockey Steve Gregory Croes , was also arrest'
1781 'John Salter Questions Bellator 188 Opponent Jason Radcliffe’s Chin\n\nAfter closing out 2016 on a five-fight win'
1871 'The CRISPR-Cas9 system was first used to edit genes in 2013 but since then its use has exploded. What makes th'
1961 'Amid fresh debate regarding Section 377 of the IPC, two 20-year-old women from Uttar Pradesh tricked their fam'
2051 "Barcelona target Philippe Coutinho has been left out of Liverpool's squad for their Champions League play-off "
2141 'Hardik Pandya plays a shot during the second day of the third Test against Sri Lanka. (Getty Images)'
2231 'Director Feroz Abbas Khan, who is behind the Mughal-e-Azam musical to be staged here next month, says that all'
2321 'Amid tepid GDP growth numbers and criticism from small businesses and export houses, Finance Minister Arun Jai'
[stdout]
counts {'OTHER': 602, 'WIKI': 1744, 'SE': 359}
wiki run end approx 1888 SE start 2346
tokens: wiki[0:1691]= 247513 mid[1691:2346]= 502249 SE[2346:]= 250238 total 1000000
1691 ' On Friday , June 17 , a fourth person , later identified as disc jockey Steve Gregory Croes , was also arrest'
1781 'John Salter Questions Bellator 188 Opponent Jason Radcliffe’s Chin\n\nAfter closing out 2016 on a five-fight win'
1871 'The CRISPR-Cas9 system was first used to edit genes in 2013 but since then its use has exploded. What makes th'
1961 'Amid fresh debate regarding Section 377 of the IPC, two 20-year-old women from Uttar Pradesh tricked their fam'
2051 "Barcelona target Philippe Coutinho has been left out of Liverpool's squad for their Champions League play-off "
2141 'Hardik Pandya plays a shot during the second day of the third Test against Sri Lanka. (Getty Images)'
2231 'Director Feroz Abbas Khan, who is behind the Mughal-e-Azam musical to be staged here next month, says that all'
2321 'Amid tepid GDP growth numbers and criticism from small businesses and export houses, Finance Minister Arun Jai'
[metadata] {"noOutputExpected": false}args
{
"command": "grep -E \"^shard|DONE\" /workspace/work/tok.log; ls -la /workspace/work/tok/ | tail -5; cat /workspace/.timer/remaining_secs",
"description": "Check tokenization progress"
}Bash result
shard 0: 18052 docs
shard 5: 17793 docs
shard 1: 18091 docs
shard 4: 18180 docs
shard 3: 18594 docs
shard 2: 18339 docs
shard 6: 11412 docs
shard 11: 8872 docs
shard 8: 8993 docs
shard 9: 8636 docs
shard 10: 8935 docs
shard 13: 8359 docs
shard 7: 8779 docs
shard 12: 8981 docs
DONE
-rw-r--r-- 1 root root 142080 Aug 3 23:27 shard_11_meta.npy
-rw-r--r-- 1 root root 31280664 Aug 3 23:27 shard_12.npy
-rw-r--r-- 1 root root 143824 Aug 3 23:27 shard_12_meta.npy
-rw-r--r-- 1 root root 31473274 Aug 3 23:27 shard_13.npy
-rw-r--r-- 1 root root 133872 Aug 3 23:27 shard_13_meta.npy
5260
[stdout]
shard 0: 18052 docs
shard 5: 17793 docs
shard 1: 18091 docs
shard 4: 18180 docs
shard 3: 18594 docs
shard 2: 18339 docs
shard 6: 11412 docs
shard 11: 8872 docs
shard 8: 8993 docs
shard 9: 8636 docs
shard 10: 8935 docs
shard 13: 8359 docs
shard 7: 8779 docs
shard 12: 8981 docs
DONE
-rw-r--r-- 1 root root 142080 Aug 3 23:27 shard_11_meta.npy
-rw-r--r-- 1 root root 31280664 Aug 3 23:27 shard_12.npy
-rw-r--r-- 1 root root 143824 Aug 3 23:27 shard_12_meta.npy
-rw-r--r-- 1 root root 31473274 Aug 3 23:27 shard_13.npy
-rw-r--r-- 1 root root 133872 Aug 3 23:27 shard_13_meta.npy
5260
[metadata] {"noOutputExpected": false}Tokenization done in 2 min. Registers are exact quarters: wiki 247k / web-prose+news 502k / StackExchange 250k tokens. Now building the scoring pipeline — I'll featurize on canonicalized GPT-2 token ids so the WikiText spacing artifacts can't leak into the classifier.
args
{
"file_path": "/workspace/work/build_index.py",
"content": "\"\"\"Consolidate the tokenized shards into one id-ordered corpus + offset index.\"\"\"\nimport numpy as np, glob\n\nmetas, streams = [], []\nfor f in sorted(glob.glob(\"/workspace/work/tok/shard_*_meta.npy\")):\n m = np.load(f)\n metas.append(m)\n streams.append(np.load(f.replace(\"_meta\", \"\")))\n\nmeta = np.concatenate(metas) # [id, ntok] in file order\nstream = np.concatenate(streams) # matching token stream\nids, ntok = meta[:, 0], meta[:, 1]\nassert ntok.sum() == len(stream), (ntok.sum(), len(stream))\n\n# reorder so row k holds doc id k\norder = np.argsort(ids, kind=\"stable\")\nassert (ids[order] == np.arange(len(ids))).all(), \"ids must be 0..N-1 exactly once\"\nstarts_file = np.concatenate([[0], np.cumsum(ntok)])[:-1]\n\nntok_s = ntok[order]\nstarts_new = np.concatenate([[0], np.cumsum(ntok_s)])[:-1]\nout = np.empty_like(stream)\nfor k, src in enumerate(order):\n out[starts_new[k]:starts_new[k] + ntok_s[k]] = stream[starts_file[src]:starts_file[src] + ntok[src]]\n\nnp.save(\"/workspace/work/corpus.npy\", out)\nnp.save(\"/workspace/work/offsets.npy\", np.concatenate([starts_new, [len(out)]]).astype(np.int64))\nprint(\"docs\", len(ntok_s), \"tokens\", len(out), \"mean\", len(out) / len(ntok_s))\n"
}Write result
File created successfully at: /workspace/work/build_index.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/work/build_index.py", "content": "\"\"\"Consolidate the tokenized shards into one id-ordered corpus + offset index.\"\"\"\nimport numpy as np, glob\n\nmetas, streams = [], []\nfor f in sorted(glob.glob(\"/workspace/work/tok/shard_*_meta.npy\")):\n m = np.load(f)\n metas.append(m)\n streams.append(np.load(f.replace(\"_meta\", \"\")))\n\nmeta = np.concatenate(metas) # [id, ntok] in file order\nstream = np.concatenate(streams) # matching token stream\nids, ntok = meta[:, 0], meta[:, 1]\nassert ntok.sum() == len(stream), (ntok.sum(), len(stream))\n\n# reorder so row k holds doc id k\norder = np.argsort(ids, kind=\"stable\")\nassert (ids[order] == np.arange(len(ids))).all(), \"ids must be 0..N-1 exactly once\"\nstarts_file = np.concatenate([[0], np.cumsum(ntok)])[:-1]\n\nntok_s = ntok[order]\nstarts_new = np.concatenate([[0], np.cumsum(ntok_s)])[:-1]\nout = np.empty_like(stream)\nfor k, src in enumerate(order):\n out[starts_new[k]:starts_new[k] + ntok_s[k]] = stream[starts_file[src]:starts_file[src] + ntok[src]]\n\nnp.save(\"/workspace/work/corpus.npy\", out)\nnp.save(\"/workspace/work/offsets.npy\", np.concatenate([starts_new, [len(out)]]).astype(np.int64))\nprint(\"docs\", len(ntok_s), \"tokens\", len(out), \"mean\", len(out) / len(ntok_s))\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "time python work/build_index.py",
"description": "Build consolidated corpus index",
"timeout": 600000
}Bash result
docs 182016 tokens 200093788 mean 1099.3197740857947
real 0m1.007s
user 0m1.730s
sys 0m0.660s
[stdout]
docs 182016 tokens 200093788 mean 1099.3197740857947
real 0m1.007s
user 0m1.730s
sys 0m0.660s
[metadata] {"noOutputExpected": false}200M tokens indexed. Now the canonical-vocab map and the register classifier.
args
{
"file_path": "/workspace/work/score.py",
"content": "\"\"\"Score every pool document for similarity to each of the 4 target registers.\n\nMethod\n------\nThe disclosed dev target (`multi_dev.npy`) is 1M GPT-2 tokens made of four equal\nquarters: encyclopedic (WikiText-style), general high-quality web prose, news, and\ntechnical Q&A (StackExchange HTML). We fit a 5-way fastText-style linear classifier\n(bag of *canonicalised* GPT-2 token ids) over\n {register 0..3} vs {random pool background}\nand score each pool doc by its per-register log-odds against the background class.\n\nCanonicalisation is the key trick: the WikiText quarter is detokenised (\" 's\",\n\" , \", \" @-@ \"), so raw GPT-2 ids would let the classifier win by detecting\nwhitespace artefacts instead of register/content. Mapping every vocab entry to\n`decoded.strip().lower()` -> letters-only (or NUM / punctuation bucket) collapses\n\" ,\" with \",\" and \" the\" with \"the\", so only content survives.\n\"\"\"\nimport os, json, numpy as np, torch\n\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\nW = \"/workspace/work\"\nEOS = 50256\nDEV = \"/workspace/data/multi_dev.npy\"\nNREG = 4\nWIN = 192 # window length in tokens for classifier examples\nSEED = 0\ndev_t = torch.device(\"cuda\")\n\n\n# ---------------------------------------------------------------- canonical vocab\ndef build_canon():\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n canon = np.zeros(50257, dtype=np.int64)\n slot, names = {}, []\n for t in range(50257):\n s = tok.decode([t]).strip().lower()\n letters = \"\".join(c for c in s if c.isalpha())\n if letters:\n key = letters\n elif any(c.isdigit() for c in s):\n key = \"<num>\"\n elif s == \"\":\n key = \"<sp>\"\n else:\n key = \"P\" + s\n if key not in slot:\n slot[key] = len(names); names.append(key)\n canon[t] = slot[key]\n return canon, names\n\n\nif os.path.exists(f\"{W}/canon.npy\"):\n canon = np.load(f\"{W}/canon.npy\")\n names = json.load(open(f\"{W}/canon_names.json\"))\nelse:\n canon, names = build_canon()\n np.save(f\"{W}/canon.npy\", canon)\n json.dump(names, open(f\"{W}/canon_names.json\", \"w\"))\nV = len(names)\nprint(\"canonical vocab\", V, flush=True)\n\ncorpus = np.load(f\"{W}/corpus.npy\")\noffs = np.load(f\"{W}/offsets.npy\")\nndoc = len(offs) - 1\ncanon_t = torch.from_numpy(canon).to(dev_t)\n\n# EOS and the wikitext \"@\" filler carry pure-format information -> mask them out.\nDROP = {canon[EOS], canon[np.int64(31)]} # <|endoftext|>, \"@\"\ndrop_mask = torch.zeros(V, dtype=torch.bool, device=dev_t)\nfor d in DROP:\n drop_mask[int(d)] = True\n\n\ndef windows(stream, win=WIN):\n \"\"\"Split a 1-D token array into non-overlapping windows (drop the remainder).\"\"\"\n n = (len(stream) // win) * win\n if n == 0:\n return np.zeros((0, win), dtype=stream.dtype)\n return stream[:n].reshape(-1, win)\n\n\ndef bags(win_mat):\n \"\"\"windows (M,WIN) of raw token ids -> L2-normalised canonical count matrix (M,V).\"\"\"\n x = torch.from_numpy(win_mat.astype(np.int64)).to(dev_t)\n c = canon_t[x]\n out = torch.zeros(len(x), V, device=dev_t)\n out.scatter_add_(1, c, torch.ones_like(c, dtype=torch.float))\n out[:, drop_mask] = 0.0\n out = torch.log1p(out) # damp repetition\n return out / out.norm(dim=1, keepdim=True).clamp_min(1e-6)\n\n\n# ---------------------------------------------------------------- register windows\ndev = np.load(DEV).astype(np.int64)\nq = len(dev) // NREG\nreg_wins = [windows(dev[i * q:(i + 1) * q]) for i in range(NREG)]\nprint(\"register windows\", [len(w) for w in reg_wins], flush=True)\n\nrng = np.random.default_rng(SEED)\nbg_docs = rng.choice(ndoc, size=6000, replace=False)\nbg = np.concatenate([windows(corpus[offs[i]:offs[i + 1]])[:3] for i in bg_docs])\nprint(\"background windows\", len(bg), flush=True)\n\nXw = [bags(w) for w in reg_wins] + [bags(bg)]\nX = torch.cat(Xw)\ny = torch.cat([torch.full((len(b),), k, device=dev_t, dtype=torch.long)\n for k, b in enumerate(Xw)])\n# class-balanced weights so the small registers are not swamped by background\ncnt = torch.bincount(y, minlength=NREG + 1).float()\ncw = (cnt.sum() / (len(cnt) * cnt)).to(dev_t)\n\ntorch.manual_seed(SEED)\nlin = torch.nn.Linear(V, NREG + 1).to(dev_t)\ntorch.nn.init.zeros_(lin.weight); torch.nn.init.zeros_(lin.bias)\noptim = torch.optim.AdamW(lin.parameters(), lr=3e-3, weight_decay=1e-4)\n# held-out split to confirm the classifier learned register, not noise\nperm = torch.randperm(len(X), device=dev_t)\nntr = int(0.9 * len(X)); tr, va = perm[:ntr], perm[ntr:]\nfor ep in range(320):\n optim.zero_grad()\n loss = torch.nn.functional.cross_entropy(lin(X[tr]), y[tr], weight=cw)\n loss.backward(); optim.step()\n if (ep + 1) % 80 == 0:\n with torch.no_grad():\n acc = (lin(X[va]).argmax(1) == y[va]).float().mean().item()\n print(f\" ep{ep+1} loss {loss.item():.4f} val_acc {acc:.3f}\", flush=True)\nwith torch.no_grad():\n pv = lin(X[va]).argmax(1)\n for k in range(NREG + 1):\n m = y[va] == k\n if m.any():\n print(f\" class {k} recall {(pv[m] == k).float().mean().item():.3f}\", flush=True)\n\n# ---------------------------------------------------------------- score the pool\nBATCH = 4096\nscores = np.zeros((ndoc, NREG), dtype=np.float32)\nnwin_used = np.zeros(ndoc, dtype=np.int32)\nrep = np.zeros(ndoc, dtype=np.float32) # lexical diversity (uniq/total wordpiece)\nMAXW = 8 # cap windows/doc: bounds cost, enough signal\nwith torch.no_grad():\n b_idx, b_win, b_own = [], [], []\n\n def flush(b_win, b_own):\n if not b_win:\n return\n mat = np.concatenate(b_win)\n own = np.concatenate(b_own)\n lg = torch.log_softmax(lin(bags(mat)), dim=1)\n lo = (lg[:, :NREG] - lg[:, NREG:NREG + 1]).cpu().numpy() # log-odds vs background\n own_t = own\n for k in range(NREG):\n np.add.at(scores[:, k], own_t, lo[:, k])\n np.add.at(nwin_used, own_t, 1)\n\n for i in range(ndoc):\n s = corpus[offs[i]:offs[i + 1]]\n c = canon[s.astype(np.int64)]\n rep[i] = len(np.unique(c)) / max(1, len(c))\n w = windows(s)[:MAXW]\n if len(w):\n b_win.append(w); b_own.append(np.full(len(w), i))\n if sum(len(x) for x in b_win) >= BATCH:\n flush(b_win, b_own); b_win, b_own = [], []\n if (i + 1) % 40000 == 0:\n print(\"scored\", i + 1, flush=True)\n flush(b_win, b_own)\n\nscores /= np.maximum(1, nwin_used)[:, None]\nnp.save(f\"{W}/reg_scores.npy\", scores)\nnp.save(f\"{W}/nwin.npy\", nwin_used)\nnp.save(f\"{W}/rep.npy\", rep)\nprint(\"saved scores\", scores.shape, flush=True)\nfor k in range(NREG):\n print(f\"register {k}: mean {scores[:,k].mean():.3f} p99 {np.percentile(scores[:,k],99):.3f}\")\n"
}Write result
File created successfully at: /workspace/work/score.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/work/score.py", "content": "\"\"\"Score every pool document for similarity to each of the 4 target registers.\n\nMethod\n------\nThe disclosed dev target (`multi_dev.npy`) is 1M GPT-2 tokens made of four equal\nquarters: encyclopedic (WikiText-style), general high-quality web prose, news, and\ntechnical Q&A (StackExchange HTML). We fit a 5-way fastText-style linear classifier\n(bag of *canonicalised* GPT-2 token ids) over\n {register 0..3} vs {random pool background}\nand score each pool doc by its per-register log-odds against the background class.\n\nCanonicalisation is the key trick: the WikiText quarter is detokenised (\" 's\",\n\" , \", \" @-@ \"), so raw GPT-2 ids would let the classifier win by detecting\nwhitespace artefacts instead of register/content. Mapping every vocab entry to\n`decoded.strip().lower()` -> letters-only (or NUM / punctuation bucket) collapses\n\" ,\" with \",\" and \" the\" with \"the\", so only content survives.\n\"\"\"\nimport os, json, numpy as np, torch\n\nos.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"\nW = \"/workspace/work\"\nEOS = 50256\nDEV = \"/workspace/data/multi_dev.npy\"\nNREG = 4\nWIN = 192 # window length in tokens for classifier examples\nSEED = 0\ndev_t = torch.device(\"cuda\")\n\n\n# ---------------------------------------------------------------- canonical vocab\ndef build_canon():\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n canon = np.zeros(50257, dtype=np.int64)\n slot, names = {}, []\n for t in range(50257):\n s = tok.decode([t]).strip().lower()\n letters = \"\".join(c for c in s if c.isalpha())\n if letters:\n key = letters\n elif any(c.isdigit() for c in s):\n key = \"<num>\"\n elif s == \"\":\n key = \"<sp>\"\n else:\n key = \"P\" + s\n if key not in slot:\n slot[key] = len(names); names.append(key)\n canon[t] = slot[key]\n return canon, names\n\n\nif os.path.exists(f\"{W}/canon.npy\"):\n canon = np.load(f\"{W}/canon.npy\")\n names = json.load(open(f\"{W}/canon_names.json\"))\nelse:\n canon, names = build_canon()\n np.save(f\"{W}/canon.npy\", canon)\n json.dump(names, open(f\"{W}/canon_names.json\", \"w\"))\nV = len(names)\nprint(\"canonical vocab\", V, flush=True)\n\ncorpus = np.load(f\"{W}/corpus.npy\")\noffs = np.load(f\"{W}/offsets.npy\")\nndoc = len(offs) - 1\ncanon_t = torch.from_numpy(canon).to(dev_t)\n\n# EOS and the wikitext \"@\" filler carry pure-format information -> mask them out.\nDROP = {canon[EOS], canon[np.int64(31)]} # <|endoftext|>, \"@\"\ndrop_mask = torch.zeros(V, dtype=torch.bool, device=dev_t)\nfor d in DROP:\n drop_mask[int(d)] = True\n\n\ndef windows(stream, win=WIN):\n \"\"\"Split a 1-D token array into non-overlapping windows (drop the remainder).\"\"\"\n n = (len(stream) // win) * win\n if n == 0:\n return np.zeros((0, win), dtype=stream.dtype)\n return stream[:n].reshape(-1, win)\n\n\ndef bags(win_mat):\n \"\"\"windows (M,WIN) of raw token ids -> L2-normalised canonical count matrix (M,V).\"\"\"\n x = torch.from_numpy(win_mat.astype(np.int64)).to(dev_t)\n c = canon_t[x]\n out = torch.zeros(len(x), V, device=dev_t)\n out.scatter_add_(1, c, torch.ones_like(c, dtype=torch.float))\n out[:, drop_mask] = 0.0\n out = torch.log1p(out) # damp repetition\n return out / out.norm(dim=1, keepdim=True).clamp_min(1e-6)\n\n\n# ---------------------------------------------------------------- register windows\ndev = np.load(DEV).astype(np.int64)\nq = len(dev) // NREG\nreg_wins = [windows(dev[i * q:(i + 1) * q]) for i in range(NREG)]\nprint(\"register windows\", [len(w) for w in reg_wins], flush=True)\n\nrng = np.random.default_rng(SEED)\nbg_docs = rng.choice(ndoc, size=6000, replace=False)\nbg = np.concatenate([windows(corpus[offs[i]:offs[i + 1]])[:3] for i in bg_docs])\nprint(\"background windows\", len(bg), flush=True)\n\nXw = [bags(w) for w in reg_wins] + [bags(bg)]\nX = torch.cat(Xw)\ny = torch.cat([torch.full((len(b),), k, device=dev_t, dtype=torch.long)\n for k, b in enumerate(Xw)])\n# class-balanced weights so the small registers are not swamped by background\ncnt = torch.bincount(y, minlength=NREG + 1).float()\ncw = (cnt.sum() / (len(cnt) * cnt)).to(dev_t)\n\ntorch.manual_seed(SEED)\nlin = torch.nn.Linear(V, NREG + 1).to(dev_t)\ntorch.nn.init.zeros_(lin.weight); torch.nn.init.zeros_(lin.bias)\noptim = torch.optim.AdamW(lin.parameters(), lr=3e-3, weight_decay=1e-4)\n# held-out split to confirm the classifier learned register, not noise\nperm = torch.randperm(len(X), device=dev_t)\nntr = int(0.9 * len(X)); tr, va = perm[:ntr], perm[ntr:]\nfor ep in range(320):\n optim.zero_grad()\n loss = torch.nn.functional.cross_entropy(lin(X[tr]), y[tr], weight=cw)\n loss.backward(); optim.step()\n if (ep + 1) % 80 == 0:\n with torch.no_grad():\n acc = (lin(X[va]).argmax(1) == y[va]).float().mean().item()\n print(f\" ep{ep+1} loss {loss.item():.4f} val_acc {acc:.3f}\", flush=True)\nwith torch.no_grad():\n pv = lin(X[va]).argmax(1)\n for k in range(NREG + 1):\n m = y[va] == k\n if m.any():\n print(f\" class {k} recall {(pv[m] == k).float().mean().item():.3f}\", flush=True)\n\n# ---------------------------------------------------------------- score the pool\nBATCH = 4096\nscores = np.zeros((ndoc, NREG), dtype=np.float32)\nnwin_used = np.zeros(ndoc, dtype=np.int32)\nrep = np.zeros(ndoc, dtype=np.float32) # lexical diversity (uniq/total wordpiece)\nMAXW = 8 # cap windows/doc: bounds cost, enough signal\nwith torch.no_grad():\n b_idx, b_win, b_own = [], [], []\n\n def flush(b_win, b_own):\n if not b_win:\n return\n mat = np.concatenate(b_win)\n own = np.concatenate(b_own)\n lg = torch.log_softmax(lin(bags(mat)), dim=1)\n lo = (lg[:, :NREG] - lg[:, NREG:NREG + 1]).cpu().numpy() # log-odds vs background\n own_t = own\n for k in range(NREG):\n np.add.at(scores[:, k], own_t, lo[:, k])\n np.add.at(nwin_used, own_t, 1)\n\n for i in range(ndoc):\n s = corpus[offs[i]:offs[i + 1]]\n c = canon[s.astype(np.int64)]\n rep[i] = len(np.unique(c)) / max(1, len(c))\n w = windows(s)[:MAXW]\n if len(w):\n b_win.append(w); b_own.append(np.full(len(w), i))\n if sum(len(x) for x in b_win) >= BATCH:\n flush(b_win, b_own); b_win, b_own = [], []\n if (i + 1) % 40000 == 0:\n print(\"scored\", i + 1, flush=True)\n flush(b_win, b_own)\n\nscores /= np.maximum(1, nwin_used)[:, None]\nnp.save(f\"{W}/reg_scores.npy\", scores)\nnp.save(f\"{W}/nwin.npy\", nwin_used)\nnp.save(f\"{W}/rep.npy\", rep)\nprint(\"saved scores\", scores.shape, flush=True)\nfor k in range(NREG):\n print(f\"register {k}: mean {scores[:,k].mean():.3f} p99 {np.percentile(scores[:,k],99):.3f}\")\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "time python work/score.py 2>&1 | tail -40",
"description": "Train register classifier and score pool",
"timeout": 900000
}Bash result
canonical vocab 31652
register windows [1302, 1302, 1302, 1302]
background windows 11567
ep80 loss 0.7925 val_acc 0.787
ep160 loss 0.4620 val_acc 0.857
ep240 loss 0.3088 val_acc 0.880
ep320 loss 0.2239 val_acc 0.902
class 0 recall 0.983
class 1 recall 0.814
class 2 recall 0.903
class 3 recall 1.000
class 4 recall 0.891
scored 40000
scored 80000
scored 120000
scored 160000
saved scores (182016, 4)
register 0: mean -2.317 p99 0.184
register 1: mean -1.624 p99 0.627
register 2: mean -2.103 p99 0.379
register 3: mean -2.598 p99 0.000
real 0m17.221s
user 0m16.100s
sys 0m1.530s
[stdout]
canonical vocab 31652
register windows [1302, 1302, 1302, 1302]
background windows 11567
ep80 loss 0.7925 val_acc 0.787
ep160 loss 0.4620 val_acc 0.857
ep240 loss 0.3088 val_acc 0.880
ep320 loss 0.2239 val_acc 0.902
class 0 recall 0.983
class 1 recall 0.814
class 2 recall 0.903
class 3 recall 1.000
class 4 recall 0.891
scored 40000
scored 80000
scored 120000
scored 160000
saved scores (182016, 4)
register 0: mean -2.317 p99 0.184
register 1: mean -1.624 p99 0.627
register 2: mean -2.103 p99 0.379
register 3: mean -2.598 p99 0.000
real 0m17.221s
user 0m16.100s
sys 0m1.530s
[metadata] {"noOutputExpected": false}Classifier: 90% val accuracy, all four registers separated from background. Let me eyeball what it's actually selecting before I commit.
args
{
"command": "python -c \"\nimport numpy as np, json\nos_=__import__('os'); os_.environ['TOKENIZERS_PARALLELISM']='false'\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained('gpt2')\nS=np.load('work/reg_scores.npy'); C=np.load('work/corpus.npy'); O=np.load('work/offsets.npy')\nnames=['WIKI','WEBPROSE','NEWS','TECH-QA']\ndef show(i,n=230):\n t=tok.decode(C[O[i]:O[i]+90].astype(np.int64))\n return t.replace(chr(10),' | ')[:n]\nfor k in range(4):\n o=np.argsort(-S[:,k])[:4]\n print('#####',names[k], 'top scores', np.round(S[o,k],2))\n for i in o: print(' ',i,'|',show(i))\nprint()\nprint('##### WORST overall (max over registers)')\nmx=S.max(1)\nfor i in np.argsort(mx)[:4]: print(' ',i, round(mx[i],2),'|',show(i))\nprint()\nntok=np.diff(O)\nprint('token stats: total',ntok.sum(),'median',np.median(ntok))\n\"\n",
"description": "Inspect top-scoring docs per register"
}Bash result
##### WIKI top scores [2.76 2.44 2.36 2.24]
82269 | Hill 303 massacre | |Hill 303 massacre| | Bodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound | |Location||Hill 303, Waegwan, South Korea| | |Date||August 17, 1950 | |Target||U.S. Army
50793 | Majesty King Peter II of Yugoslavia was the firstborn son of King Alexander I and Queen Maria of Yugoslavia. King Peter II was born in Belgrade 6 September 1923 his Godparents were King George VI and Queen Elizabeth (later Queen
3475 | Anglo-Dutch Wars, also called Dutch Wars, Dutch Engelse Oorlogen, four 17th- and 18th-century naval conflicts between England and the Dutch Republic. The first three wars, stemming from commercial rivalry, established England’s na
14644 | The War In 1861 | In response to popular cries of On to Richmond, a Union army under Brigadier General Irvin McDowell began on July 16 a slow advance from near Washington toward the Confederate capital. On July 21, the Union force
##### WEBPROSE top scores [2.85 2.51 2.35 2.28]
47889 | GAINESVILLE, Fla. (AP) — Tim Kaine is shrugging off any possibility that he could be embarrassed by the release of hacked emails. | WikiLeaks, which has been posting stolen emails from Hillary Clinton's campaign manager John Podes
66305 | Bradford West Respect MP George Galloway has defended his controversial claim that a sex assault allegation against WikiLeaks campaigner Julian Assange amounted to no more than bad "sexual etiquette". | Mr Galloway provoked a furi
12295 | Here’s what you need to know for Tuesday morning. | Yesterday the Japanese government detected a ballistic missile launch from North Korea, and warned its citizens to prepare for a possible attack as the missile flew over northern
12336 | So the Free Syrian Army are fighting Assad but also Daesh. | Assad loyalists are also fighting both the FSA and Daesh. | Meanwhile, just about everyone is trying to kill the Kurds. Who, ironically, are probably the most effective
##### NEWS top scores [3.33 3.26 3.22 3.19]
103973 | pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crowds, he said, "We have sent funds under various national schemes to Uttar Pradesh. Rs 24000 crore have been giv
81859 | |Rediff India Abroad Home | All the sections| | Bihar: Vigilante justice resurfaces, three people lynched | February 18, 2008 17:20 IST | Fresh incidents of vigilante justice have been reported from Bihar with three suspected thie
28143 | <|endoftext|>Ahead of its Foreign Minister's visit to Bangalore, China on Tuesday described the Kashmir issue as a question "left over by history" and highlighted the need for India and Pakistan to "properly" resolve it through di
37522 | Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's national elections. | "I congratulate Prime Minister Modi on the electoral victory of BJP and allies. Look forward t
##### TECH-QA top scores [3.26 3.18 3.15 3.14]
167171 | , dibs, shotgun | '; var startBlockContent = ' | '; var endBlockContent = ' | '; var endBlock = ' | '; widgetContent = startBlockContent + widgetContent + endBlockContent; if (widgetTitle && true) { widgetContent = startBlockHeade
169305 | Sitemap | RSS | Trademark<|endoftext|>Separated by a Common Language: scatological adjectives | '; var startBlockContent = ' | '; var endBlockContent = ' | '; var endBlock = ' | '; widgetContent = startBlockContent + widgetConten
171615 | opy<|endoftext|>K FASHION V-neck Collarless Blazer | Men's Clothing | GOBIZKOREA.COM | '; //} addCnt.push(ctgCd+","+gCnt); gCnt = 0; } rootCnt++; dtlCnt = 0; dtl += ' | '+ac.ctgry_nm_eng_dflt+' | '; c
162020 | o de regularización de 10.000 vehículos públicos ' : ''; var trtd = ' | '+posttitle+' | '+removeHtmlTag(postcontent,summaryPost)+'... | '; document.write(trtd); j++; } } function showrecentposts1a(json) { j = (showRandomImg) ? Mat
##### WORST overall (max over registers)
115858 -6.31 | Yang<|endoftext|>Free Classifieds Ads Online - Buynow-us | Search | Post your ad! | My Account | Sign In | Register | Favorites | All Categories Vehicles Cars Motorcycles Boats RVs Commercial Trucks Aircrafts Parts & Accessories
138514 -6.31 | Yang<|endoftext|>Free Classifieds Ads Online - Buynow-us | Search | Post your ad! | My Account | Sign In | Register | Favorites | All Categories Vehicles Cars Motorcycles Boats RVs Commercial Trucks Aircrafts Parts & Accessories
137433 -6.17 | ug<|endoftext|>Post Taged with Breckenridge Mountain Map — | Sendiksonoakland.com | Show Menu | Search | All posts tagged Breckenridge Mountain Map | Posted in Bathrooms by Alessa on December 14, 2018 | Gorgeous Rustic Copper Bath
114777 -6.17 | ug<|endoftext|>Post Taged with Breckenridge Mountain Map — | Sendiksonoakland.com | Show Menu | Search | All posts tagged Breckenridge Mountain Map | Posted in Bathrooms by Alessa on December 14, 2018 | Gorgeous Rustic Copper Bath
token stats: total 200093788 median 540.0
[stdout]
##### WIKI top scores [2.76 2.44 2.36 2.24]
82269 | Hill 303 massacre | |Hill 303 massacre| | Bodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound | |Location||Hill 303, Waegwan, South Korea| | |Date||August 17, 1950 | |Target||U.S. Army
50793 | Majesty King Peter II of Yugoslavia was the firstborn son of King Alexander I and Queen Maria of Yugoslavia. King Peter II was born in Belgrade 6 September 1923 his Godparents were King George VI and Queen Elizabeth (later Queen
3475 | Anglo-Dutch Wars, also called Dutch Wars, Dutch Engelse Oorlogen, four 17th- and 18th-century naval conflicts between England and the Dutch Republic. The first three wars, stemming from commercial rivalry, established England’s na
14644 | The War In 1861 | In response to popular cries of On to Richmond, a Union army under Brigadier General Irvin McDowell began on July 16 a slow advance from near Washington toward the Confederate capital. On July 21, the Union force
##### WEBPROSE top scores [2.85 2.51 2.35 2.28]
47889 | GAINESVILLE, Fla. (AP) — Tim Kaine is shrugging off any possibility that he could be embarrassed by the release of hacked emails. | WikiLeaks, which has been posting stolen emails from Hillary Clinton's campaign manager John Podes
66305 | Bradford West Respect MP George Galloway has defended his controversial claim that a sex assault allegation against WikiLeaks campaigner Julian Assange amounted to no more than bad "sexual etiquette". | Mr Galloway provoked a furi
12295 | Here’s what you need to know for Tuesday morning. | Yesterday the Japanese government detected a ballistic missile launch from North Korea, and warned its citizens to prepare for a possible attack as the missile flew over northern
12336 | So the Free Syrian Army are fighting Assad but also Daesh. | Assad loyalists are also fighting both the FSA and Daesh. | Meanwhile, just about everyone is trying to kill the Kurds. Who, ironically, are probably the most effective
##### NEWS top scores [3.33 3.26 3.22 3.19]
103973 | pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crowds, he said, "We have sent funds under various national schemes to Uttar Pradesh. Rs 24000 crore have been giv
81859 | |Rediff India Abroad Home | All the sections| | Bihar: Vigilante justice resurfaces, three people lynched | February 18, 2008 17:20 IST | Fresh incidents of vigilante justice have been reported from Bihar with three suspected thie
28143 | <|endoftext|>Ahead of its Foreign Minister's visit to Bangalore, China on Tuesday described the Kashmir issue as a question "left over by history" and highlighted the need for India and Pakistan to "properly" resolve it through di
37522 | Prime Minister Imran Khan on Thursday congratulated India’s premier Narendra Modi on his victory in the country's national elections. | "I congratulate Prime Minister Modi on the electoral victory of BJP and allies. Look forward t
##### TECH-QA top scores [3.26 3.18 3.15 3.14]
167171 | , dibs, shotgun | '; var startBlockContent = ' | '; var endBlockContent = ' | '; var endBlock = ' | '; widgetContent = startBlockContent + widgetContent + endBlockContent; if (widgetTitle && true) { widgetContent = startBlockHeade
169305 | Sitemap | RSS | Trademark<|endoftext|>Separated by a Common Language: scatological adjectives | '; var startBlockContent = ' | '; var endBlockContent = ' | '; var endBlock = ' | '; widgetContent = startBlockContent + widgetConten
171615 | opy<|endoftext|>K FASHION V-neck Collarless Blazer | Men's Clothing | GOBIZKOREA.COM | '; //} addCnt.push(ctgCd+","+gCnt); gCnt = 0; } rootCnt++; dtlCnt = 0; dtl += ' | '+ac.ctgry_nm_eng_dflt+' | '; c
162020 | o de regularización de 10.000 vehículos públicos ' : ''; var trtd = ' | '+posttitle+' | '+removeHtmlTag(postcontent,summaryPost)+'... | '; document.write(trtd); j++; } } function showrecentposts1a(json) { j = (showRandomImg) ? Mat
##### WORST overall (max over registers)
115858 -6.31 | Yang<|endoftext|>Free Classifieds Ads Online - Buynow-us | Search | Post your ad! | My Account | Sign In | Register | Favorites | All Categories Vehicles Cars Motorcycles Boats RVs Commercial Trucks Aircrafts Parts & Accessories
138514 -6.31 | Yang<|endoftext|>Free Classifieds Ads Online - Buynow-us | Search | Post your ad! | My Account | Sign In | Register | Favorites | All Categories Vehicles Cars Motorcycles Boats RVs Commercial Trucks Aircrafts Parts & Accessories
137433 -6.17 | ug<|endoftext|>Post Taged with Breckenridge Mountain Map — | Sendiksonoakland.com | Show Menu | Search | All posts tagged Breckenridge Mountain Map | Posted in Bathrooms by Alessa on December 14, 2018 | Gorgeous Rustic Copper Bath
114777 -6.17 | ug<|endoftext|>Post Taged with Breckenridge Mountain Map — | Sendiksonoakland.com | Show Menu | Search | All posts tagged Breckenridge Mountain Map | Posted in Bathrooms by Alessa on December 14, 2018 | Gorgeous Rustic Copper Bath
token stats: total 200093788 median 540.0
[metadata] {"noOutputExpected": false}Crucial finding: the TECH-QA register is matching JavaScript boilerplate — the classifier learned SE's HTML markup rather than technical language. Also the pool has exact duplicates (115858≡138514). Let me fix both: strip HTML from the SE target before fitting, add a quality gate, and dedup.
args
{
"file_path": "/workspace/submission/curate.py",
"content": "\"\"\"Curate /workspace/data/pool.jsonl -> submission/selection.json (priority ordered).\n\nSTATED CRITERION\n================\nThe training budget (12M tokens) is ~6% of the pool, and the evaluation target is a\n*known* mixture: four equal quarters of (0) encyclopedic WikiText-style prose,\n(1) general high-quality web prose, (2) news, (3) technical Q&A. Under a fixed token\nbudget, cross-entropy on such a target is minimised by training on data drawn from the\nsame mixture. So we (a) learn what each target register looks like, (b) reject\nmachine-generated web boilerplate that no register contains, and (c) fill the budget\nwith a register-balanced, deduplicated selection, best-scoring documents first.\n\nConcretely, a document is selected if it passes a register-agnostic quality gate, is\nnot a near-duplicate of an already-accepted document, and ranks highly for one of the\nfour registers under a linear register classifier.\n\n score_r(doc) = mean over the doc's 192-token windows of\n [ log P(register r | window) - log P(background | window) ]\n\nwhere P is a 5-way multinomial logistic regression over bags of *canonicalised* GPT-2\ntoken ids, fit on windows of the disclosed dev target (4 register classes) against\nwindows of a random pool sample (the \"background\" class).\n\nTwo details matter and are the difference between this working and not working:\n\n* CANONICALISATION. The encyclopedic quarter of the target is detokenised WikiText\n (\" 's\", \" , \", \" @-@ \"). Raw GPT-2 ids would let the classifier separate target from\n pool by *whitespace artefacts* rather than by register, and the resulting score would\n be meaningless on pool text. We therefore map every vocab entry through\n `decode(t).strip().lower()` -> letters-only, else <num> / <sp> / punctuation bucket.\n This collapses \" ,\" with \",\" and \" The\" with \"the\", leaving only content.\n\n* HTML STRIPPING OF THE TECHNICAL REGISTER. The technical quarter is StackExchange with\n markup (`<p>`, `<pre><code>`, `>`). Fit on raw markup, the classifier's top-ranked\n pool documents are scraped JavaScript/template boilerplate -- symbol soup that shares\n the markup but none of the language. We strip tags and entities before fitting, so\n register 3 selects genuine technical prose.\n\nUsage: python curate.py [--out selection.json]\nRequires the pool pre-tokenised by tokenize_pool.py + build_index.py (GPT-2 BPE);\nboth are pure preprocessing (no labels, no external data).\n\"\"\"\nimport argparse, html, json, os, re, sys\nimport numpy as np\nimport torch\n\nos.environ.setdefault(\"TOKENIZERS_PARALLELISM\", \"false\")\n\nW = \"/workspace/work\"\nDEV = \"/workspace/data/multi_dev.npy\"\nEOS = 50256\nNREG = 4\nWIN = 192 # classifier window length, in tokens\nMAXW = 8 # windows scored per document (cost bound; ~1.5k tokens)\nSEED = 0\nBUDGET = 12_000_000\n\nap = argparse.ArgumentParser()\nap.add_argument(\"--out\", default=\"/workspace/submission/selection.json\")\nap.add_argument(\"--budget_mult\", type=float, default=2.5,\n help=\"emit this many times the token budget, so ids beyond it are spare\")\nap.add_argument(\"--shares\", default=\"0.25,0.25,0.25,0.25\",\n help=\"target token share per register (wiki,webprose,news,techqa)\")\nap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nargs = ap.parse_args()\ndev_t = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n\n# ------------------------------------------------------------------ canonical vocab\ndef build_canon():\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n canon, slot, names = np.zeros(50257, dtype=np.int64), {}, []\n for t in range(50257):\n s = tok.decode([t]).strip().lower()\n letters = \"\".join(c for c in s if c.isalpha())\n key = letters if letters else (\"<num>\" if any(c.isdigit() for c in s)\n else (\"<sp>\" if s == \"\" else \"P\" + s))\n if key not in slot:\n slot[key] = len(names); names.append(key)\n canon[t] = slot[key]\n return canon, names\n\n\nif os.path.exists(f\"{W}/canon.npy\"):\n canon = np.load(f\"{W}/canon.npy\"); names = json.load(open(f\"{W}/canon_names.json\"))\nelse:\n canon, names = build_canon()\n np.save(f\"{W}/canon.npy\", canon); json.dump(names, open(f\"{W}/canon_names.json\", \"w\"))\nV = len(names)\nname2id = {n: i for i, n in enumerate(names)}\n\ncorpus = np.load(f\"{W}/corpus.npy\")\noffs = np.load(f\"{W}/offsets.npy\")\nndoc = len(offs) - 1\nntok_doc = np.diff(offs)\ncanon_t = torch.from_numpy(canon).to(dev_t)\nprint(f\"pool: {ndoc} docs, {ntok_doc.sum()} tokens\", flush=True)\n\n# format-only tokens carry no register information -> zero them out\ndrop_mask = torch.zeros(V, dtype=torch.bool, device=dev_t)\nfor k in (canon[EOS], canon[31]): # <|endoftext|>, \"@\" (WikiText filler)\n drop_mask[int(k)] = True\n\n\ndef windows(stream, win=WIN):\n n = (len(stream) // win) * win\n return (stream[:n].reshape(-1, win) if n else\n np.zeros((0, win), dtype=stream.dtype))\n\n\ndef bags(win_mat):\n x = torch.from_numpy(np.ascontiguousarray(win_mat).astype(np.int64)).to(dev_t)\n c = canon_t[x]\n out = torch.zeros(len(x), V, device=dev_t)\n out.scatter_add_(1, c, torch.ones_like(c, dtype=torch.float))\n out[:, drop_mask] = 0.0\n out = torch.log1p(out) # damp within-window repetition\n return out / out.norm(dim=1, keepdim=True).clamp_min(1e-6)\n\n\n# ------------------------------------------------- target registers (quarters of dev)\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).astype(np.int64)\nq = len(dev) // NREG\nreg_wins = []\nfor i in range(NREG):\n seg = dev[i * q:(i + 1) * q]\n txt = tok.decode(seg)\n if re.search(r\"</?(p|pre|code|blockquote)[ >]\", txt): # the technical quarter\n txt = html.unescape(re.sub(r\"<[^>\\n]{1,60}>\", \" \", txt))\n seg = np.array(tok(txt, add_special_tokens=False).input_ids, dtype=np.int64)\n print(\" stripped markup from register\", i, flush=True)\n reg_wins.append(windows(seg))\nprint(\"register windows\", [len(w) for w in reg_wins], flush=True)\n\nrng = np.random.default_rng(SEED)\nbg_docs = rng.choice(ndoc, size=6000, replace=False)\nbg = np.concatenate([windows(corpus[offs[i]:offs[i + 1]])[:3] for i in bg_docs])\n\nXw = [bags(w) for w in reg_wins] + [bags(bg)]\nX = torch.cat(Xw)\ny = torch.cat([torch.full((len(b),), k, device=dev_t, dtype=torch.long)\n for k, b in enumerate(Xw)])\ncnt = torch.bincount(y, minlength=NREG + 1).float()\ncw = cnt.sum() / (len(cnt) * cnt) # class-balanced\n\ntorch.manual_seed(SEED)\nlin = torch.nn.Linear(V, NREG + 1).to(dev_t)\ntorch.nn.init.zeros_(lin.weight); torch.nn.init.zeros_(lin.bias)\noptim = torch.optim.AdamW(lin.parameters(), lr=3e-3, weight_decay=1e-4)\nperm = torch.randperm(len(X), device=dev_t)\nntr = int(0.9 * len(X)); tr, va = perm[:ntr], perm[ntr:]\nfor ep in range(320):\n optim.zero_grad()\n torch.nn.functional.cross_entropy(lin(X[tr]), y[tr], weight=cw).backward()\n optim.step()\nwith torch.no_grad():\n pv = lin(X[va]).argmax(1)\n print(\"classifier val acc %.3f | per-class recall %s\" % (\n (pv == y[va]).float().mean().item(),\n \" \".join(f\"{k}:{(pv[y[va]==k]==k).float().mean().item():.2f}\" for k in range(NREG + 1))),\n flush=True)\n\n# ------------------------------------------------------------ score + gate the pool\nPUNCT = np.array([i for i, n in enumerate(names) if n.startswith(\"P\")], dtype=np.int64)\nis_punct = np.zeros(V, dtype=bool); is_punct[PUNCT] = True\nis_word = np.array([not n.startswith(\"P\") and n not in (\"<num>\", \"<sp>\")\n for n in names], dtype=bool)\nSP, NUM = name2id[\"<sp>\"], name2id[\"<num>\"]\n\nscores = np.zeros((ndoc, NREG), dtype=np.float32)\nnwin = np.zeros(ndoc, dtype=np.int32)\nfeat = np.zeros((ndoc, 4), dtype=np.float32) # uniq, word, punct+sp, num fractions\nsig = np.zeros((ndoc, 16), dtype=np.uint64) # near-duplicate min-hash signature\nPRIME = np.uint64(1099511628211)\n\nbuf_w, buf_o, nbuf = [], [], 0\n\n\ndef flush():\n global buf_w, buf_o, nbuf\n if not buf_w:\n return\n mat, own = np.concatenate(buf_w), np.concatenate(buf_o)\n with torch.no_grad():\n lg = torch.log_softmax(lin(bags(mat)), dim=1)\n lo = (lg[:, :NREG] - lg[:, NREG:NREG + 1]).cpu().numpy()\n for k in range(NREG):\n np.add.at(scores[:, k], own, lo[:, k])\n np.add.at(nwin, own, 1)\n buf_w, buf_o, nbuf = [], [], 0\n\nfor i in range(ndoc):\n c = canon[corpus[offs[i]:offs[i + 1]].astype(np.int64)]\n n = max(1, len(c))\n bc = np.bincount(c, minlength=V)\n feat[i] = (np.count_nonzero(bc) / n, bc[is_word].sum() / n,\n (bc[is_punct].sum() + bc[SP]) / n, bc[NUM] / n)\n # 5-gram rolling hash -> 16 smallest values = LSH signature for near-dup detection\n cu = c.astype(np.uint64)\n if len(cu) >= 5:\n h = np.zeros(len(cu) - 4, dtype=np.uint64)\n for j in range(5):\n h = h * PRIME + cu[j:len(cu) - 4 + j]\n h ^= h >> np.uint64(29)\n sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(\n np.sort(np.unique(h)), (0, 16 - len(np.unique(h))), constant_values=0)\n w = windows(corpus[offs[i]:offs[i + 1]])[:MAXW]\n if len(w):\n buf_w.append(w); buf_o.append(np.full(len(w), i)); nbuf += len(w)\n if nbuf >= 4096:\n flush()\nflush()\nscores /= np.maximum(1, nwin)[:, None]\n\nuniq_f, word_f, punct_f, num_f = feat.T\n# Register-agnostic quality gate. Thresholds are the properties every one of the four\n# target registers shares (running prose) and that scraped nav/JS/spam pages violate.\ngate = ((ntok_doc >= 256) & # at least one full training block of context\n (word_f >= 0.55) & # mostly words, not symbols/markup\n (punct_f <= 0.32) & # not menu/template/symbol soup\n (num_f <= 0.10) & # not price lists / tables / logs\n (uniq_f >= 0.20)) # not a boilerplate loop\nif args.no_gate:\n gate = ntok_doc >= 256\nprint(f\"gate keeps {gate.sum()} / {ndoc} docs \"\n f\"({ntok_doc[gate].sum()/1e6:.0f}M tokens)\", flush=True)\n\n# ------------------------------------------------------------------------ selection\nshares = np.array([float(x) for x in args.shares.split(\",\")]); shares /= shares.sum()\nbest_reg = scores.argmax(1)\ntarget_tok = shares * BUDGET * args.budget_mult\n\n# rank candidates within each register by that register's score\norder = {k: [i for i in np.argsort(-scores[:, k]) if gate[i] and best_reg[i] == k]\n for k in range(NREG)}\nprint(\"candidates/register\", {k: len(v) for k, v in order.items()}, flush=True)\n\nclaimed = {} # min-hash bucket -> owner\nselected, got = [], np.zeros(NREG, dtype=np.int64)\nptr = [0] * NREG\nndup = 0\n# Interleave registers in proportion to their share so the *prefix* of the list --\n# which is what the budget actually consumes -- is already register-balanced.\nwhile True:\n deficit = [(got[k] / target_tok[k], k) for k in range(NREG) if ptr[k] < len(order[k])]\n if not deficit:\n break\n _, k = min(deficit)\n i = order[k][ptr[k]]; ptr[k] += 1\n if not args.no_dedup:\n hits = [b for b in sig[i] if b and b in claimed]\n if len(hits) >= 4: # >=4/16 shared 5-grams -> near-dup\n ndup += 1\n continue\n for b in sig[i]:\n if b:\n claimed.setdefault(b, i)\n selected.append(int(i)); got[k] += ntok_doc[i]\n if got.sum() >= BUDGET * args.budget_mult:\n break\n\nprint(f\"selected {len(selected)} docs, {got.sum()/1e6:.1f}M tokens; \"\n f\"per-register Mtok {np.round(got/1e6,2).tolist()}; dropped {ndup} near-dups\", flush=True)\ncum = np.cumsum(ntok_doc[np.array(selected)])\ninfill = np.searchsorted(cum, BUDGET) + 1\nprint(f\"first {infill} docs fill the {BUDGET/1e6:.0f}M budget\", flush=True)\nos.makedirs(os.path.dirname(args.out), exist_ok=True)\njson.dump(selected, open(args.out, \"w\"))\nprint(\"wrote\", args.out, flush=True)\n"
}Write result
File created successfully at: /workspace/submission/curate.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/submission/curate.py", "content": "\"\"\"Curate /workspace/data/pool.jsonl -> submission/selection.json (priority ordered).\n\nSTATED CRITERION\n================\nThe training budget (12M tokens) is ~6% of the pool, and the evaluation target is a\n*known* mixture: four equal quarters of (0) encyclopedic WikiText-style prose,\n(1) general high-quality web prose, (2) news, (3) technical Q&A. Under a fixed token\nbudget, cross-entropy on such a target is minimised by training on data drawn from the\nsame mixture. So we (a) learn what each target register looks like, (b) reject\nmachine-generated web boilerplate that no register contains, and (c) fill the budget\nwith a register-balanced, deduplicated selection, best-scoring documents first.\n\nConcretely, a document is selected if it passes a register-agnostic quality gate, is\nnot a near-duplicate of an already-accepted document, and ranks highly for one of the\nfour registers under a linear register classifier.\n\n score_r(doc) = mean over the doc's 192-token windows of\n [ log P(register r | window) - log P(background | window) ]\n\nwhere P is a 5-way multinomial logistic regression over bags of *canonicalised* GPT-2\ntoken ids, fit on windows of the disclosed dev target (4 register classes) against\nwindows of a random pool sample (the \"background\" class).\n\nTwo details matter and are the difference between this working and not working:\n\n* CANONICALISATION. The encyclopedic quarter of the target is detokenised WikiText\n (\" 's\", \" , \", \" @-@ \"). Raw GPT-2 ids would let the classifier separate target from\n pool by *whitespace artefacts* rather than by register, and the resulting score would\n be meaningless on pool text. We therefore map every vocab entry through\n `decode(t).strip().lower()` -> letters-only, else <num> / <sp> / punctuation bucket.\n This collapses \" ,\" with \",\" and \" The\" with \"the\", leaving only content.\n\n* HTML STRIPPING OF THE TECHNICAL REGISTER. The technical quarter is StackExchange with\n markup (`<p>`, `<pre><code>`, `>`). Fit on raw markup, the classifier's top-ranked\n pool documents are scraped JavaScript/template boilerplate -- symbol soup that shares\n the markup but none of the language. We strip tags and entities before fitting, so\n register 3 selects genuine technical prose.\n\nUsage: python curate.py [--out selection.json]\nRequires the pool pre-tokenised by tokenize_pool.py + build_index.py (GPT-2 BPE);\nboth are pure preprocessing (no labels, no external data).\n\"\"\"\nimport argparse, html, json, os, re, sys\nimport numpy as np\nimport torch\n\nos.environ.setdefault(\"TOKENIZERS_PARALLELISM\", \"false\")\n\nW = \"/workspace/work\"\nDEV = \"/workspace/data/multi_dev.npy\"\nEOS = 50256\nNREG = 4\nWIN = 192 # classifier window length, in tokens\nMAXW = 8 # windows scored per document (cost bound; ~1.5k tokens)\nSEED = 0\nBUDGET = 12_000_000\n\nap = argparse.ArgumentParser()\nap.add_argument(\"--out\", default=\"/workspace/submission/selection.json\")\nap.add_argument(\"--budget_mult\", type=float, default=2.5,\n help=\"emit this many times the token budget, so ids beyond it are spare\")\nap.add_argument(\"--shares\", default=\"0.25,0.25,0.25,0.25\",\n help=\"target token share per register (wiki,webprose,news,techqa)\")\nap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nargs = ap.parse_args()\ndev_t = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n\n\n# ------------------------------------------------------------------ canonical vocab\ndef build_canon():\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n canon, slot, names = np.zeros(50257, dtype=np.int64), {}, []\n for t in range(50257):\n s = tok.decode([t]).strip().lower()\n letters = \"\".join(c for c in s if c.isalpha())\n key = letters if letters else (\"<num>\" if any(c.isdigit() for c in s)\n else (\"<sp>\" if s == \"\" else \"P\" + s))\n if key not in slot:\n slot[key] = len(names); names.append(key)\n canon[t] = slot[key]\n return canon, names\n\n\nif os.path.exists(f\"{W}/canon.npy\"):\n canon = np.load(f\"{W}/canon.npy\"); names = json.load(open(f\"{W}/canon_names.json\"))\nelse:\n canon, names = build_canon()\n np.save(f\"{W}/canon.npy\", canon); json.dump(names, open(f\"{W}/canon_names.json\", \"w\"))\nV = len(names)\nname2id = {n: i for i, n in enumerate(names)}\n\ncorpus = np.load(f\"{W}/corpus.npy\")\noffs = np.load(f\"{W}/offsets.npy\")\nndoc = len(offs) - 1\nntok_doc = np.diff(offs)\ncanon_t = torch.from_numpy(canon).to(dev_t)\nprint(f\"pool: {ndoc} docs, {ntok_doc.sum()} tokens\", flush=True)\n\n# format-only tokens carry no register information -> zero them out\ndrop_mask = torch.zeros(V, dtype=torch.bool, device=dev_t)\nfor k in (canon[EOS], canon[31]): # <|endoftext|>, \"@\" (WikiText filler)\n drop_mask[int(k)] = True\n\n\ndef windows(stream, win=WIN):\n n = (len(stream) // win) * win\n return (stream[:n].reshape(-1, win) if n else\n np.zeros((0, win), dtype=stream.dtype))\n\n\ndef bags(win_mat):\n x = torch.from_numpy(np.ascontiguousarray(win_mat).astype(np.int64)).to(dev_t)\n c = canon_t[x]\n out = torch.zeros(len(x), V, device=dev_t)\n out.scatter_add_(1, c, torch.ones_like(c, dtype=torch.float))\n out[:, drop_mask] = 0.0\n out = torch.log1p(out) # damp within-window repetition\n return out / out.norm(dim=1, keepdim=True).clamp_min(1e-6)\n\n\n# ------------------------------------------------- target registers (quarters of dev)\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).astype(np.int64)\nq = len(dev) // NREG\nreg_wins = []\nfor i in range(NREG):\n seg = dev[i * q:(i + 1) * q]\n txt = tok.decode(seg)\n if re.search(r\"</?(p|pre|code|blockquote)[ >]\", txt): # the technical quarter\n txt = html.unescape(re.sub(r\"<[^>\\n]{1,60}>\", \" \", txt))\n seg = np.array(tok(txt, add_special_tokens=False).input_ids, dtype=np.int64)\n print(\" stripped markup from register\", i, flush=True)\n reg_wins.append(windows(seg))\nprint(\"register windows\", [len(w) for w in reg_wins], flush=True)\n\nrng = np.random.default_rng(SEED)\nbg_docs = rng.choice(ndoc, size=6000, replace=False)\nbg = np.concatenate([windows(corpus[offs[i]:offs[i + 1]])[:3] for i in bg_docs])\n\nXw = [bags(w) for w in reg_wins] + [bags(bg)]\nX = torch.cat(Xw)\ny = torch.cat([torch.full((len(b),), k, device=dev_t, dtype=torch.long)\n for k, b in enumerate(Xw)])\ncnt = torch.bincount(y, minlength=NREG + 1).float()\ncw = cnt.sum() / (len(cnt) * cnt) # class-balanced\n\ntorch.manual_seed(SEED)\nlin = torch.nn.Linear(V, NREG + 1).to(dev_t)\ntorch.nn.init.zeros_(lin.weight); torch.nn.init.zeros_(lin.bias)\noptim = torch.optim.AdamW(lin.parameters(), lr=3e-3, weight_decay=1e-4)\nperm = torch.randperm(len(X), device=dev_t)\nntr = int(0.9 * len(X)); tr, va = perm[:ntr], perm[ntr:]\nfor ep in range(320):\n optim.zero_grad()\n torch.nn.functional.cross_entropy(lin(X[tr]), y[tr], weight=cw).backward()\n optim.step()\nwith torch.no_grad():\n pv = lin(X[va]).argmax(1)\n print(\"classifier val acc %.3f | per-class recall %s\" % (\n (pv == y[va]).float().mean().item(),\n \" \".join(f\"{k}:{(pv[y[va]==k]==k).float().mean().item():.2f}\" for k in range(NREG + 1))),\n flush=True)\n\n# ------------------------------------------------------------ score + gate the pool\nPUNCT = np.array([i for i, n in enumerate(names) if n.startswith(\"P\")], dtype=np.int64)\nis_punct = np.zeros(V, dtype=bool); is_punct[PUNCT] = True\nis_word = np.array([not n.startswith(\"P\") and n not in (\"<num>\", \"<sp>\")\n for n in names], dtype=bool)\nSP, NUM = name2id[\"<sp>\"], name2id[\"<num>\"]\n\nscores = np.zeros((ndoc, NREG), dtype=np.float32)\nnwin = np.zeros(ndoc, dtype=np.int32)\nfeat = np.zeros((ndoc, 4), dtype=np.float32) # uniq, word, punct+sp, num fractions\nsig = np.zeros((ndoc, 16), dtype=np.uint64) # near-duplicate min-hash signature\nPRIME = np.uint64(1099511628211)\n\nbuf_w, buf_o, nbuf = [], [], 0\n\n\ndef flush():\n global buf_w, buf_o, nbuf\n if not buf_w:\n return\n mat, own = np.concatenate(buf_w), np.concatenate(buf_o)\n with torch.no_grad():\n lg = torch.log_softmax(lin(bags(mat)), dim=1)\n lo = (lg[:, :NREG] - lg[:, NREG:NREG + 1]).cpu().numpy()\n for k in range(NREG):\n np.add.at(scores[:, k], own, lo[:, k])\n np.add.at(nwin, own, 1)\n buf_w, buf_o, nbuf = [], [], 0\n\nfor i in range(ndoc):\n c = canon[corpus[offs[i]:offs[i + 1]].astype(np.int64)]\n n = max(1, len(c))\n bc = np.bincount(c, minlength=V)\n feat[i] = (np.count_nonzero(bc) / n, bc[is_word].sum() / n,\n (bc[is_punct].sum() + bc[SP]) / n, bc[NUM] / n)\n # 5-gram rolling hash -> 16 smallest values = LSH signature for near-dup detection\n cu = c.astype(np.uint64)\n if len(cu) >= 5:\n h = np.zeros(len(cu) - 4, dtype=np.uint64)\n for j in range(5):\n h = h * PRIME + cu[j:len(cu) - 4 + j]\n h ^= h >> np.uint64(29)\n sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(\n np.sort(np.unique(h)), (0, 16 - len(np.unique(h))), constant_values=0)\n w = windows(corpus[offs[i]:offs[i + 1]])[:MAXW]\n if len(w):\n buf_w.append(w); buf_o.append(np.full(len(w), i)); nbuf += len(w)\n if nbuf >= 4096:\n flush()\nflush()\nscores /= np.maximum(1, nwin)[:, None]\n\nuniq_f, word_f, punct_f, num_f = feat.T\n# Register-agnostic quality gate. Thresholds are the properties every one of the four\n# target registers shares (running prose) and that scraped nav/JS/spam pages violate.\ngate = ((ntok_doc >= 256) & # at least one full training block of context\n (word_f >= 0.55) & # mostly words, not symbols/markup\n (punct_f <= 0.32) & # not menu/template/symbol soup\n (num_f <= 0.10) & # not price lists / tables / logs\n (uniq_f >= 0.20)) # not a boilerplate loop\nif args.no_gate:\n gate = ntok_doc >= 256\nprint(f\"gate keeps {gate.sum()} / {ndoc} docs \"\n f\"({ntok_doc[gate].sum()/1e6:.0f}M tokens)\", flush=True)\n\n# ------------------------------------------------------------------------ selection\nshares = np.array([float(x) for x in args.shares.split(\",\")]); shares /= shares.sum()\nbest_reg = scores.argmax(1)\ntarget_tok = shares * BUDGET * args.budget_mult\n\n# rank candidates within each register by that register's score\norder = {k: [i for i in np.argsort(-scores[:, k]) if gate[i] and best_reg[i] == k]\n for k in range(NREG)}\nprint(\"candidates/register\", {k: len(v) for k, v in order.items()}, flush=True)\n\nclaimed = {} # min-hash bucket -> owner\nselected, got = [], np.zeros(NREG, dtype=np.int64)\nptr = [0] * NREG\nndup = 0\n# Interleave registers in proportion to their share so the *prefix* of the list --\n# which is what the budget actually consumes -- is already register-balanced.\nwhile True:\n deficit = [(got[k] / target_tok[k], k) for k in range(NREG) if ptr[k] < len(order[k])]\n if not deficit:\n break\n _, k = min(deficit)\n i = order[k][ptr[k]]; ptr[k] += 1\n if not args.no_dedup:\n hits = [b for b in sig[i] if b and b in claimed]\n if len(hits) >= 4: # >=4/16 shared 5-grams -> near-dup\n ndup += 1\n continue\n for b in sig[i]:\n if b:\n claimed.setdefault(b, i)\n selected.append(int(i)); got[k] += ntok_doc[i]\n if got.sum() >= BUDGET * args.budget_mult:\n break\n\nprint(f\"selected {len(selected)} docs, {got.sum()/1e6:.1f}M tokens; \"\n f\"per-register Mtok {np.round(got/1e6,2).tolist()}; dropped {ndup} near-dups\", flush=True)\ncum = np.cumsum(ntok_doc[np.array(selected)])\ninfill = np.searchsorted(cum, BUDGET) + 1\nprint(f\"first {infill} docs fill the {BUDGET/1e6:.0f}M budget\", flush=True)\nos.makedirs(os.path.dirname(args.out), exist_ok=True)\njson.dump(selected, open(args.out, \"w\"))\nprint(\"wrote\", args.out, flush=True)\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "time python submission/curate.py --out /workspace/work/sel_balanced.json 2>&1 | tail -20; cat /workspace/.timer/remaining_secs",
"description": "Run curation, balanced variant",
"timeout": 900000
}Bash result
pool: 182016 docs, 200093788 tokens
Token indices sequence length is longer than the specified maximum sequence length for this model (214539 > 1024). Running this sequence through the model will result in indexing errors
stripped markup from register 3
register windows [1302, 1302, 1302, 1117]
classifier val acc 0.901 | per-class recall 0:0.97 1:0.81 2:0.90 3:1.00 4:0.89
Traceback (most recent call last):
File "/workspace/submission/curate.py", line 209, in <module>
sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(
ValueError: could not broadcast input array from shape (15,) into shape (16,)
real 0m24.726s
user 0m22.710s
sys 0m1.960s
5010
[stdout]
pool: 182016 docs, 200093788 tokens
Token indices sequence length is longer than the specified maximum sequence length for this model (214539 > 1024). Running this sequence through the model will result in indexing errors
stripped markup from register 3
register windows [1302, 1302, 1302, 1117]
classifier val acc 0.901 | per-class recall 0:0.97 1:0.81 2:0.90 3:1.00 4:0.89
Traceback (most recent call last):
File "/workspace/submission/curate.py", line 209, in <module>
sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(
ValueError: could not broadcast input array from shape (15,) into shape (16,)
real 0m24.726s
user 0m22.710s
sys 0m1.960s
5010
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": " sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(\n np.sort(np.unique(h)), (0, 16 - len(np.unique(h))), constant_values=0)",
"new_string": " u = np.sort(np.unique(h))[:16]\n sig[i] = np.pad(u, (0, 16 - len(u)), constant_values=0)"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": " sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(\n np.sort(np.unique(h)), (0, 16 - len(np.unique(h))), constant_values=0)", "newString": " u = np.sort(np.unique(h))[:16]\n sig[i] = np.pad(u, (0, 16 - len(u)), constant_values=0)", "originalFile": null, "structuredPatch": [{"oldStart": 206, "oldLines": 8, "newStart": 206, "newLines": 8, "lines": [" for j in range(5):", " h = h * PRIME + cu[j:len(cu) - 4 + j]", " h ^= h >> np.uint64(29)", "- sig[i] = np.sort(np.unique(h))[:16] if len(h) >= 16 else np.pad(", "- np.sort(np.unique(h)), (0, 16 - len(np.unique(h))), constant_values=0)", "+ u = np.sort(np.unique(h))[:16]", "+ sig[i] = np.pad(u, (0, 16 - len(u)), constant_values=0)", " w = windows(corpus[offs[i]:offs[i + 1]])[:MAXW]", " if len(w):", " buf_w.append(w); buf_o.append(np.full(len(w), i)); nbuf += len(w)"]}], "userModified": false, "replaceAll": false}Meanwhile, let me write the fast packer (reusing the pre-tokenized corpus) and verify it matches the official pack_selection.py byte-for-byte.
args
{
"file_path": "/workspace/work/fastpack.py",
"content": "\"\"\"Pack a selection.json into a 12M-token train array using the pre-tokenised corpus.\n\nEquivalent to the provided pack_selection.py (same GPT-2 BPE, same per-doc EOS,\nsame truncation) but ~100x faster because the pool is already tokenised.\n\"\"\"\nimport json, sys, numpy as np\n\nSEL = sys.argv[1]\nOUT = sys.argv[2] if len(sys.argv) > 2 else \"/tmp/dev_train.npy\"\nBUDGET = int(sys.argv[3]) if len(sys.argv) > 3 else 12_000_000\n\ncorpus = np.load(\"/workspace/work/corpus.npy\")\noffs = np.load(\"/workspace/work/offsets.npy\")\nsel = json.load(open(SEL))\nparts, tot = [], 0\nfor i in sel:\n parts.append(corpus[offs[i]:offs[i + 1]])\n tot += offs[i + 1] - offs[i]\n if tot >= BUDGET:\n break\narr = np.concatenate(parts)[:BUDGET].astype(np.uint16)\nnp.save(OUT, arr)\nprint(f\"packed {len(arr)} tokens from {len(parts)} of {len(sel)} ids -> {OUT}\")\n"
}Write result
File created successfully at: /workspace/work/fastpack.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/work/fastpack.py", "content": "\"\"\"Pack a selection.json into a 12M-token train array using the pre-tokenised corpus.\n\nEquivalent to the provided pack_selection.py (same GPT-2 BPE, same per-doc EOS,\nsame truncation) but ~100x faster because the pool is already tokenised.\n\"\"\"\nimport json, sys, numpy as np\n\nSEL = sys.argv[1]\nOUT = sys.argv[2] if len(sys.argv) > 2 else \"/tmp/dev_train.npy\"\nBUDGET = int(sys.argv[3]) if len(sys.argv) > 3 else 12_000_000\n\ncorpus = np.load(\"/workspace/work/corpus.npy\")\noffs = np.load(\"/workspace/work/offsets.npy\")\nsel = json.load(open(SEL))\nparts, tot = [], 0\nfor i in sel:\n parts.append(corpus[offs[i]:offs[i + 1]])\n tot += offs[i + 1] - offs[i]\n if tot >= BUDGET:\n break\narr = np.concatenate(parts)[:BUDGET].astype(np.uint16)\nnp.save(OUT, arr)\nprint(f\"packed {len(arr)} tokens from {len(parts)} of {len(sel)} ids -> {OUT}\")\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "python submission/curate.py --out /workspace/work/sel_balanced.json 2>&1 | grep -v \"Token indices\" | tail -12",
"description": "Rerun curation after bugfix",
"timeout": 900000
}Bash result
pool: 182016 docs, 200093788 tokens
stripped markup from register 3
register windows [1302, 1302, 1302, 1117]
classifier val acc 0.901 | per-class recall 0:0.97 1:0.81 2:0.90 3:1.00 4:0.89
gate keeps 116205 / 182016 docs (119M tokens)
candidates/register {0: 14140, 1: 75959, 2: 15999, 3: 10107}
selected 33877 docs, 30.0M tokens; per-register Mtok [8.27, 8.26, 8.26, 5.23]; dropped 7462 near-dups
first 14174 docs fill the 12M budget
wrote /workspace/work/sel_balanced.json
[stdout]
pool: 182016 docs, 200093788 tokens
stripped markup from register 3
register windows [1302, 1302, 1302, 1117]
classifier val acc 0.901 | per-class recall 0:0.97 1:0.81 2:0.90 3:1.00 4:0.89
gate keeps 116205 / 182016 docs (119M tokens)
candidates/register {0: 14140, 1: 75959, 2: 15999, 3: 10107}
selected 33877 docs, 30.0M tokens; per-register Mtok [8.27, 8.26, 8.26, 5.23]; dropped 7462 near-dups
first 14174 docs fill the 12M budget
wrote /workspace/work/sel_balanced.json
[metadata] {"noOutputExpected": false}Register 3 exhausted its candidates. Let me verify my fast packer matches the official one, check what TECH-QA now selects, and launch the first training run.
args
{
"command": "python -c \"\nimport json,numpy as np\nsel=json.load(open('work/sel_balanced.json'))\njson.dump(sel[:60], open('/tmp/tiny.json','w'))\n\" && python work/fastpack.py /tmp/tiny.json /tmp/fp.npy 300000 && TOKENIZERS_PARALLELISM=false python pack_selection.py /tmp/tiny.json /tmp/op.npy 2>&1 | tail -2 && python -c \"\nimport numpy as np\na=np.load('/tmp/fp.npy'); b=np.load('/tmp/op.npy')\nprint('fast',a.shape,'official',b.shape)\nn=min(len(a),len(b)); print('IDENTICAL prefix:', np.array_equal(a[:n],b[:n]), n)\n\"",
"description": "Verify fast packer matches official packer",
"timeout": 600000
}Bash result
packed 33121 tokens from 60 of 60 ids -> /tmp/fp.npy
Token indices sequence length is longer than the specified maximum sequence length for this model (5414 > 1024). Running this sequence through the model will result in indexing errors
packed 33121 tokens from 60 ids -> /tmp/op.npy
fast (33121,) official (33121,)
IDENTICAL prefix: True 33121
[stdout]
packed 33121 tokens from 60 of 60 ids -> /tmp/fp.npy
Token indices sequence length is longer than the specified maximum sequence length for this model (5414 > 1024). Running this sequence through the model will result in indexing errors
packed 33121 tokens from 60 ids -> /tmp/op.npy
fast (33121,) official (33121,)
IDENTICAL prefix: True 33121
[metadata] {"noOutputExpected": false}Byte-identical. Now let me evaluate the balanced selection against a random baseline.
args
{
"command": "python -c \"\nimport numpy as np, json\n# do-nothing reference: uniform random documents\no=np.load('work/offsets.npy'); nt=np.diff(o); n=len(nt)\nr=np.random.default_rng(1234).permutation(n)\ncum=np.cumsum(nt[r]); k=int(np.searchsorted(cum,30_000_000))+1\njson.dump([int(i) for i in r[:k]], open('work/sel_random.json','w'))\nprint('random baseline docs',k)\n\"\ncat > work/eval.sh <<'EOF'\n#!/bin/bash\nset -e\ncd /workspace\nfor name in \"$@\"; do\n python work/fastpack.py work/sel_${name}.json /tmp/tr_${name}.npy > /dev/null\n echo -n \"[$name] \"\n python train_nano.py --train_npy /tmp/tr_${name}.npy --heldout_npy data/multi_dev.npy \\\n --out_json /tmp/res_${name}.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL\ndone\nEOF\nchmod +x work/eval.sh\nnohup bash work/eval.sh balanced random > work/eval1.log 2>&1 &\necho launched",
"description": "Launch balanced vs random training evals"
}Bash result
random baseline docs 27678
launched
[stdout]
random baseline docs 27678
launched
[metadata] {"noOutputExpected": false}While that trains, let me inspect what the fixed TECH-QA register now selects and the register mix inside the budget prefix.
args
{
"command": "python -c \"\nimport numpy as np, json, os\nos.environ['TOKENIZERS_PARALLELISM']='false'\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained('gpt2')\nC=np.load('work/corpus.npy'); O=np.load('work/offsets.npy')\nsel=json.load(open('work/sel_balanced.json'))\nnt=np.diff(O)\ncum=np.cumsum(nt[np.array(sel)]); k=int(np.searchsorted(cum,12_000_000))+1\nprint('budget prefix docs',k,'tokens',cum[k-1])\nS=np.load('work/reg_scores.npy') # from score.py (pre-fix) - recompute best reg is fine for mix\n\" 2>&1|tail -3\necho \"=== rerun curate with score dump for inspection ===\"\npython - <<'EOF'\nimport numpy as np, json, os, subprocess\nos.environ['TOKENIZERS_PARALLELISM']='false'\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained('gpt2')\nC=np.load('/workspace/work/corpus.npy'); O=np.load('/workspace/work/offsets.npy')\nsel=json.load(open('/workspace/work/sel_balanced.json'))\n# the interleave order is round-robin by deficit; sample the tech-register picks by\n# finding docs whose position in list is late-ish and printing a spread\ndef show(i,n=200):\n return tok.decode(C[O[i]:O[i]+80].astype(np.int64)).replace(chr(10),' | ')[:n]\nprint('--- first 8 selected (highest priority) ---')\nfor i in sel[:8]: print(' ',i,'|',show(i))\nprint('--- selected at positions 12000-12004 ---')\nfor i in sel[12000:12004]: print(' ',i,'|',show(i))\nEOF",
"description": "Inspect selected documents",
"timeout": 600000
}Bash result
budget prefix docs 14174 tokens 12000949
=== rerun curate with score dump for inspection ===
--- first 8 selected (highest priority) ---
82269 | Hill 303 massacre | |Hill 303 massacre| | Bodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound | |Location||Hill 303, Waegwan, South Korea| | |Date||August
47889 | GAINESVILLE, Fla. (AP) — Tim Kaine is shrugging off any possibility that he could be embarrassed by the release of hacked emails. | WikiLeaks, which has been posting stolen emails from Hillary Clinton
103973 | pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crowds, he said, "We have sent funds under various national schemes to Uttar Pradesh
64048 | <|endoftext|>Sometimes it's not enough for Publising features to be enabled. | Deployment Manifest generated from Export-SPWeb uses string value "PublisingPages" (15.0.4797.1001) for BaseTemplate of P
86830 | Greg Stein <firstname.lastname@example.org> writes: | > During the walk for the commit, I need the PATH for the "current" node so | > that I can pass it to svn_wc_prop_*(). But the best that I can com
37499 | Uttar Pradesh Special Task Force (STF) on Saturday arrested a wanted animal smuggler, carrying a reward of Rs 50,000, who is allegedly involved in the murder case of three priests. | The accused, iden
66305 | Bradford West Respect MP George Galloway has defended his controversial claim that a sex assault allegation against WikiLeaks campaigner Julian Assange amounted to no more than bad "sexual etiquette".
22765 | I think that the main page is referring to modules developed with any language, but are executed as dll's, lib's, vi's, exe's, and the like. Was there a specific scripting language that you were looki
--- selected at positions 12000-12004 ---
57148 | <|endoftext|>Alexander the Paphlagonian | “At about the same time with Apuleius (note: the Numidian writer in Latin, circa 124 – 170 AD) lived Alexander the Paphlagonian, of whom so extraordinary an a
169515 | ity | Vehicles Spares and Maintenance<|endoftext|>Can we build a gravity transformer? | Page 2 | Naked Science Forum | The Naked Scientists | Toggle navigation | Login | Register | Podcasts | The Nake
113597 | common interview questions are more telling than you'd think.Sebastiaan ter Burg/FlickrSome of the trickiest job interview questions are surprisingly common and straightforward. | But while they may
47504 | However, experts believe it is a company-specific issue (merger news), which has nothing to do with the improving prospects for the cement makers. | Analysts are advising Grasim investors to flock to
[stdout]
budget prefix docs 14174 tokens 12000949
=== rerun curate with score dump for inspection ===
--- first 8 selected (highest priority) ---
82269 | Hill 303 massacre | |Hill 303 massacre| | Bodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound | |Location||Hill 303, Waegwan, South Korea| | |Date||August
47889 | GAINESVILLE, Fla. (AP) — Tim Kaine is shrugging off any possibility that he could be embarrassed by the release of hacked emails. | WikiLeaks, which has been posting stolen emails from Hillary Clinton
103973 | pur: Prime Minister Manmohan Singh on Friday kicked off the Congress campaign from Kanpur on Friday. Addressing the crowds, he said, "We have sent funds under various national schemes to Uttar Pradesh
64048 | <|endoftext|>Sometimes it's not enough for Publising features to be enabled. | Deployment Manifest generated from Export-SPWeb uses string value "PublisingPages" (15.0.4797.1001) for BaseTemplate of P
86830 | Greg Stein <firstname.lastname@example.org> writes: | > During the walk for the commit, I need the PATH for the "current" node so | > that I can pass it to svn_wc_prop_*(). But the best that I can com
37499 | Uttar Pradesh Special Task Force (STF) on Saturday arrested a wanted animal smuggler, carrying a reward of Rs 50,000, who is allegedly involved in the murder case of three priests. | The accused, iden
66305 | Bradford West Respect MP George Galloway has defended his controversial claim that a sex assault allegation against WikiLeaks campaigner Julian Assange amounted to no more than bad "sexual etiquette".
22765 | I think that the main page is referring to modules developed with any language, but are executed as dll's, lib's, vi's, exe's, and the like. Was there a specific scripting language that you were looki
--- selected at positions 12000-12004 ---
57148 | <|endoftext|>Alexander the Paphlagonian | “At about the same time with Apuleius (note: the Numidian writer in Latin, circa 124 – 170 AD) lived Alexander the Paphlagonian, of whom so extraordinary an a
169515 | ity | Vehicles Spares and Maintenance<|endoftext|>Can we build a gravity transformer? | Page 2 | Naked Science Forum | The Naked Scientists | Toggle navigation | Login | Register | Podcasts | The Nake
113597 | common interview questions are more telling than you'd think.Sebastiaan ter Burg/FlickrSome of the trickiest job interview questions are surprisingly common and straightforward. | But while they may
47504 | However, experts believe it is a company-specific issue (merger news), which has nothing to do with the improving prospects for the cement makers. | Analysts are advising Grasim investors to flock to
[metadata] {"noOutputExpected": false}TECH-QA now picks genuine technical Q&A (SVN dev list, SharePoint, programming discussion) — the markup fix worked. Checking training results:
args
{
"command": "sleep 60; cat work/eval1.log | grep -E \"HELDOUT|\\[\" ; echo \"---timer\"; cat /workspace/.timer/remaining_secs",
"description": "Check eval results",
"timeout": 300000
}Bash result
[balanced] HELDOUT_PPL 331.028 loss 5.8022 train_s 56 /tmp/tr_balanced.npy
[random]
---timer
4829
[stdout]
[balanced] HELDOUT_PPL 331.028 loss 5.8022 train_s 56 /tmp/tr_balanced.npy
[random]
---timer
4829
[metadata] {"noOutputExpected": false}Balanced = 331.0. While the baseline finishes, let me test a high-value hypothesis: 25% of the target is WikiText-detokenized (space-before-punctuation). Do any pool docs share that surface form?
args
{
"command": "python - <<'EOF'\nimport numpy as np, os, json\nos.environ['TOKENIZERS_PARALLELISM']='false'\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained('gpt2')\nimport string\nPUNCT=set(string.punctuation)\nsp=np.zeros(50257,bool); bare=np.zeros(50257,bool)\nfor t in range(50257):\n s=tok.decode([t])\n core=s.strip()\n if core and all(c in PUNCT for c in core):\n if s.startswith(' '): sp[t]=True\n else: bare[t]=True\nprint('spaced-punct vocab entries',sp.sum(),'bare',bare.sum())\nC=np.load('/workspace/work/corpus.npy'); O=np.load('/workspace/work/offsets.npy')\nD=np.load('/workspace/data/multi_dev.npy').astype(np.int64)\ndef ratio(x):\n a=sp[x].sum(); b=bare[x].sum()\n return a/max(1,a+b), (a+b)/max(1,len(x))\nq=len(D)//4\nfor i in range(4):\n r,d=ratio(D[i*q:(i+1)*q]); print(f'dev quarter {i}: spaced-punct ratio {r:.3f} punct density {d:.3f}')\n# pool per doc\nn=len(O)-1\ncs_sp=np.cumsum(sp[C.astype(np.int64)]); cs_b=np.cumsum(bare[C.astype(np.int64)])\ndef seg(cs,i): return cs[O[i+1]-1]-(cs[O[i]-1] if O[i]>0 else 0)\nA=np.zeros(n); B=np.zeros(n)\npad_sp=np.concatenate([[0],cs_sp]); pad_b=np.concatenate([[0],cs_b])\nA=pad_sp[O[1:]]-pad_sp[O[:-1]]; B=pad_b[O[1:]]-pad_b[O[:-1]]\nrat=A/np.maximum(1,A+B)\nnp.save('/workspace/work/spratio.npy',rat)\nprint('pool spaced-punct ratio: mean %.3f p50 %.3f p90 %.3f p99 %.3f p99.9 %.3f max %.3f'%(\n rat.mean(),*[np.percentile(rat,p) for p in (50,90,99,99.9)],rat.max()))\nprint('docs with ratio>0.5:',(rat>0.5).sum(),' >0.3:',(rat>0.3).sum(),' >0.2:',(rat>0.2).sum())\nntok=np.diff(O)\nprint('tokens in docs ratio>0.3: %.1fM'%(ntok[rat>0.3].sum()/1e6))\no=np.argsort(-rat)[:3]\nfor i in o: print(' ',i,round(rat[i],3),'|',tok.decode(C[O[i]:O[i]+70].astype(np.int64)).replace(chr(10),' | ')[:190])\nEOF",
"description": "Measure WikiText spacing style in pool vs dev",
"timeout": 600000
}Bash result
spaced-punct vocab entries 231 bare 526
dev quarter 0: spaced-punct ratio 0.849 punct density 0.144
dev quarter 1: spaced-punct ratio 0.104 punct density 0.106
dev quarter 2: spaced-punct ratio 0.065 punct density 0.111
dev quarter 3: spaced-punct ratio 0.139 punct density 0.266
pool spaced-punct ratio: mean 0.125 p50 0.088 p90 0.296 p99 0.603 p99.9 1.000 max 1.000
docs with ratio>0.5: 3371 >0.3: 17444 >0.2: 36242
tokens in docs ratio>0.3: 27.7M
181618 1.0 | Plan<|endoftext|>Enhance Your Education | Home | Event Photos | Reserve Your Event Space | © 2017 Moraee International - Salon Moraee Education : Learn : Advance Training Careers<|endoftext
175903 1.0 | where Beanbag Slipcovers Childrens Fresh Rn My First Chai – Schwienbacher<|endoftext|>FUMBBL :: Online Blood Bowl<|endoftext|>UMBBL | Game | Community | Forums | Help | 26 coaches online • S
156512 1.0 | Mining - Mining & FIFO Jobs - Webjobz | Menu | Skip to content | Search for Jobs | Search Resumes | FAQ and Help | Job Seeker Panel | Employer Panel | Post a jobFind a job | Job<|endoftext|
[stdout]
spaced-punct vocab entries 231 bare 526
dev quarter 0: spaced-punct ratio 0.849 punct density 0.144
dev quarter 1: spaced-punct ratio 0.104 punct density 0.106
dev quarter 2: spaced-punct ratio 0.065 punct density 0.111
dev quarter 3: spaced-punct ratio 0.139 punct density 0.266
pool spaced-punct ratio: mean 0.125 p50 0.088 p90 0.296 p99 0.603 p99.9 1.000 max 1.000
docs with ratio>0.5: 3371 >0.3: 17444 >0.2: 36242
tokens in docs ratio>0.3: 27.7M
181618 1.0 | Plan<|endoftext|>Enhance Your Education | Home | Event Photos | Reserve Your Event Space | © 2017 Moraee International - Salon Moraee Education : Learn : Advance Training Careers<|endoftext
175903 1.0 | where Beanbag Slipcovers Childrens Fresh Rn My First Chai – Schwienbacher<|endoftext|>FUMBBL :: Online Blood Bowl<|endoftext|>UMBBL | Game | Community | Forums | Help | 26 coaches online • S
156512 1.0 | Mining - Mining & FIFO Jobs - Webjobz | Menu | Skip to content | Search for Jobs | Search Resumes | FAQ and Help | Job Seeker Panel | Employer Panel | Post a jobFind a job | Job<|endoftext|
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cat work/eval1.log | grep -E \"HELDOUT\"; echo \"--- timer\"; cat /workspace/.timer/remaining_secs",
"description": "Check random baseline result"
}Bash result
[balanced] HELDOUT_PPL 331.028 loss 5.8022 train_s 56 /tmp/tr_balanced.npy
[random] HELDOUT_PPL 477.843 loss 6.1693 train_s 56 /tmp/tr_random.npy
--- timer
4769
[stdout]
[balanced] HELDOUT_PPL 331.028 loss 5.8022 train_s 56 /tmp/tr_balanced.npy
[random] HELDOUT_PPL 477.843 loss 6.1693 train_s 56 /tmp/tr_random.npy
--- timer
4769
[metadata] {"noOutputExpected": false}Random baseline 477.8 vs balanced 331.0 — 31% better already. Now let me add principled surface-style matching (the wiki quarter's spacing is an observable I can match) and cache the scoring so variants are cheap.
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": "ap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nargs = ap.parse_args()",
"new_string": "ap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nap.add_argument(\"--style\", type=float, default=1.0,\n help=\"weight on matching each register's spaced-punctuation ratio\")\nap.add_argument(\"--cache\", default=f\"{W}/cache\")\nargs = ap.parse_args()"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": "ap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nargs = ap.parse_args()", "newString": "ap.add_argument(\"--no_gate\", action=\"store_true\")\nap.add_argument(\"--no_dedup\", action=\"store_true\")\nap.add_argument(\"--style\", type=float, default=1.0,\n help=\"weight on matching each register's spaced-punctuation ratio\")\nap.add_argument(\"--cache\", default=f\"{W}/cache\")\nargs = ap.parse_args()", "originalFile": null, "structuredPatch": [{"oldStart": 63, "oldLines": 6, "newStart": 63, "newLines": 9, "lines": [" help=\"target token share per register (wiki,webprose,news,techqa)\")", " ap.add_argument(\"--no_gate\", action=\"store_true\")", " ap.add_argument(\"--no_dedup\", action=\"store_true\")", "+ap.add_argument(\"--style\", type=float, default=1.0,", "+ help=\"weight on matching each register's spaced-punctuation ratio\")", "+ap.add_argument(\"--cache\", default=f\"{W}/cache\")", " args = ap.parse_args()", " dev_t = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")", " "]}], "userModified": false, "replaceAll": false}