claude-opus-4-8
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-4-8
- metric
- held-out multi-domain perplexity
- actions
- 84
- duration
- 1h 6m
- recorded
- 2026-08-04
run iddclm-nano-curation__GKXSUa7
Everything claude-opus-4-8 did on this task, in order.
download JSON ↗I'll start by understanding my environment: the time budget, the data, and the training pipeline.
args
{
"command": "cat /workspace/.timer/remaining_secs 2>/dev/null; echo \"---\"; ls -la /workspace/; echo \"---\"; ls -la /workspace/data/; echo \"---\"; ls -la /workspace/submission/ 2>/dev/null || echo \"no submission dir\"",
"description": "Check timer, list workspace and data"
}Bash result
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 4 04:45 .
drwxr-xr-x 1 root root 160 Aug 4 04:45 ..
drwxr-xr-x 1 root root 60 Aug 4 04:45 .live
drwxr-xr-x 1 root root 100 Aug 4 04:45 .metrics
drwxr-xr-x 1 root root 140 Aug 4 04:45 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 04:45 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 04:45 ..
[stdout]
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 4 04:45 .
drwxr-xr-x 1 root root 160 Aug 4 04:45 ..
drwxr-xr-x 1 root root 60 Aug 4 04:45 .live
drwxr-xr-x 1 root root 100 Aug 4 04:45 .metrics
drwxr-xr-x 1 root root 140 Aug 4 04:45 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 04:45 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 04:45 ..
[metadata] {"noOutputExpected": false}args
{
"command": "nvidia-smi 2>/dev/null | head -20; echo \"---GPU---\"; python3 -c \"import torch; print('torch', torch.__version__, 'cuda', torch.cuda.is_available())\" 2>/dev/null; echo \"---\"; head -c 2000 /workspace/data/pool.jsonl",
"description": "Check GPU and peek at pool data"
}Bash result
Tue Aug 4 04:45:21 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 80GB HBM3 On | 00000000:45:00.0 Off | 0 |
| N/A 30C P0 72W / 700W | 5MiB / 81559MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
---GPU---
torch 2.8.0+cu128 cuda True
---
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
[stdout]
Tue Aug 4 04:45:21 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05 Driver Version: 580.95.05 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA H100 80GB HBM3 On | 00000000:45:00.0 Off | 0 |
| N/A 30C P0 72W / 700W | 5MiB / 81559MiB | 0% Default |
| | | Disabled |
+-----------------------------------------+------------------------+----------------------+
+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| No running processes found |
+-----------------------------------------------------------------------------------------+
---GPU---
torch 2.8.0+cu128 cuda True
---
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
[metadata] {"noOutputExpected": false}~90 minutes of budget (5390s). Let me read the training pipeline and helper scripts to understand exactly how selection is consumed and evaluated.
args
{
"command": "cd /workspace; echo \"===== train_nano.py =====\"; cat train_nano.py; echo; echo \"===== model.py =====\"; cat model.py",
"description": "Read training script and model"
}Bash result
===== train_nano.py =====
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
===== model.py =====
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
[stdout]
===== train_nano.py =====
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
===== model.py =====
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; echo \"===== pack_selection.py =====\"; cat pack_selection.py; echo; echo \"===== dev_eval.sh =====\"; cat dev_eval.sh; echo; echo \"===== timer.sh =====\"; cat timer.sh",
"description": "Read helper scripts"
}Bash result
===== pack_selection.py =====
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
===== dev_eval.sh =====
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
===== timer.sh =====
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[stdout]
===== pack_selection.py =====
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
===== dev_eval.sh =====
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
===== timer.sh =====
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[metadata] {"noOutputExpected": false}Now I understand the pipeline. Key facts:
- Time: ~90 min total. Each
dev_eval= tokenize 12M tokens + train 3000 iters (~60s) ≈ a few min. I can afford maybe 4-6 training runs. - Training: 30M GPT, 3000 iters × batch 32 × block 256 ≈ 24.6M tokens seen over a 12M-token set (~2 epochs). Reports held-out PPL on
multi_dev.npy. - Target: broad HQ English (Wikipedia + HQ web + news + technical Q&A). Crucially, `multi_dev.npy` is tokenized target-domain text — I can decode it to get real target text to drive domain-matching selection.
Let me explore the pool and decode the dev target in parallel.
args
{
"command": "cd /workspace; python3 -c \"import tiktoken; print('tiktoken OK')\" 2>&1 | head -1; python3 -c \"import sklearn; print('sklearn', sklearn.__version__)\" 2>&1 | head -1; echo \"---count docs---\"; wc -l /workspace/data/pool.jsonl; echo \"---dev npy---\"; python3 -c \"\nimport numpy as np\na=np.load('/workspace/data/multi_dev.npy')\nprint('dev tokens', a.shape, a.dtype, 'min',a.min(),'max',a.max())\nimport numpy as np\n# count EOS (50256) occurrences => number of doc boundaries\nprint('EOS(50256) count', int((a==50256).sum()))\n\"",
"description": "Check libs, count docs, inspect dev npy"
}Bash result
Traceback (most recent call last):
Traceback (most recent call last):
---count docs---
182016 /workspace/data/pool.jsonl
---dev npy---
dev tokens (1000000,) uint16 min 0 max 50256
EOS(50256) count 2704
[stdout]
Traceback (most recent call last):
Traceback (most recent call last):
---count docs---
182016 /workspace/data/pool.jsonl
---dev npy---
dev tokens (1000000,) uint16 min 0 max 50256
EOS(50256) count 2704
[metadata] {"noOutputExpected": false}tiktoken and sklearn are absent. Pool = 182,016 docs; dev = 1M tokens / ~2704 docs (avg ~370 tok/doc). Let me check available libraries and decode the dev target to see the domains concretely.
args
{
"command": "cd /workspace; python3 -c \"\nimport pkgutil, importlib\nfor m in ['tokenizers','transformers','numpy','torch','scipy','sklearn','pandas','regex','datasets','fasttext']:\n try:\n mod=importlib.import_module(m); print(m, getattr(mod,'__version__','?'))\n except Exception as e:\n print(m, 'MISSING')\n\"",
"description": "Check available python libraries"
}Bash result
tokenizers 0.22.2
transformers 4.57.1
numpy 1.26.4
torch 2.8.0+cu128
scipy MISSING
sklearn MISSING
pandas MISSING
regex 2026.7.19
datasets MISSING
fasttext MISSING
[stdout]
tokenizers 0.22.2
transformers 4.57.1
numpy 1.26.4
torch 2.8.0+cu128
scipy MISSING
sklearn MISSING
pandas MISSING
regex 2026.7.19
datasets MISSING
fasttext MISSING
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np, time\ntok=AutoTokenizer.from_pretrained('gpt2')\nprint('is_fast', tok.is_fast)\na=np.load('/workspace/data/multi_dev.npy').astype(np.int64)\n# split on EOS 50256 into docs\ndocs=[]; cur=[]\nfor t in a:\n if t==50256:\n if cur: docs.append(cur); cur=[]\n else: cur.append(t)\nif cur: docs.append(cur)\nprint('num dev docs', len(docs))\nimport statistics\nlens=[len(d) for d in docs]\nprint('len tokens: min',min(lens),'median',int(statistics.median(lens)),'max',max(lens),'mean',int(sum(lens)/len(lens)))\n# decode first 6 docs to see domains\nfor i in range(6):\n txt=tok.decode(docs[i])\n print('==== DEV DOC',i,'len',len(docs[i]),'====')\n print(txt[:600].replace(chr(10),' ⏎ '))\n print()\n\"",
"description": "Decode dev target docs to inspect domains"
}Bash result
is_fast True
num dev docs 2705
len tokens: min 20 median 192 max 23346 mean 368
==== DEV DOC 0 len 206 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
==== DEV DOC 1 len 125 ====
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
==== DEV DOC 2 len 223 ====
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
==== DEV DOC 3 len 95 ====
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused . ⏎
==== DEV DOC 4 len 201 ====
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
==== DEV DOC 5 len 131 ====
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
[stdout]
is_fast True
num dev docs 2705
len tokens: min 20 median 192 max 23346 mean 368
==== DEV DOC 0 len 206 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
==== DEV DOC 1 len 125 ====
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
==== DEV DOC 2 len 223 ====
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
==== DEV DOC 3 len 95 ====
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused . ⏎
==== DEV DOC 4 len 201 ====
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
==== DEV DOC 5 len 131 ====
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
[metadata] {"noOutputExpected": false}The dev docs are WikiText-style formatted (note @,@ , @-@ , spaced punctuation). The first docs are all Wikipedia (Shackleton). Since the target is 4-register, let me sample across the whole dev set to see all domains, and benchmark tokenizer throughput (I'll need token counts for the whole pool).
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np\ntok=AutoTokenizer.from_pretrained('gpt2')\na=np.load('/workspace/data/multi_dev.npy').astype(np.int64)\ndocs=[]; cur=[]\nfor t in a:\n if t==50256:\n if cur: docs.append(cur); cur=[]\n else: cur.append(t)\nif cur: docs.append(cur)\nN=len(docs)\nfor idx in [700, 1300, 1400, 1500, 2000, 2600, 2700]:\n if idx<N:\n txt=tok.decode(docs[idx])\n print('==== DEV DOC',idx,'len',len(docs[idx]),'====')\n print(txt[:500].replace(chr(10),' ⏎ '))\n print()\n\"",
"description": "Sample dev docs across positions to see all registers"
}Bash result
==== DEV DOC 700 len 274 ====
In 1968 , Lennon was told The Dairy Cottage was too cramped for them all , so he told Birch to buy a house , and he found a 4 @-@ bedroom house in Gateacre Park Drive , Liverpool . Lennon told Birch to furnish and decorate it , and to send all the bills to him . The Dykinses heard nothing from Lennon for years , until he phoned Baird in 1975 , and asked for mementos of his childhood life , such as his school tie and photographs . He sent £ 3 @,@ 000 to cover the cost of shipping and as a gift ,
==== DEV DOC 1300 len 176 ====
In 1925 , US 112 was originally proposed to run from Oshkosh to Fremont , Wisconsin , on what later became U.S. Route 110 . When it was initially designated in November 1926 , US 112 made a sharp turn to the southwest to connect to US 20 in Elkhart , Indiana . In 1931 , a new trunkline highway was designated between M @-@ 60 at Niles and US 112 at Union . This highway was numbered M @-@ 151 . In 1933 , the section of US 112 from Union to Elkhart was renumbered US 112S . M @-@ 151 and US 112S ea
==== DEV DOC 1400 len 78 ====
Grissom is often regarded as well @-@ educated , but unusual in his approach toward his work and social life . In the series , some of his comments and actions can be seen to dumbfound his co @-@ workers and superiors . His relationship with his subordinates in the office is portrayed as being a father figure to the team , but very professional in his work . ⏎
==== DEV DOC 1500 len 111 ====
In his 2002 State of the Union Address , Bush referred to an axis of evil including Iraq , Iran and North Korea . After the September 11 attacks on New York , Bush launched the War on Terror , in which the United States military and a small international coalition invaded Afghanistan . In 2003 , Bush then launched the invasion of Iraq , searching for Weapons of Mass Destruction , which he described as being part of the War on Terrorism . Those invasions led to the toppling of the Taliban regime
==== DEV DOC 2000 len 151 ====
Socialite Paris Hilton looked happy during her holiday in Formentera with her boyfriend Chris Zylka as they enjoyed a boating session together.Hilton was spotted in a lacy red dress as she held hands with the actor as they crossed the beach and took a ride in a small boat on Friday, reports dailymail.co.uk.Hilton accessorised her ensemble with a massive red-rimmed pair of sunglasses, as well as with a black bag.She was seen strolling around while resting her hand on Zylka's shoulders.Credit: @ P
==== DEV DOC 2600 len 589 ====
<p>I was wondering if there is a more elegant way to do IN() queries with Spring's JDBCTemplate. Currently I do something like that:</p> ⏎ ⏎ <pre><code>StringBuilder jobTypeInClauseBuilder = new StringBuilder(); ⏎ for(int i = 0; i < jobTypes.length; i++) { ⏎ Type jobType = jobTypes[i]; ⏎ ⏎ if(i != 0) { ⏎ jobTypeInClauseBuilder.append(','); ⏎ } ⏎ ⏎ jobTypeInClauseBuilder.append(jobType.convert()); ⏎ } ⏎ </code></pre> ⏎ ⏎ <p>Which is quite painful since if I have nine lines just for building th
==== DEV DOC 2700 len 638 ====
<p>I'm working on a project which makes heavy use of extension methods to convert strings from the UI-layer into their appropriate object-types in the code-layer: <em>(pseudo-code)</em></p> ⏎ ⏎ <pre><code>// C# ⏎ ⏎ /* ⏎ Converts a String from the UI-layer, formatted according to a user-defined ⏎ UI-culture preference (in this case from da-DK) into a Double ⏎ */ ⏎ ⏎ Double d = "1.000,50".fromWebStringToDouble(); // 1000.5 ⏎ </code></pre> ⏎ ⏎ <p><br /> ⏎ To unify the programming experience across C# and JavaScript, I w
[stdout]
==== DEV DOC 700 len 274 ====
In 1968 , Lennon was told The Dairy Cottage was too cramped for them all , so he told Birch to buy a house , and he found a 4 @-@ bedroom house in Gateacre Park Drive , Liverpool . Lennon told Birch to furnish and decorate it , and to send all the bills to him . The Dykinses heard nothing from Lennon for years , until he phoned Baird in 1975 , and asked for mementos of his childhood life , such as his school tie and photographs . He sent £ 3 @,@ 000 to cover the cost of shipping and as a gift ,
==== DEV DOC 1300 len 176 ====
In 1925 , US 112 was originally proposed to run from Oshkosh to Fremont , Wisconsin , on what later became U.S. Route 110 . When it was initially designated in November 1926 , US 112 made a sharp turn to the southwest to connect to US 20 in Elkhart , Indiana . In 1931 , a new trunkline highway was designated between M @-@ 60 at Niles and US 112 at Union . This highway was numbered M @-@ 151 . In 1933 , the section of US 112 from Union to Elkhart was renumbered US 112S . M @-@ 151 and US 112S ea
==== DEV DOC 1400 len 78 ====
Grissom is often regarded as well @-@ educated , but unusual in his approach toward his work and social life . In the series , some of his comments and actions can be seen to dumbfound his co @-@ workers and superiors . His relationship with his subordinates in the office is portrayed as being a father figure to the team , but very professional in his work . ⏎
==== DEV DOC 1500 len 111 ====
In his 2002 State of the Union Address , Bush referred to an axis of evil including Iraq , Iran and North Korea . After the September 11 attacks on New York , Bush launched the War on Terror , in which the United States military and a small international coalition invaded Afghanistan . In 2003 , Bush then launched the invasion of Iraq , searching for Weapons of Mass Destruction , which he described as being part of the War on Terrorism . Those invasions led to the toppling of the Taliban regime
==== DEV DOC 2000 len 151 ====
Socialite Paris Hilton looked happy during her holiday in Formentera with her boyfriend Chris Zylka as they enjoyed a boating session together.Hilton was spotted in a lacy red dress as she held hands with the actor as they crossed the beach and took a ride in a small boat on Friday, reports dailymail.co.uk.Hilton accessorised her ensemble with a massive red-rimmed pair of sunglasses, as well as with a black bag.She was seen strolling around while resting her hand on Zylka's shoulders.Credit: @ P
==== DEV DOC 2600 len 589 ====
<p>I was wondering if there is a more elegant way to do IN() queries with Spring's JDBCTemplate. Currently I do something like that:</p> ⏎ ⏎ <pre><code>StringBuilder jobTypeInClauseBuilder = new StringBuilder(); ⏎ for(int i = 0; i < jobTypes.length; i++) { ⏎ Type jobType = jobTypes[i]; ⏎ ⏎ if(i != 0) { ⏎ jobTypeInClauseBuilder.append(','); ⏎ } ⏎ ⏎ jobTypeInClauseBuilder.append(jobType.convert()); ⏎ } ⏎ </code></pre> ⏎ ⏎ <p>Which is quite painful since if I have nine lines just for building th
==== DEV DOC 2700 len 638 ====
<p>I'm working on a project which makes heavy use of extension methods to convert strings from the UI-layer into their appropriate object-types in the code-layer: <em>(pseudo-code)</em></p> ⏎ ⏎ <pre><code>// C# ⏎ ⏎ /* ⏎ Converts a String from the UI-layer, formatted according to a user-defined ⏎ UI-culture preference (in this case from da-DK) into a Double ⏎ */ ⏎ ⏎ Double d = "1.000,50".fromWebStringToDouble(); // 1000.5 ⏎ </code></pre> ⏎ ⏎ <p><br /> ⏎ To unify the programming experience across C# and JavaScript, I w
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=true; python3 -c \"\nfrom transformers import AutoTokenizer\nimport json, time\ntok=AutoTokenizer.from_pretrained('gpt2')\ntexts=[]\nwith open('/workspace/data/pool.jsonl') as f:\n for i,line in enumerate(f):\n if i>=3000: break\n texts.append(json.loads(line)['text'])\nt0=time.time()\nenc=tok(texts, add_special_tokens=False)\ndt=time.time()-t0\nntok=sum(len(x) for x in enc.input_ids)\nprint(f'{len(texts)} docs, {ntok} tokens in {dt:.2f}s -> {ntok/dt/1e6:.2f}M tok/s, {len(texts)/dt:.0f} docs/s')\nprint(f'projected full pool 182016 docs: {182016/(len(texts)/dt):.0f}s')\nimport statistics\nlens=[len(x) for x in enc.input_ids]\nprint('pool sample token lens: min',min(lens),'median',int(statistics.median(lens)),'mean',int(sum(lens)/len(lens)),'max',max(lens))\n\"",
"description": "Benchmark tokenizer throughput on 3000 pool docs"
}Bash result
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
3000 docs, 2086327 tokens in 0.87s -> 2.38M tok/s, 3429 docs/s
projected full pool 182016 docs: 53s
pool sample token lens: min 46 median 386 mean 695 max 30220
[stdout]
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
3000 docs, 2086327 tokens in 0.87s -> 2.38M tok/s, 3429 docs/s
projected full pool 182016 docs: 53s
pool sample token lens: min 46 median 386 mean 695 max 30220
[metadata] {"noOutputExpected": false}The picture is clear now. The dev target is domain-blocked: Wikipedia/encyclopedic (WikiText-style, ~docs 0–1500+), then news (Paris Hilton/dailymail, ~2000), then technical Q&A (StackOverflow HTML <p>/<pre><code>, ~2600+). Tokenizer is fast (full pool ≈ 53s). Let me set up task tracking and check RAM.
args
{
"command": "free -g 2>/dev/null | head -3; echo \"---\"; grep MemTotal /proc/meminfo; echo \"---timer---\"; cat /workspace/.timer/remaining_secs; echo \"---nproc---\"; nproc",
"description": "Check RAM, timer, CPU count"
}Bash result
total used free shared buff/cache available
Mem: 1024 0 1023 0 0 1023
Swap: 0 0 0
---
MemTotal: 1073741824 kB
---timer---
5169
---nproc---
16
[stdout]
total used free shared buff/cache available
Mem: 1024 0 1023 0 0 1023
Swap: 0 0 0
---
MemTotal: 1073741824 kB
---timer---
5169
---nproc---
16
[metadata] {"noOutputExpected": false}1 TB RAM, 16 CPUs, 5169s left. Let me create tasks, then in one pass: characterize the dev registers precisely and tokenize+cache the whole pool (counts) for fast iteration.
args
{
"subject": "Characterize dev target registers & proportions",
"description": "Decode multi_dev.npy, classify docs into Wikipedia/news/web/techQA, measure proportions to guide balanced selection.",
"activeForm": "Characterizing dev target registers"
}TaskCreate result
Task #1 created successfully: Characterize dev target registers & proportions
[metadata] {"task": {"id": "1", "subject": "Characterize dev target registers & proportions"}}args
{
"subject": "Build domain-matching curation script (curate.py)",
"description": "NB log-count-ratio classifier: positives=decoded dev docs, negatives=random pool. Score all pool docs, rank, select top to fill budget. Produce selection.json reproducibly.",
"activeForm": "Building curation script"
}TaskCreate result
Task #2 created successfully: Build domain-matching curation script (curate.py)
[metadata] {"task": {"id": "2", "subject": "Build domain-matching curation script (curate.py)"}}args
{
"subject": "Validate selection via dev_eval vs random baseline",
"description": "Run frozen train_nano on random baseline and on curated selection; compare held-out PPL. Iterate on selection if time permits.",
"activeForm": "Validating selection with dev_eval"
}TaskCreate result
Task #3 created successfully: Validate selection via dev_eval vs random baseline
[metadata] {"task": {"id": "3", "subject": "Validate selection via dev_eval vs random baseline"}}args
{
"subject": "Write claim.md deliverable",
"description": "Hypothesis / Mechanism (observable other than final perplexity) / Falsification / Transfer.",
"activeForm": "Writing claim.md"
}TaskCreate result
Task #4 created successfully: Write claim.md deliverable
[metadata] {"task": {"id": "4", "subject": "Write claim.md deliverable"}}encrypted chain-of-thought — recorded, not exposed
args
{
"taskId": "1",
"status": "in_progress"
}TaskUpdate result
Updated task #1 status
[metadata] {"success": true, "taskId": "1", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}args
{
"command": "mkdir -p /workspace/artifacts /workspace/submission; cd /workspace; export TOKENIZERS_PARALLELISM=true; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np, json, time, re\ntok=AutoTokenizer.from_pretrained('gpt2')\n\n# ---- DEV register analysis ----\na=np.load('/workspace/data/multi_dev.npy').astype(np.int64)\ndocs=[]; cur=[]\nfor t in a:\n if t==50256:\n if cur: docs.append(cur); cur=[]\n else: cur.append(t)\nif cur: docs.append(cur)\ndef reg(txt):\n if '<p>' in txt or '<pre>' in txt or '<code>' in txt or '</p>' in txt: return 'techQA'\n if ' @-@ ' in txt or ' @,@ ' in txt or ' @.@ ' in txt: return 'wiki'\n # spaced punctuation typical of wikitext: many ' . ' ' , '\n sp = txt.count(' . ')+txt.count(' , ')+txt.count(\\\" 's \\\")\n if sp > len(txt)/200: return 'wiki_like'\n return 'web_news'\nfrom collections import Counter\nc=Counter(); pos={}\ndecoded=[tok.decode(d) for d in docs]\ntokcount=Counter()\nfor i,txt in enumerate(decoded):\n r=reg(txt); c[r]+=1; tokcount[r]+=len(docs[i])\nprint('DEV register doc counts:', dict(c))\nprint('DEV register TOKEN counts:', dict(tokcount), 'total', sum(tokcount.values()))\nfor k in tokcount: print(f' {k}: {100*tokcount[k]/sum(tokcount.values()):.1f}% of tokens')\n# show boundaries: where does each register appear by index\nidxreg=[reg(t) for t in decoded]\n# print run-length blocks\nprev=None; start=0\nfor i,r in enumerate(idxreg+['END']):\n if r!=prev:\n if prev is not None: print(f' docs [{start}:{i}] = {prev} ({i-start} docs)')\n prev=r; start=i\n\" 2>&1 | grep -v \"longer than the specified\\|indexing errors\"",
"description": "Analyze dev register composition and block structure"
}Bash result
DEV register doc counts: {'wiki': 966, 'wiki_like': 741, 'web_news': 639, 'techQA': 359}
DEV register TOKEN counts: {'wiki': 163490, 'wiki_like': 84383, 'web_news': 499543, 'techQA': 249880} total 997296
wiki: 16.4% of tokens
wiki_like: 8.5% of tokens
web_news: 50.1% of tokens
techQA: 25.1% of tokens
docs [0:1] = wiki (1 docs)
docs [1:2] = wiki_like (1 docs)
docs [2:3] = wiki (1 docs)
docs [3:4] = wiki_like (1 docs)
docs [4:5] = wiki (1 docs)
docs [5:9] = wiki_like (4 docs)
docs [9:10] = wiki (1 docs)
docs [10:12] = wiki_like (2 docs)
docs [12:14] = wiki (2 docs)
docs [14:16] = wiki_like (2 docs)
docs [16:17] = wiki (1 docs)
docs [17:21] = wiki_like (4 docs)
docs [21:23] = wiki (2 docs)
docs [23:25] = wiki_like (2 docs)
docs [25:26] = wiki (1 docs)
docs [26:28] = wiki_like (2 docs)
docs [28:31] = wiki (3 docs)
docs [31:33] = wiki_like (2 docs)
docs [33:35] = wiki (2 docs)
docs [35:36] = wiki_like (1 docs)
docs [36:38] = wiki (2 docs)
docs [38:39] = wiki_like (1 docs)
docs [39:43] = wiki (4 docs)
docs [43:45] = wiki_like (2 docs)
docs [45:46] = wiki (1 docs)
docs [46:48] = wiki_like (2 docs)
docs [48:50] = wiki (2 docs)
docs [50:52] = wiki_like (2 docs)
docs [52:55] = wiki (3 docs)
docs [55:60] = wiki_like (5 docs)
docs [60:62] = wiki (2 docs)
docs [62:64] = wiki_like (2 docs)
docs [64:70] = wiki (6 docs)
docs [70:72] = wiki_like (2 docs)
docs [72:73] = wiki (1 docs)
docs [73:74] = wiki_like (1 docs)
docs [74:76] = wiki (2 docs)
docs [76:77] = wiki_like (1 docs)
docs [77:78] = wiki (1 docs)
docs [78:90] = wiki_like (12 docs)
docs [90:91] = wiki (1 docs)
docs [91:92] = wiki_like (1 docs)
docs [92:95] = wiki (3 docs)
docs [95:96] = wiki_like (1 docs)
docs [96:97] = wiki (1 docs)
docs [97:98] = wiki_like (1 docs)
docs [98:125] = wiki (27 docs)
docs [125:126] = wiki_like (1 docs)
docs [126:136] = wiki (10 docs)
docs [136:137] = wiki_like (1 docs)
docs [137:157] = wiki (20 docs)
docs [157:159] = wiki_like (2 docs)
docs [159:160] = wiki (1 docs)
docs [160:161] = wiki_like (1 docs)
docs [161:162] = wiki (1 docs)
docs [162:166] = wiki_like (4 docs)
docs [166:172] = wiki (6 docs)
docs [172:175] = wiki_like (3 docs)
docs [175:189] = wiki (14 docs)
docs [189:194] = wiki_like (5 docs)
docs [194:196] = wiki (2 docs)
docs [196:202] = wiki_like (6 docs)
docs [202:208] = wiki (6 docs)
docs [208:210] = wiki_like (2 docs)
docs [210:212] = wiki (2 docs)
docs [212:214] = wiki_like (2 docs)
docs [214:219] = wiki (5 docs)
docs [219:220] = wiki_like (1 docs)
docs [220:226] = wiki (6 docs)
docs [226:227] = wiki_like (1 docs)
docs [227:232] = wiki (5 docs)
docs [232:233] = wiki_like (1 docs)
docs [233:235] = wiki (2 docs)
docs [235:236] = wiki_like (1 docs)
docs [236:238] = wiki (2 docs)
docs [238:239] = wiki_like (1 docs)
docs [239:243] = wiki (4 docs)
docs [243:244] = wiki_like (1 docs)
docs [244:246] = wiki (2 docs)
docs [246:247] = wiki_like (1 docs)
docs [247:249] = wiki (2 docs)
docs [249:250] = wiki_like (1 docs)
docs [250:251] = wiki (1 docs)
docs [251:253] = wiki_like (2 docs)
docs [253:254] = wiki (1 docs)
docs [254:255] = wiki_like (1 docs)
docs [255:256] = wiki (1 docs)
docs [256:257] = wiki_like (1 docs)
docs [257:260] = wiki (3 docs)
docs [260:261] = wiki_like (1 docs)
docs [261:262] = wiki (1 docs)
docs [262:263] = wiki_like (1 docs)
docs [263:265] = wiki (2 docs)
docs [265:266] = wiki_like (1 docs)
docs [266:272] = wiki (6 docs)
docs [272:273] = wiki_like (1 docs)
docs [273:276] = wiki (3 docs)
docs [276:277] = wiki_like (1 docs)
docs [277:281] = wiki (4 docs)
docs [281:282] = wiki_like (1 docs)
docs [282:284] = wiki (2 docs)
docs [284:285] = wiki_like (1 docs)
docs [285:286] = wiki (1 docs)
docs [286:287] = wiki_like (1 docs)
docs [287:288] = wiki (1 docs)
docs [288:289] = wiki_like (1 docs)
docs [289:290] = wiki (1 docs)
docs [290:292] = wiki_like (2 docs)
docs [292:294] = wiki (2 docs)
docs [294:296] = wiki_like (2 docs)
docs [296:297] = wiki (1 docs)
docs [297:300] = wiki_like (3 docs)
docs [300:303] = wiki (3 docs)
docs [303:305] = wiki_like (2 docs)
docs [305:306] = wiki (1 docs)
docs [306:310] = wiki_like (4 docs)
docs [310:314] = wiki (4 docs)
docs [314:315] = wiki_like (1 docs)
docs [315:318] = wiki (3 docs)
docs [318:319] = wiki_like (1 docs)
docs [319:321] = wiki (2 docs)
docs [321:322] = wiki_like (1 docs)
docs [322:323] = wiki (1 docs)
docs [323:324] = wiki_like (1 docs)
docs [324:326] = wiki (2 docs)
docs [326:331] = wiki_like (5 docs)
docs [331:332] = wiki (1 docs)
docs [332:336] = wiki_like (4 docs)
docs [336:341] = wiki (5 docs)
docs [341:342] = wiki_like (1 docs)
docs [342:345] = wiki (3 docs)
docs [345:347] = wiki_like (2 docs)
docs [347:350] = wiki (3 docs)
docs [350:351] = wiki_like (1 docs)
docs [351:352] = wiki (1 docs)
docs [352:356] = wiki_like (4 docs)
docs [356:357] = wiki (1 docs)
docs [357:359] = wiki_like (2 docs)
docs [359:362] = wiki (3 docs)
docs [362:363] = wiki_like (1 docs)
docs [363:366] = wiki (3 docs)
docs [366:368] = wiki_like (2 docs)
docs [368:369] = wiki (1 docs)
docs [369:371] = wiki_like (2 docs)
docs [371:372] = wiki (1 docs)
docs [372:381] = wiki_like (9 docs)
docs [381:382] = wiki (1 docs)
docs [382:383] = wiki_like (1 docs)
docs [383:384] = wiki (1 docs)
docs [384:391] = wiki_like (7 docs)
docs [391:392] = wiki (1 docs)
docs [392:397] = wiki_like (5 docs)
docs [397:398] = wiki (1 docs)
docs [398:399] = wiki_like (1 docs)
docs [399:402] = wiki (3 docs)
docs [402:404] = wiki_like (2 docs)
docs [404:409] = wiki (5 docs)
docs [409:411] = wiki_like (2 docs)
docs [411:412] = wiki (1 docs)
docs [412:416] = wiki_like (4 docs)
docs [416:417] = wiki (1 docs)
docs [417:420] = wiki_like (3 docs)
docs [420:422] = wiki (2 docs)
docs [422:425] = wiki_like (3 docs)
docs [425:427] = wiki (2 docs)
docs [427:428] = wiki_like (1 docs)
docs [428:429] = wiki (1 docs)
docs [429:430] = wiki_like (1 docs)
docs [430:431] = wiki (1 docs)
docs [431:435] = wiki_like (4 docs)
docs [435:436] = wiki (1 docs)
docs [436:437] = wiki_like (1 docs)
docs [437:439] = wiki (2 docs)
docs [439:440] = wiki_like (1 docs)
docs [440:441] = wiki (1 docs)
docs [441:452] = wiki_like (11 docs)
docs [452:454] = wiki (2 docs)
docs [454:460] = wiki_like (6 docs)
docs [460:462] = wiki (2 docs)
docs [462:463] = wiki_like (1 docs)
docs [463:465] = wiki (2 docs)
docs [465:468] = wiki_like (3 docs)
docs [468:470] = wiki (2 docs)
docs [470:471] = wiki_like (1 docs)
docs [471:472] = wiki (1 docs)
docs [472:476] = wiki_like (4 docs)
docs [476:479] = wiki (3 docs)
docs [479:481] = wiki_like (2 docs)
docs [481:482] = wiki (1 docs)
docs [482:483] = web_news (1 docs)
docs [483:484] = wiki (1 docs)
docs [484:485] = wiki_like (1 docs)
docs [485:486] = wiki (1 docs)
docs [486:487] = wiki_like (1 docs)
docs [487:497] = wiki (10 docs)
docs [497:498] = wiki_like (1 docs)
docs [498:499] = wiki (1 docs)
docs [499:501] = wiki_like (2 docs)
docs [501:502] = wiki (1 docs)
docs [502:503] = wiki_like (1 docs)
docs [503:505] = wiki (2 docs)
docs [505:517] = wiki_like (12 docs)
docs [517:518] = wiki (1 docs)
docs [518:520] = wiki_like (2 docs)
docs [520:524] = wiki (4 docs)
docs [524:525] = wiki_like (1 docs)
docs [525:526] = wiki (1 docs)
docs [526:528] = wiki_like (2 docs)
docs [528:539] = wiki (11 docs)
docs [539:540] = wiki_like (1 docs)
docs [540:545] = wiki (5 docs)
docs [545:547] = wiki_like (2 docs)
docs [547:561] = wiki (14 docs)
docs [561:563] = wiki_like (2 docs)
docs [563:577] = wiki (14 docs)
docs [577:578] = wiki_like (1 docs)
docs [578:581] = wiki (3 docs)
docs [581:582] = wiki_like (1 docs)
docs [582:585] = wiki (3 docs)
docs [585:586] = wiki_like (1 docs)
docs [586:588] = wiki (2 docs)
docs [588:589] = wiki_like (1 docs)
docs [589:596] = wiki (7 docs)
docs [596:597] = wiki_like (1 docs)
docs [597:598] = wiki (1 docs)
docs [598:600] = wiki_like (2 docs)
docs [600:607] = wiki (7 docs)
docs [607:610] = wiki_like (3 docs)
docs [610:611] = wiki (1 docs)
docs [611:612] = wiki_like (1 docs)
docs [612:616] = wiki (4 docs)
docs [616:617] = wiki_like (1 docs)
docs [617:620] = wiki (3 docs)
docs [620:630] = wiki_like (10 docs)
docs [630:632] = wiki (2 docs)
docs [632:633] = wiki_like (1 docs)
docs [633:634] = wiki (1 docs)
docs [634:637] = wiki_like (3 docs)
docs [637:647] = wiki (10 docs)
docs [647:649] = wiki_like (2 docs)
docs [649:652] = wiki (3 docs)
docs [652:655] = wiki_like (3 docs)
docs [655:656] = wiki (1 docs)
docs [656:660] = wiki_like (4 docs)
docs [660:662] = wiki (2 docs)
docs [662:663] = wiki_like (1 docs)
docs [663:664] = wiki (1 docs)
docs [664:665] = wiki_like (1 docs)
docs [665:667] = wiki (2 docs)
docs [667:668] = wiki_like (1 docs)
docs [668:669] = wiki (1 docs)
docs [669:670] = web_news (1 docs)
docs [670:672] = wiki (2 docs)
docs [672:673] = web_news (1 docs)
docs [673:675] = wiki_like (2 docs)
docs [675:678] = wiki (3 docs)
docs [678:679] = wiki_like (1 docs)
docs [679:681] = wiki (2 docs)
docs [681:682] = wiki_like (1 docs)
docs [682:684] = wiki (2 docs)
docs [684:685] = wiki_like (1 docs)
docs [685:689] = wiki (4 docs)
docs [689:691] = wiki_like (2 docs)
docs [691:693] = wiki (2 docs)
docs [693:695] = wiki_like (2 docs)
docs [695:699] = wiki (4 docs)
docs [699:700] = wiki_like (1 docs)
docs [700:704] = wiki (4 docs)
docs [704:706] = wiki_like (2 docs)
docs [706:710] = wiki (4 docs)
docs [710:712] = wiki_like (2 docs)
docs [712:715] = wiki (3 docs)
docs [715:719] = wiki_like (4 docs)
docs [719:720] = wiki (1 docs)
docs [720:722] = wiki_like (2 docs)
docs [722:723] = wiki (1 docs)
docs [723:727] = wiki_like (4 docs)
docs [727:728] = wiki (1 docs)
docs [728:729] = wiki_like (1 docs)
docs [729:730] = wiki (1 docs)
docs [730:732] = wiki_like (2 docs)
docs [732:733] = wiki (1 docs)
docs [733:738] = wiki_like (5 docs)
docs [738:739] = wiki (1 docs)
docs [739:740] = wiki_like (1 docs)
docs [740:744] = wiki (4 docs)
docs [744:745] = wiki_like (1 docs)
docs [745:750] = wiki (5 docs)
docs [750:752] = wiki_like (2 docs)
docs [752:758] = wiki (6 docs)
docs [758:759] = wiki_like (1 docs)
docs [759:760] = web_news (1 docs)
docs [760:761] = wiki_like (1 docs)
docs [761:762] = web_news (1 docs)
docs [762:776] = wiki (14 docs)
docs [776:777] = wiki_like (1 docs)
docs [777:781] = wiki (4 docs)
docs [781:783] = wiki_like (2 docs)
docs [783:787] = wiki (4 docs)
docs [787:788] = wiki_like (1 docs)
docs [788:796] = wiki (8 docs)
docs [796:798] = wiki_like (2 docs)
docs [798:801] = wiki (3 docs)
docs [801:804] = wiki_like (3 docs)
docs [804:808] = wiki (4 docs)
docs [808:810] = wiki_like (2 docs)
docs [810:819] = wiki (9 docs)
docs [819:820] = wiki_like (1 docs)
docs [820:822] = wiki (2 docs)
docs [822:824] = wiki_like (2 docs)
docs [824:827] = wiki (3 docs)
docs [827:828] = wiki_like (1 docs)
docs [828:834] = wiki (6 docs)
docs [834:835] = wiki_like (1 docs)
docs [835:837] = wiki (2 docs)
docs [837:838] = wiki_like (1 docs)
docs [838:842] = wiki (4 docs)
docs [842:845] = wiki_like (3 docs)
docs [845:847] = wiki (2 docs)
docs [847:849] = wiki_like (2 docs)
docs [849:851] = wiki (2 docs)
docs [851:852] = wiki_like (1 docs)
docs [852:863] = wiki (11 docs)
docs [863:864] = wiki_like (1 docs)
docs [864:868] = wiki (4 docs)
docs [868:870] = wiki_like (2 docs)
docs [870:881] = wiki (11 docs)
docs [881:882] = wiki_like (1 docs)
docs [882:884] = wiki (2 docs)
docs [884:892] = wiki_like (8 docs)
docs [892:895] = wiki (3 docs)
docs [895:898] = wiki_like (3 docs)
docs [898:899] = wiki (1 docs)
docs [899:900] = wiki_like (1 docs)
docs [900:902] = wiki (2 docs)
docs [902:903] = wiki_like (1 docs)
docs [903:904] = wiki (1 docs)
docs [904:905] = wiki_like (1 docs)
docs [905:906] = wiki (1 docs)
docs [906:909] = wiki_like (3 docs)
docs [909:911] = wiki (2 docs)
docs [911:919] = wiki_like (8 docs)
docs [919:922] = wiki (3 docs)
docs [922:930] = wiki_like (8 docs)
docs [930:934] = wiki (4 docs)
docs [934:937] = wiki_like (3 docs)
docs [937:944] = wiki (7 docs)
docs [944:949] = wiki_like (5 docs)
docs [949:951] = wiki (2 docs)
docs [951:956] = wiki_like (5 docs)
docs [956:966] = wiki (10 docs)
docs [966:967] = wiki_like (1 docs)
docs [967:968] = wiki (1 docs)
docs [968:969] = wiki_like (1 docs)
docs [969:971] = wiki (2 docs)
docs [971:972] = wiki_like (1 docs)
docs [972:974] = wiki (2 docs)
docs [974:975] = wiki_like (1 docs)
docs [975:977] = wiki (2 docs)
docs [977:980] = wiki_like (3 docs)
docs [980:985] = wiki (5 docs)
docs [985:990] = wiki_like (5 docs)
docs [990:994] = wiki (4 docs)
docs [994:995] = wiki_like (1 docs)
docs [995:996] = wiki (1 docs)
docs [996:1000] = wiki_like (4 docs)
docs [1000:1002] = wiki (2 docs)
docs [1002:1003] = wiki_like (1 docs)
docs [1003:1005] = wiki (2 docs)
docs [1005:1007] = wiki_like (2 docs)
docs [1007:1009] = wiki (2 docs)
docs [1009:1011] = wiki_like (2 docs)
docs [1011:1013] = wiki (2 docs)
docs [1013:1016] = wiki_like (3 docs)
docs [1016:1018] = wiki (2 docs)
docs [1018:1020] = wiki_like (2 docs)
docs [1020:1023] = wiki (3 docs)
docs [1023:1025] = wiki_like (2 docs)
docs [1025:1026] = wiki (1 docs)
docs [1026:1027] = wiki_like (1 docs)
docs [1027:1030] = wiki (3 docs)
docs [1030:1032] = wiki_like (2 docs)
docs [1032:1033] = wiki (1 docs)
docs [1033:1034] = wiki_like (1 docs)
docs [1034:1038] = wiki (4 docs)
docs [1038:1041] = wiki_like (3 docs)
docs [1041:1042] = wiki (1 docs)
docs [1042:1046] = wiki_like (4 docs)
docs [1046:1048] = wiki (2 docs)
docs [1048:1050] = wiki_like (2 docs)
docs [1050:1052] = wiki (2 docs)
docs [1052:1055] = wiki_like (3 docs)
docs [1055:1061] = wiki (6 docs)
docs [1061:1062] = wiki_like (1 docs)
docs [1062:1063] = wiki (1 docs)
docs [1063:1066] = wiki_like (3 docs)
docs [1066:1070] = wiki (4 docs)
docs [1070:1072] = wiki_like (2 docs)
docs [1072:1073] = wiki (1 docs)
docs [1073:1074] = wiki_like (1 docs)
docs [1074:1076] = wiki (2 docs)
docs [1076:1077] = wiki_like (1 docs)
docs [1077:1078] = wiki (1 docs)
docs [1078:1081] = wiki_like (3 docs)
docs [1081:1082] = wiki (1 docs)
docs [1082:1086] = wiki_like (4 docs)
docs [1086:1087] = wiki (1 docs)
docs [1087:1089] = wiki_like (2 docs)
docs [1089:1094] = wiki (5 docs)
docs [1094:1097] = wiki_like (3 docs)
docs [1097:1098] = wiki (1 docs)
docs [1098:1107] = wiki_like (9 docs)
docs [1107:1110] = wiki (3 docs)
docs [1110:1111] = wiki_like (1 docs)
docs [1111:1113] = wiki (2 docs)
docs [1113:1114] = wiki_like (1 docs)
docs [1114:1117] = wiki (3 docs)
docs [1117:1118] = wiki_like (1 docs)
docs [1118:1122] = wiki (4 docs)
docs [1122:1123] = wiki_like (1 docs)
docs [1123:1124] = wiki (1 docs)
docs [1124:1128] = wiki_like (4 docs)
docs [1128:1129] = wiki (1 docs)
docs [1129:1133] = wiki_like (4 docs)
docs [1133:1134] = wiki (1 docs)
docs [1134:1135] = wiki_like (1 docs)
docs [1135:1137] = wiki (2 docs)
docs [1137:1138] = wiki_like (1 docs)
docs [1138:1139] = wiki (1 docs)
docs [1139:1141] = wiki_like (2 docs)
docs [1141:1143] = wiki (2 docs)
docs [1143:1144] = wiki_like (1 docs)
docs [1144:1147] = wiki (3 docs)
docs [1147:1149] = wiki_like (2 docs)
docs [1149:1156] = wiki (7 docs)
docs [1156:1158] = wiki_like (2 docs)
docs [1158:1162] = wiki (4 docs)
docs [1162:1164] = wiki_like (2 docs)
docs [1164:1165] = wiki (1 docs)
docs [1165:1166] = wiki_like (1 docs)
docs [1166:1167] = wiki (1 docs)
docs [1167:1169] = wiki_like (2 docs)
docs [1169:1174] = wiki (5 docs)
docs [1174:1175] = wiki_like (1 docs)
docs [1175:1176] = wiki (1 docs)
docs [1176:1180] = wiki_like (4 docs)
docs [1180:1183] = wiki (3 docs)
docs [1183:1184] = wiki_like (1 docs)
docs [1184:1185] = wiki (1 docs)
docs [1185:1186] = wiki_like (1 docs)
docs [1186:1187] = wiki (1 docs)
docs [1187:1191] = wiki_like (4 docs)
docs [1191:1193] = wiki (2 docs)
docs [1193:1194] = wiki_like (1 docs)
docs [1194:1198] = wiki (4 docs)
docs [1198:1202] = wiki_like (4 docs)
docs [1202:1204] = wiki (2 docs)
docs [1204:1205] = wiki_like (1 docs)
docs [1205:1206] = wiki (1 docs)
docs [1206:1209] = wiki_like (3 docs)
docs [1209:1213] = wiki (4 docs)
docs [1213:1216] = wiki_like (3 docs)
docs [1216:1217] = wiki (1 docs)
docs [1217:1219] = wiki_like (2 docs)
docs [1219:1221] = wiki (2 docs)
docs [1221:1222] = wiki_like (1 docs)
docs [1222:1229] = wiki (7 docs)
docs [1229:1230] = wiki_like (1 docs)
docs [1230:1233] = wiki (3 docs)
docs [1233:1236] = wiki_like (3 docs)
docs [1236:1237] = wiki (1 docs)
docs [1237:1242] = wiki_like (5 docs)
docs [1242:1243] = wiki (1 docs)
docs [1243:1247] = wiki_like (4 docs)
docs [1247:1254] = wiki (7 docs)
docs [1254:1257] = wiki_like (3 docs)
docs [1257:1259] = wiki (2 docs)
docs [1259:1262] = wiki_like (3 docs)
docs [1262:1263] = wiki (1 docs)
docs [1263:1265] = wiki_like (2 docs)
docs [1265:1266] = wiki (1 docs)
docs [1266:1267] = wiki_like (1 docs)
docs [1267:1290] = wiki (23 docs)
docs [1290:1291] = wiki_like (1 docs)
docs [1291:1292] = wiki (1 docs)
docs [1292:1293] = wiki_like (1 docs)
docs [1293:1296] = wiki (3 docs)
docs [1296:1297] = wiki_like (1 docs)
docs [1297:1301] = wiki (4 docs)
docs [1301:1302] = wiki_like (1 docs)
docs [1302:1304] = wiki (2 docs)
docs [1304:1306] = wiki_like (2 docs)
docs [1306:1308] = wiki (2 docs)
docs [1308:1311] = wiki_like (3 docs)
docs [1311:1316] = wiki (5 docs)
docs [1316:1317] = wiki_like (1 docs)
docs [1317:1318] = wiki (1 docs)
docs [1318:1320] = wiki_like (2 docs)
docs [1320:1322] = wiki (2 docs)
docs [1322:1323] = wiki_like (1 docs)
docs [1323:1324] = wiki (1 docs)
docs [1324:1327] = wiki_like (3 docs)
docs [1327:1328] = wiki (1 docs)
docs [1328:1329] = wiki_like (1 docs)
docs [1329:1331] = wiki (2 docs)
docs [1331:1332] = wiki_like (1 docs)
docs [1332:1334] = wiki (2 docs)
docs [1334:1335] = wiki_like (1 docs)
docs [1335:1337] = wiki (2 docs)
docs [1337:1340] = wiki_like (3 docs)
docs [1340:1342] = wiki (2 docs)
docs [1342:1343] = wiki_like (1 docs)
docs [1343:1347] = wiki (4 docs)
docs [1347:1348] = wiki_like (1 docs)
docs [1348:1351] = wiki (3 docs)
docs [1351:1352] = wiki_like (1 docs)
docs [1352:1353] = wiki (1 docs)
docs [1353:1354] = wiki_like (1 docs)
docs [1354:1356] = wiki (2 docs)
docs [1356:1361] = wiki_like (5 docs)
docs [1361:1363] = wiki (2 docs)
docs [1363:1364] = wiki_like (1 docs)
docs [1364:1369] = wiki (5 docs)
docs [1369:1374] = wiki_like (5 docs)
docs [1374:1375] = wiki (1 docs)
docs [1375:1376] = wiki_like (1 docs)
docs [1376:1377] = wiki (1 docs)
docs [1377:1379] = wiki_like (2 docs)
docs [1379:1380] = wiki (1 docs)
docs [1380:1383] = wiki_like (3 docs)
docs [1383:1384] = wiki (1 docs)
docs [1384:1389] = wiki_like (5 docs)
docs [1389:1391] = wiki (2 docs)
docs [1391:1393] = wiki_like (2 docs)
docs [1393:1394] = wiki (1 docs)
docs [1394:1395] = wiki_like (1 docs)
docs [1395:1398] = wiki (3 docs)
docs [1398:1400] = wiki_like (2 docs)
docs [1400:1401] = wiki (1 docs)
docs [1401:1402] = wiki_like (1 docs)
docs [1402:1403] = wiki (1 docs)
docs [1403:1404] = wiki_like (1 docs)
docs [1404:1406] = wiki (2 docs)
docs [1406:1407] = wiki_like (1 docs)
docs [1407:1408] = wiki (1 docs)
docs [1408:1409] = wiki_like (1 docs)
docs [1409:1412] = wiki (3 docs)
docs [1412:1413] = wiki_like (1 docs)
docs [1413:1414] = wiki (1 docs)
docs [1414:1415] = wiki_like (1 docs)
docs [1415:1416] = wiki (1 docs)
docs [1416:1419] = wiki_like (3 docs)
docs [1419:1420] = wiki (1 docs)
docs [1420:1421] = wiki_like (1 docs)
docs [1421:1422] = wiki (1 docs)
docs [1422:1424] = wiki_like (2 docs)
docs [1424:1426] = wiki (2 docs)
docs [1426:1429] = wiki_like (3 docs)
docs [1429:1433] = wiki (4 docs)
docs [1433:1434] = wiki_like (1 docs)
docs [1434:1435] = wiki (1 docs)
docs [1435:1436] = wiki_like (1 docs)
docs [1436:1437] = wiki (1 docs)
docs [1437:1438] = wiki_like (1 docs)
docs [1438:1439] = wiki (1 docs)
docs [1439:1440] = wiki_like (1 docs)
docs [1440:1441] = wiki (1 docs)
docs [1441:1444] = wiki_like (3 docs)
docs [1444:1447] = wiki (3 docs)
docs [1447:1448] = wiki_like (1 docs)
docs [1448:1450] = wiki (2 docs)
docs [1450:1451] = wiki_like (1 docs)
docs [1451:1453] = wiki (2 docs)
docs [1453:1456] = wiki_like (3 docs)
docs [1456:1457] = wiki (1 docs)
docs [1457:1458] = wiki_like (1 docs)
docs [1458:1460] = wiki (2 docs)
docs [1460:1461] = wiki_like (1 docs)
docs [1461:1470] = wiki (9 docs)
docs [1470:1471] = wiki_like (1 docs)
docs [1471:1472] = wiki (1 docs)
docs [1472:1473] = wiki_like (1 docs)
docs [1473:1475] = wiki (2 docs)
docs [1475:1477] = wiki_like (2 docs)
docs [1477:1478] = wiki (1 docs)
docs [1478:1479] = wiki_like (1 docs)
docs [1479:1480] = wiki (1 docs)
docs [1480:1481] = wiki_like (1 docs)
docs [1481:1483] = wiki (2 docs)
docs [1483:1484] = wiki_like (1 docs)
docs [1484:1487] = wiki (3 docs)
docs [1487:1488] = wiki_like (1 docs)
docs [1488:1489] = wiki (1 docs)
docs [1489:1490] = wiki_like (1 docs)
docs [1490:1492] = wiki (2 docs)
docs [1492:1499] = wiki_like (7 docs)
docs [1499:1500] = wiki (1 docs)
docs [1500:1501] = wiki_like (1 docs)
docs [1501:1503] = wiki (2 docs)
docs [1503:1504] = wiki_like (1 docs)
docs [1504:1505] = wiki (1 docs)
docs [1505:1507] = wiki_like (2 docs)
docs [1507:1508] = wiki (1 docs)
docs [1508:1510] = wiki_like (2 docs)
docs [1510:1513] = wiki (3 docs)
docs [1513:1516] = wiki_like (3 docs)
docs [1516:1521] = wiki (5 docs)
docs [1521:1524] = wiki_like (3 docs)
docs [1524:1525] = wiki (1 docs)
docs [1525:1526] = wiki_like (1 docs)
docs [1526:1527] = wiki (1 docs)
docs [1527:1529] = wiki_like (2 docs)
docs [1529:1530] = wiki (1 docs)
docs [1530:1531] = wiki_like (1 docs)
docs [1531:1533] = wiki (2 docs)
docs [1533:1534] = wiki_like (1 docs)
docs [1534:1536] = wiki (2 docs)
docs [1536:1537] = wiki_like (1 docs)
docs [1537:1538] = wiki (1 docs)
docs [1538:1546] = wiki_like (8 docs)
docs [1546:1548] = wiki (2 docs)
docs [1548:1549] = wiki_like (1 docs)
docs [1549:1550] = wiki (1 docs)
docs [1550:1551] = wiki_like (1 docs)
docs [1551:1552] = wiki (1 docs)
docs [1552:1556] = wiki_like (4 docs)
docs [1556:1558] = wiki (2 docs)
docs [1558:1560] = wiki_like (2 docs)
docs [1560:1561] = wiki (1 docs)
docs [1561:1565] = wiki_like (4 docs)
docs [1565:1567] = wiki (2 docs)
docs [1567:1573] = wiki_like (6 docs)
docs [1573:1574] = wiki (1 docs)
docs [1574:1577] = wiki_like (3 docs)
docs [1577:1579] = wiki (2 docs)
docs [1579:1581] = wiki_like (2 docs)
docs [1581:1582] = wiki (1 docs)
docs [1582:1584] = wiki_like (2 docs)
docs [1584:1593] = wiki (9 docs)
docs [1593:1595] = wiki_like (2 docs)
docs [1595:1596] = wiki (1 docs)
docs [1596:1597] = wiki_like (1 docs)
docs [1597:1598] = web_news (1 docs)
docs [1598:1599] = wiki (1 docs)
docs [1599:1600] = wiki_like (1 docs)
docs [1600:1601] = wiki (1 docs)
docs [1601:1602] = wiki_like (1 docs)
docs [1602:1607] = wiki (5 docs)
docs [1607:1610] = wiki_like (3 docs)
docs [1610:1612] = wiki (2 docs)
docs [1612:1617] = wiki_like (5 docs)
docs [1617:1618] = wiki (1 docs)
docs [1618:1619] = wiki_like (1 docs)
docs [1619:1627] = wiki (8 docs)
docs [1627:1628] = wiki_like (1 docs)
docs [1628:1630] = wiki (2 docs)
docs [1630:1632] = wiki_like (2 docs)
docs [1632:1635] = wiki (3 docs)
docs [1635:1636] = wiki_like (1 docs)
docs [1636:1637] = wiki (1 docs)
docs [1637:1638] = wiki_like (1 docs)
docs [1638:1640] = wiki (2 docs)
docs [1640:1641] = wiki_like (1 docs)
docs [1641:1646] = wiki (5 docs)
docs [1646:1647] = wiki_like (1 docs)
docs [1647:1655] = wiki (8 docs)
docs [1655:1656] = wiki_like (1 docs)
docs [1656:1659] = wiki (3 docs)
docs [1659:1660] = wiki_like (1 docs)
docs [1660:1665] = wiki (5 docs)
docs [1665:1667] = wiki_like (2 docs)
docs [1667:1668] = wiki (1 docs)
docs [1668:1671] = wiki_like (3 docs)
docs [1671:1676] = wiki (5 docs)
docs [1676:1678] = wiki_like (2 docs)
docs [1678:1679] = wiki (1 docs)
docs [1679:1680] = wiki_like (1 docs)
docs [1680:1683] = wiki (3 docs)
docs [1683:1684] = wiki_like (1 docs)
docs [1684:1685] = wiki (1 docs)
docs [1685:1689] = wiki_like (4 docs)
docs [1689:1690] = wiki (1 docs)
docs [1690:1694] = wiki_like (4 docs)
docs [1694:1695] = wiki (1 docs)
docs [1695:1696] = wiki_like (1 docs)
docs [1696:1699] = wiki (3 docs)
docs [1699:1700] = wiki_like (1 docs)
docs [1700:1701] = wiki (1 docs)
docs [1701:1703] = wiki_like (2 docs)
docs [1703:1706] = wiki (3 docs)
docs [1706:1713] = wiki_like (7 docs)
docs [1713:2346] = web_news (633 docs)
docs [2346:2705] = techQA (359 docs)
[stdout]
DEV register doc counts: {'wiki': 966, 'wiki_like': 741, 'web_news': 639, 'techQA': 359}
DEV register TOKEN counts: {'wiki': 163490, 'wiki_like': 84383, 'web_news': 499543, 'techQA': 249880} total 997296
wiki: 16.4% of tokens
wiki_like: 8.5% of tokens
web_news: 50.1% of tokens
techQA: 25.1% of tokens
docs [0:1] = wiki (1 docs)
docs [1:2] = wiki_like (1 docs)
docs [2:3] = wiki (1 docs)
docs [3:4] = wiki_like (1 docs)
docs [4:5] = wiki (1 docs)
docs [5:9] = wiki_like (4 docs)
docs [9:10] = wiki (1 docs)
docs [10:12] = wiki_like (2 docs)
docs [12:14] = wiki (2 docs)
docs [14:16] = wiki_like (2 docs)
docs [16:17] = wiki (1 docs)
docs [17:21] = wiki_like (4 docs)
docs [21:23] = wiki (2 docs)
docs [23:25] = wiki_like (2 docs)
docs [25:26] = wiki (1 docs)
docs [26:28] = wiki_like (2 docs)
docs [28:31] = wiki (3 docs)
docs [31:33] = wiki_like (2 docs)
docs [33:35] = wiki (2 docs)
docs [35:36] = wiki_like (1 docs)
docs [36:38] = wiki (2 docs)
docs [38:39] = wiki_like (1 docs)
docs [39:43] = wiki (4 docs)
docs [43:45] = wiki_like (2 docs)
docs [45:46] = wiki (1 docs)
docs [46:48] = wiki_like (2 docs)
docs [48:50] = wiki (2 docs)
docs [50:52] = wiki_like (2 docs)
docs [52:55] = wiki (3 docs)
docs [55:60] = wiki_like (5 docs)
docs [60:62] = wiki (2 docs)
docs [62:64] = wiki_like (2 docs)
docs [64:70] = wiki (6 docs)
docs [70:72] = wiki_like (2 docs)
docs [72:73] = wiki (1 docs)
docs [73:74] = wiki_like (1 docs)
docs [74:76] = wiki (2 docs)
docs [76:77] = wiki_like (1 docs)
docs [77:78] = wiki (1 docs)
docs [78:90] = wiki_like (12 docs)
docs [90:91] = wiki (1 docs)
docs [91:92] = wiki_like (1 docs)
docs [92:95] = wiki (3 docs)
docs [95:96] = wiki_like (1 docs)
docs [96:97] = wiki (1 docs)
docs [97:98] = wiki_like (1 docs)
docs [98:125] = wiki (27 docs)
docs [125:126] = wiki_like (1 docs)
docs [126:136] = wiki (10 docs)
docs [136:137] = wiki_like (1 docs)
docs [137:157] = wiki (20 docs)
docs [157:159] = wiki_like (2 docs)
docs [159:160] = wiki (1 docs)
docs [160:161] = wiki_like (1 docs)
docs [161:162] = wiki (1 docs)
docs [162:166] = wiki_like (4 docs)
docs [166:172] = wiki (6 docs)
docs [172:175] = wiki_like (3 docs)
docs [175:189] = wiki (14 docs)
docs [189:194] = wiki_like (5 docs)
docs [194:196] = wiki (2 docs)
docs [196:202] = wiki_like (6 docs)
docs [202:208] = wiki (6 docs)
docs [208:210] = wiki_like (2 docs)
docs [210:212] = wiki (2 docs)
docs [212:214] = wiki_like (2 docs)
docs [214:219] = wiki (5 docs)
docs [219:220] = wiki_like (1 docs)
docs [220:226] = wiki (6 docs)
docs [226:227] = wiki_like (1 docs)
docs [227:232] = wiki (5 docs)
docs [232:233] = wiki_like (1 docs)
docs [233:235] = wiki (2 docs)
docs [235:236] = wiki_like (1 docs)
docs [236:238] = wiki (2 docs)
docs [238:239] = wiki_like (1 docs)
docs [239:243] = wiki (4 docs)
docs [243:244] = wiki_like (1 docs)
docs [244:246] = wiki (2 docs)
docs [246:247] = wiki_like (1 docs)
docs [247:249] = wiki (2 docs)
docs [249:250] = wiki_like (1 docs)
docs [250:251] = wiki (1 docs)
docs [251:253] = wiki_like (2 docs)
docs [253:254] = wiki (1 docs)
docs [254:255] = wiki_like (1 docs)
docs [255:256] = wiki (1 docs)
docs [256:257] = wiki_like (1 docs)
docs [257:260] = wiki (3 docs)
docs [260:261] = wiki_like (1 docs)
docs [261:262] = wiki (1 docs)
docs [262:263] = wiki_like (1 docs)
docs [263:265] = wiki (2 docs)
docs [265:266] = wiki_like (1 docs)
docs [266:272] = wiki (6 docs)
docs [272:273] = wiki_like (1 docs)
docs [273:276] = wiki (3 docs)
docs [276:277] = wiki_like (1 docs)
docs [277:281] = wiki (4 docs)
docs [281:282] = wiki_like (1 docs)
docs [282:284] = wiki (2 docs)
docs [284:285] = wiki_like (1 docs)
docs [285:286] = wiki (1 docs)
docs [286:287] = wiki_like (1 docs)
docs [287:288] = wiki (1 docs)
docs [288:289] = wiki_like (1 docs)
docs [289:290] = wiki (1 docs)
docs [290:292] = wiki_like (2 docs)
docs [292:294] = wiki (2 docs)
docs [294:296] = wiki_like (2 docs)
docs [296:297] = wiki (1 docs)
docs [297:300] = wiki_like (3 docs)
docs [300:303] = wiki (3 docs)
docs [303:305] = wiki_like (2 docs)
docs [305:306] = wiki (1 docs)
docs [306:310] = wiki_like (4 docs)
docs [310:314] = wiki (4 docs)
docs [314:315] = wiki_like (1 docs)
docs [315:318] = wiki (3 docs)
docs [318:319] = wiki_like (1 docs)
docs [319:321] = wiki (2 docs)
docs [321:322] = wiki_like (1 docs)
docs [322:323] = wiki (1 docs)
docs [323:324] = wiki_like (1 docs)
docs [324:326] = wiki (2 docs)
docs [326:331] = wiki_like (5 docs)
docs [331:332] = wiki (1 docs)
docs [332:336] = wiki_like (4 docs)
docs [336:341] = wiki (5 docs)
docs [341:342] = wiki_like (1 docs)
docs [342:345] = wiki (3 docs)
docs [345:347] = wiki_like (2 docs)
docs [347:350] = wiki (3 docs)
docs [350:351] = wiki_like (1 docs)
docs [351:352] = wiki (1 docs)
docs [352:356] = wiki_like (4 docs)
docs [356:357] = wiki (1 docs)
docs [357:359] = wiki_like (2 docs)
docs [359:362] = wiki (3 docs)
docs [362:363] = wiki_like (1 docs)
docs [363:366] = wiki (3 docs)
docs [366:368] = wiki_like (2 docs)
docs [368:369] = wiki (1 docs)
docs [369:371] = wiki_like (2 docs)
docs [371:372] = wiki (1 docs)
docs [372:381] = wiki_like (9 docs)
docs [381:382] = wiki (1 docs)
docs [382:383] = wiki_like (1 docs)
docs [383:384] = wiki (1 docs)
docs [384:391] = wiki_like (7 docs)
docs [391:392] = wiki (1 docs)
docs [392:397] = wiki_like (5 docs)
docs [397:398] = wiki (1 docs)
docs [398:399] = wiki_like (1 docs)
docs [399:402] = wiki (3 docs)
docs [402:404] = wiki_like (2 docs)
docs [404:409] = wiki (5 docs)
docs [409:411] = wiki_like (2 docs)
docs [411:412] = wiki (1 docs)
docs [412:416] = wiki_like (4 docs)
docs [416:417] = wiki (1 docs)
docs [417:420] = wiki_like (3 docs)
docs [420:422] = wiki (2 docs)
docs [422:425] = wiki_like (3 docs)
docs [425:427] = wiki (2 docs)
docs [427:428] = wiki_like (1 docs)
docs [428:429] = wiki (1 docs)
docs [429:430] = wiki_like (1 docs)
docs [430:431] = wiki (1 docs)
docs [431:435] = wiki_like (4 docs)
docs [435:436] = wiki (1 docs)
docs [436:437] = wiki_like (1 docs)
docs [437:439] = wiki (2 docs)
docs [439:440] = wiki_like (1 docs)
docs [440:441] = wiki (1 docs)
docs [441:452] = wiki_like (11 docs)
docs [452:454] = wiki (2 docs)
docs [454:460] = wiki_like (6 docs)
docs [460:462] = wiki (2 docs)
docs [462:463] = wiki_like (1 docs)
docs [463:465] = wiki (2 docs)
docs [465:468] = wiki_like (3 docs)
docs [468:470] = wiki (2 docs)
docs [470:471] = wiki_like (1 docs)
docs [471:472] = wiki (1 docs)
docs [472:476] = wiki_like (4 docs)
docs [476:479] = wiki (3 docs)
docs [479:481] = wiki_like (2 docs)
docs [481:482] = wiki (1 docs)
docs [482:483] = web_news (1 docs)
docs [483:484] = wiki (1 docs)
docs [484:485] = wiki_like (1 docs)
docs [485:486] = wiki (1 docs)
docs [486:487] = wiki_like (1 docs)
docs [487:497] = wiki (10 docs)
docs [497:498] = wiki_like (1 docs)
docs [498:499] = wiki (1 docs)
docs [499:501] = wiki_like (2 docs)
docs [501:502] = wiki (1 docs)
docs [502:503] = wiki_like (1 docs)
docs [503:505] = wiki (2 docs)
docs [505:517] = wiki_like (12 docs)
docs [517:518] = wiki (1 docs)
docs [518:520] = wiki_like (2 docs)
docs [520:524] = wiki (4 docs)
docs [524:525] = wiki_like (1 docs)
docs [525:526] = wiki (1 docs)
docs [526:528] = wiki_like (2 docs)
docs [528:539] = wiki (11 docs)
docs [539:540] = wiki_like (1 docs)
docs [540:545] = wiki (5 docs)
docs [545:547] = wiki_like (2 docs)
docs [547:561] = wiki (14 docs)
docs [561:563] = wiki_like (2 docs)
docs [563:577] = wiki (14 docs)
docs [577:578] = wiki_like (1 docs)
docs [578:581] = wiki (3 docs)
docs [581:582] = wiki_like (1 docs)
docs [582:585] = wiki (3 docs)
docs [585:586] = wiki_like (1 docs)
docs [586:588] = wiki (2 docs)
docs [588:589] = wiki_like (1 docs)
docs [589:596] = wiki (7 docs)
docs [596:597] = wiki_like (1 docs)
docs [597:598] = wiki (1 docs)
docs [598:600] = wiki_like (2 docs)
docs [600:607] = wiki (7 docs)
docs [607:610] = wiki_like (3 docs)
docs [610:611] = wiki (1 docs)
docs [611:612] = wiki_like (1 docs)
docs [612:616] = wiki (4 docs)
docs [616:617] = wiki_like (1 docs)
docs [617:620] = wiki (3 docs)
docs [620:630] = wiki_like (10 docs)
docs [630:632] = wiki (2 docs)
docs [632:633] = wiki_like (1 docs)
docs [633:634] = wiki (1 docs)
docs [634:637] = wiki_like (3 docs)
docs [637:647] = wiki (10 docs)
docs [647:649] = wiki_like (2 docs)
docs [649:652] = wiki (3 docs)
docs [652:655] = wiki_like (3 docs)
docs [655:656] = wiki (1 docs)
docs [656:660] = wiki_like (4 docs)
docs [660:662] = wiki (2 docs)
docs [662:663] = wiki_like (1 docs)
docs [663:664] = wiki (1 docs)
docs [664:665] = wiki_like (1 docs)
docs [665:667] = wiki (2 docs)
docs [667:668] = wiki_like (1 docs)
docs [668:669] = wiki (1 docs)
docs [669:670] = web_news (1 docs)
docs [670:672] = wiki (2 docs)
docs [672:673] = web_news (1 docs)
docs [673:675] = wiki_like (2 docs)
docs [675:678] = wiki (3 docs)
docs [678:679] = wiki_like (1 docs)
docs [679:681] = wiki (2 docs)
docs [681:682] = wiki_like (1 docs)
docs [682:684] = wiki (2 docs)
docs [684:685] = wiki_like (1 docs)
docs [685:689] = wiki (4 docs)
docs [689:691] = wiki_like (2 docs)
docs [691:693] = wiki (2 docs)
docs [693:695] = wiki_like (2 docs)
docs [695:699] = wiki (4 docs)
docs [699:700] = wiki_like (1 docs)
docs [700:704] = wiki (4 docs)
docs [704:706] = wiki_like (2 docs)
docs [706:710] = wiki (4 docs)
docs [710:712] = wiki_like (2 docs)
docs [712:715] = wiki (3 docs)
docs [715:719] = wiki_like (4 docs)
docs [719:720] = wiki (1 docs)
docs [720:722] = wiki_like (2 docs)
docs [722:723] = wiki (1 docs)
docs [723:727] = wiki_like (4 docs)
docs [727:728] = wiki (1 docs)
docs [728:729] = wiki_like (1 docs)
docs [729:730] = wiki (1 docs)
docs [730:732] = wiki_like (2 docs)
docs [732:733] = wiki (1 docs)
docs [733:738] = wiki_like (5 docs)
docs [738:739] = wiki (1 docs)
docs [739:740] = wiki_like (1 docs)
docs [740:744] = wiki (4 docs)
docs [744:745] = wiki_like (1 docs)
docs [745:750] = wiki (5 docs)
docs [750:752] = wiki_like (2 docs)
docs [752:758] = wiki (6 docs)
docs [758:759] = wiki_like (1 docs)
docs [759:760] = web_news (1 docs)
docs [760:761] = wiki_like (1 docs)
docs [761:762] = web_news (1 docs)
docs [762:776] = wiki (14 docs)
docs [776:777] = wiki_like (1 docs)
docs [777:781] = wiki (4 docs)
docs [781:783] = wiki_like (2 docs)
docs [783:787] = wiki (4 docs)
docs [787:788] = wiki_like (1 docs)
docs [788:796] = wiki (8 docs)
docs [796:798] = wiki_like (2 docs)
docs [798:801] = wiki (3 docs)
docs [801:804] = wiki_like (3 docs)
docs [804:808] = wiki (4 docs)
docs [808:810] = wiki_like (2 docs)
docs [810:819] = wiki (9 docs)
docs [819:820] = wiki_like (1 docs)
docs [820:822] = wiki (2 docs)
docs [822:824] = wiki_like (2 docs)
docs [824:827] = wiki (3 docs)
docs [827:828] = wiki_like (1 docs)
docs [828:834] = wiki (6 docs)
docs [834:835] = wiki_like (1 docs)
docs [835:837] = wiki (2 docs)
docs [837:838] = wiki_like (1 docs)
docs [838:842] = wiki (4 docs)
docs [842:845] = wiki_like (3 docs)
docs [845:847] = wiki (2 docs)
docs [847:849] = wiki_like (2 docs)
docs [849:851] = wiki (2 docs)
docs [851:852] = wiki_like (1 docs)
docs [852:863] = wiki (11 docs)
docs [863:864] = wiki_like (1 docs)
docs [864:868] = wiki (4 docs)
docs [868:870] = wiki_like (2 docs)
docs [870:881] = wiki (11 docs)
docs [881:882] = wiki_like (1 docs)
docs [882:884] = wiki (2 docs)
docs [884:892] = wiki_like (8 docs)
docs [892:895] = wiki (3 docs)
docs [895:898] = wiki_like (3 docs)
docs [898:899] = wiki (1 docs)
docs [899:900] = wiki_like (1 docs)
docs [900:902] = wiki (2 docs)
docs [902:903] = wiki_like (1 docs)
docs [903:904] = wiki (1 docs)
docs [904:905] = wiki_like (1 docs)
docs [905:906] = wiki (1 docs)
docs [906:909] = wiki_like (3 docs)
docs [909:911] = wiki (2 docs)
docs [911:919] = wiki_like (8 docs)
docs [919:922] = wiki (3 docs)
docs [922:930] = wiki_like (8 docs)
docs [930:934] = wiki (4 docs)
docs [934:937] = wiki_like (3 docs)
docs [937:944] = wiki (7 docs)
docs [944:949] = wiki_like (5 docs)
docs [949:951] = wiki (2 docs)
docs [951:956] = wiki_like (5 docs)
docs [956:966] = wiki (10 docs)
docs [966:967] = wiki_like (1 docs)
docs [967:968] = wiki (1 docs)
docs [968:969] = wiki_like (1 docs)
docs [969:971] = wiki (2 docs)
docs [971:972] = wiki_like (1 docs)
docs [972:974] = wiki (2 docs)
docs [974:975] = wiki_like (1 docs)
docs [975:977] = wiki (2 docs)
docs [977:980] = wiki_like (3 docs)
docs [980:985] = wiki (5 docs)
docs [985:990] = wiki_like (5 docs)
docs [990:994] = wiki (4 docs)
docs [994:995] = wiki_like (1 docs)
docs [995:996] = wiki (1 docs)
docs [996:1000] = wiki_like (4 docs)
docs [1000:1002] = wiki (2 docs)
docs [1002:1003] = wiki_like (1 docs)
docs [1003:1005] = wiki (2 docs)
docs [1005:1007] = wiki_like (2 docs)
docs [1007:1009] = wiki (2 docs)
docs [1009:1011] = wiki_like (2 docs)
docs [1011:1013] = wiki (2 docs)
docs [1013:1016] = wiki_like (3 docs)
docs [1016:1018] = wiki (2 docs)
docs [1018:1020] = wiki_like (2 docs)
docs [1020:1023] = wiki (3 docs)
docs [1023:1025] = wiki_like (2 docs)
docs [1025:1026] = wiki (1 docs)
docs [1026:1027] = wiki_like (1 docs)
docs [1027:1030] = wiki (3 docs)
docs [1030:1032] = wiki_like (2 docs)
docs [1032:1033] = wiki (1 docs)
docs [1033:1034] = wiki_like (1 docs)
docs [1034:1038] = wiki (4 docs)
docs [1038:1041] = wiki_like (3 docs)
docs [1041:1042] = wiki (1 docs)
docs [1042:1046] = wiki_like (4 docs)
docs [1046:1048] = wiki (2 docs)
docs [1048:1050] = wiki_like (2 docs)
docs [1050:1052] = wiki (2 docs)
docs [1052:1055] = wiki_like (3 docs)
docs [1055:1061] = wiki (6 docs)
docs [1061:1062] = wiki_like (1 docs)
docs [1062:1063] = wiki (1 docs)
docs [1063:1066] = wiki_like (3 docs)
docs [1066:1070] = wiki (4 docs)
docs [1070:1072] = wiki_like (2 docs)
docs [1072:1073] = wiki (1 docs)
docs [1073:1074] = wiki_like (1 docs)
docs [1074:1076] = wiki (2 docs)
docs [1076:1077] = wiki_like (1 docs)
docs [1077:1078] = wiki (1 docs)
docs [1078:1081] = wiki_like (3 docs)
docs [1081:1082] = wiki (1 docs)
docs [1082:1086] = wiki_like (4 docs)
docs [1086:1087] = wiki (1 docs)
docs [1087:1089] = wiki_like (2 docs)
docs [1089:1094] = wiki (5 docs)
docs [1094:1097] = wiki_like (3 docs)
docs [1097:1098] = wiki (1 docs)
docs [1098:1107] = wiki_like (9 docs)
docs [1107:1110] = wiki (3 docs)
docs [1110:1111] = wiki_like (1 docs)
docs [1111:1113] = wiki (2 docs)
docs [1113:1114] = wiki_like (1 docs)
docs [1114:1117] = wiki (3 docs)
docs [1117:1118] = wiki_like (1 docs)
docs [1118:1122] = wiki (4 docs)
docs [1122:1123] = wiki_like (1 docs)
docs [1123:1124] = wiki (1 docs)
docs [1124:1128] = wiki_like (4 docs)
docs [1128:1129] = wiki (1 docs)
docs [1129:1133] = wiki_like (4 docs)
docs [1133:1134] = wiki (1 docs)
docs [1134:1135] = wiki_like (1 docs)
docs [1135:1137] = wiki (2 docs)
docs [1137:1138] = wiki_like (1 docs)
docs [1138:1139] = wiki (1 docs)
docs [1139:1141] = wiki_like (2 docs)
docs [1141:1143] = wiki (2 docs)
docs [1143:1144] = wiki_like (1 docs)
docs [1144:1147] = wiki (3 docs)
docs [1147:1149] = wiki_like (2 docs)
docs [1149:1156] = wiki (7 docs)
docs [1156:1158] = wiki_like (2 docs)
docs [1158:1162] = wiki (4 docs)
docs [1162:1164] = wiki_like (2 docs)
docs [1164:1165] = wiki (1 docs)
docs [1165:1166] = wiki_like (1 docs)
docs [1166:1167] = wiki (1 docs)
docs [1167:1169] = wiki_like (2 docs)
docs [1169:1174] = wiki (5 docs)
docs [1174:1175] = wiki_like (1 docs)
docs [1175:1176] = wiki (1 docs)
docs [1176:1180] = wiki_like (4 docs)
docs [1180:1183] = wiki (3 docs)
docs [1183:1184] = wiki_like (1 docs)
docs [1184:1185] = wiki (1 docs)
docs [1185:1186] = wiki_like (1 docs)
docs [1186:1187] = wiki (1 docs)
docs [1187:1191] = wiki_like (4 docs)
docs [1191:1193] = wiki (2 docs)
docs [1193:1194] = wiki_like (1 docs)
docs [1194:1198] = wiki (4 docs)
docs [1198:1202] = wiki_like (4 docs)
docs [1202:1204] = wiki (2 docs)
docs [1204:1205] = wiki_like (1 docs)
docs [1205:1206] = wiki (1 docs)
docs [1206:1209] = wiki_like (3 docs)
docs [1209:1213] = wiki (4 docs)
docs [1213:1216] = wiki_like (3 docs)
docs [1216:1217] = wiki (1 docs)
docs [1217:1219] = wiki_like (2 docs)
docs [1219:1221] = wiki (2 docs)
docs [1221:1222] = wiki_like (1 docs)
docs [1222:1229] = wiki (7 docs)
docs [1229:1230] = wiki_like (1 docs)
docs [1230:1233] = wiki (3 docs)
docs [1233:1236] = wiki_like (3 docs)
docs [1236:1237] = wiki (1 docs)
docs [1237:1242] = wiki_like (5 docs)
docs [1242:1243] = wiki (1 docs)
docs [1243:1247] = wiki_like (4 docs)
docs [1247:1254] = wiki (7 docs)
docs [1254:1257] = wiki_like (3 docs)
docs [1257:1259] = wiki (2 docs)
docs [1259:1262] = wiki_like (3 docs)
docs [1262:1263] = wiki (1 docs)
docs [1263:1265] = wiki_like (2 docs)
docs [1265:1266] = wiki (1 docs)
docs [1266:1267] = wiki_like (1 docs)
docs [1267:1290] = wiki (23 docs)
docs [1290:1291] = wiki_like (1 docs)
docs [1291:1292] = wiki (1 docs)
docs [1292:1293] = wiki_like (1 docs)
docs [1293:1296] = wiki (3 docs)
docs [1296:1297] = wiki_like (1 docs)
docs [1297:1301] = wiki (4 docs)
docs [1301:1302] = wiki_like (1 docs)
docs [1302:1304] = wiki (2 docs)
docs [1304:1306] = wiki_like (2 docs)
docs [1306:1308] = wiki (2 docs)
docs [1308:1311] = wiki_like (3 docs)
docs [1311:1316] = wiki (5 docs)
docs [1316:1317] = wiki_like (1 docs)
docs [1317:1318] = wiki (1 docs)
docs [1318:1320] = wiki_like (2 docs)
docs [1320:1322] = wiki (2 docs)
docs [1322:1323] = wiki_like (1 docs)
docs [1323:1324] = wiki (1 docs)
docs [1324:1327] = wiki_like (3 docs)
docs [1327:1328] = wiki (1 docs)
docs [1328:1329] = wiki_like (1 docs)
docs [1329:1331] = wiki (2 docs)
docs [1331:1332] = wiki_like (1 docs)
docs [1332:1334] = wiki (2 docs)
docs [1334:1335] = wiki_like (1 docs)
docs [1335:1337] = wiki (2 docs)
docs [1337:1340] = wiki_like (3 docs)
docs [1340:1342] = wiki (2 docs)
docs [1342:1343] = wiki_like (1 docs)
docs [1343:1347] = wiki (4 docs)
docs [1347:1348] = wiki_like (1 docs)
docs [1348:1351] = wiki (3 docs)
docs [1351:1352] = wiki_like (1 docs)
docs [1352:1353] = wiki (1 docs)
docs [1353:1354] = wiki_like (1 docs)
docs [1354:1356] = wiki (2 docs)
docs [1356:1361] = wiki_like (5 docs)
docs [1361:1363] = wiki (2 docs)
docs [1363:1364] = wiki_like (1 docs)
docs [1364:1369] = wiki (5 docs)
docs [1369:1374] = wiki_like (5 docs)
docs [1374:1375] = wiki (1 docs)
docs [1375:1376] = wiki_like (1 docs)
docs [1376:1377] = wiki (1 docs)
docs [1377:1379] = wiki_like (2 docs)
docs [1379:1380] = wiki (1 docs)
docs [1380:1383] = wiki_like (3 docs)
docs [1383:1384] = wiki (1 docs)
docs [1384:1389] = wiki_like (5 docs)
docs [1389:1391] = wiki (2 docs)
docs [1391:1393] = wiki_like (2 docs)
docs [1393:1394] = wiki (1 docs)
docs [1394:1395] = wiki_like (1 docs)
docs [1395:1398] = wiki (3 docs)
docs [1398:1400] = wiki_like (2 docs)
docs [1400:1401] = wiki (1 docs)
docs [1401:1402] = wiki_like (1 docs)
docs [1402:1403] = wiki (1 docs)
docs [1403:1404] = wiki_like (1 docs)
docs [1404:1406] = wiki (2 docs)
docs [1406:1407] = wiki_like (1 docs)
docs [1407:1408] = wiki (1 docs)
docs [1408:1409] = wiki_like (1 docs)
docs [1409:1412] = wiki (3 docs)
docs [1412:1413] = wiki_like (1 docs)
docs [1413:1414] = wiki (1 docs)
docs [1414:1415] = wiki_like (1 docs)
docs [1415:1416] = wiki (1 docs)
docs [1416:1419] = wiki_like (3 docs)
docs [1419:1420] = wiki (1 docs)
docs [1420:1421] = wiki_like (1 docs)
docs [1421:1422] = wiki (1 docs)
docs [1422:1424] = wiki_like (2 docs)
docs [1424:1426] = wiki (2 docs)
docs [1426:1429] = wiki_like (3 docs)
docs [1429:1433] = wiki (4 docs)
docs [1433:1434] = wiki_like (1 docs)
docs [1434:1435] = wiki (1 docs)
docs [1435:1436] = wiki_like (1 docs)
docs [1436:1437] = wiki (1 docs)
docs [1437:1438] = wiki_like (1 docs)
docs [1438:1439] = wiki (1 docs)
docs [1439:1440] = wiki_like (1 docs)
docs [1440:1441] = wiki (1 docs)
docs [1441:1444] = wiki_like (3 docs)
docs [1444:1447] = wiki (3 docs)
docs [1447:1448] = wiki_like (1 docs)
docs [1448:1450] = wiki (2 docs)
docs [1450:1451] = wiki_like (1 docs)
docs [1451:1453] = wiki (2 docs)
docs [1453:1456] = wiki_like (3 docs)
docs [1456:1457] = wiki (1 docs)
docs [1457:1458] = wiki_like (1 docs)
docs [1458:1460] = wiki (2 docs)
docs [1460:1461] = wiki_like (1 docs)
docs [1461:1470] = wiki (9 docs)
docs [1470:1471] = wiki_like (1 docs)
docs [1471:1472] = wiki (1 docs)
docs [1472:1473] = wiki_like (1 docs)
docs [1473:1475] = wiki (2 docs)
docs [1475:1477] = wiki_like (2 docs)
docs [1477:1478] = wiki (1 docs)
docs [1478:1479] = wiki_like (1 docs)
docs [1479:1480] = wiki (1 docs)
docs [1480:1481] = wiki_like (1 docs)
docs [1481:1483] = wiki (2 docs)
docs [1483:1484] = wiki_like (1 docs)
docs [1484:1487] = wiki (3 docs)
docs [1487:1488] = wiki_like (1 docs)
docs [1488:1489] = wiki (1 docs)
docs [1489:1490] = wiki_like (1 docs)
docs [1490:1492] = wiki (2 docs)
docs [1492:1499] = wiki_like (7 docs)
docs [1499:1500] = wiki (1 docs)
docs [1500:1501] = wiki_like (1 docs)
docs [1501:1503] = wiki (2 docs)
docs [1503:1504] = wiki_like (1 docs)
docs [1504:1505] = wiki (1 docs)
docs [1505:1507] = wiki_like (2 docs)
docs [1507:1508] = wiki (1 docs)
docs [1508:1510] = wiki_like (2 docs)
docs [1510:1513] = wiki (3 docs)
docs [1513:1516] = wiki_like (3 docs)
docs [1516:1521] = wiki (5 docs)
docs [1521:1524] = wiki_like (3 docs)
docs [1524:1525] = wiki (1 docs)
docs [1525:1526] = wiki_like (1 docs)
docs [1526:1527] = wiki (1 docs)
docs [1527:1529] = wiki_like (2 docs)
docs [1529:1530] = wiki (1 docs)
docs [1530:1531] = wiki_like (1 docs)
docs [1531:1533] = wiki (2 docs)
docs [1533:1534] = wiki_like (1 docs)
docs [1534:1536] = wiki (2 docs)
docs [1536:1537] = wiki_like (1 docs)
docs [1537:1538] = wiki (1 docs)
docs [1538:1546] = wiki_like (8 docs)
docs [1546:1548] = wiki (2 docs)
docs [1548:1549] = wiki_like (1 docs)
docs [1549:1550] = wiki (1 docs)
docs [1550:1551] = wiki_like (1 docs)
docs [1551:1552] = wiki (1 docs)
docs [1552:1556] = wiki_like (4 docs)
docs [1556:1558] = wiki (2 docs)
docs [1558:1560] = wiki_like (2 docs)
docs [1560:1561] = wiki (1 docs)
docs [1561:1565] = wiki_like (4 docs)
docs [1565:1567] = wiki (2 docs)
docs [1567:1573] = wiki_like (6 docs)
docs [1573:1574] = wiki (1 docs)
docs [1574:1577] = wiki_like (3 docs)
docs [1577:1579] = wiki (2 docs)
docs [1579:1581] = wiki_like (2 docs)
docs [1581:1582] = wiki (1 docs)
docs [1582:1584] = wiki_like (2 docs)
docs [1584:1593] = wiki (9 docs)
docs [1593:1595] = wiki_like (2 docs)
docs [1595:1596] = wiki (1 docs)
docs [1596:1597] = wiki_like (1 docs)
docs [1597:1598] = web_news (1 docs)
docs [1598:1599] = wiki (1 docs)
docs [1599:1600] = wiki_like (1 docs)
docs [1600:1601] = wiki (1 docs)
docs [1601:1602] = wiki_like (1 docs)
docs [1602:1607] = wiki (5 docs)
docs [1607:1610] = wiki_like (3 docs)
docs [1610:1612] = wiki (2 docs)
docs [1612:1617] = wiki_like (5 docs)
docs [1617:1618] = wiki (1 docs)
docs [1618:1619] = wiki_like (1 docs)
docs [1619:1627] = wiki (8 docs)
docs [1627:1628] = wiki_like (1 docs)
docs [1628:1630] = wiki (2 docs)
docs [1630:1632] = wiki_like (2 docs)
docs [1632:1635] = wiki (3 docs)
docs [1635:1636] = wiki_like (1 docs)
docs [1636:1637] = wiki (1 docs)
docs [1637:1638] = wiki_like (1 docs)
docs [1638:1640] = wiki (2 docs)
docs [1640:1641] = wiki_like (1 docs)
docs [1641:1646] = wiki (5 docs)
docs [1646:1647] = wiki_like (1 docs)
docs [1647:1655] = wiki (8 docs)
docs [1655:1656] = wiki_like (1 docs)
docs [1656:1659] = wiki (3 docs)
docs [1659:1660] = wiki_like (1 docs)
docs [1660:1665] = wiki (5 docs)
docs [1665:1667] = wiki_like (2 docs)
docs [1667:1668] = wiki (1 docs)
docs [1668:1671] = wiki_like (3 docs)
docs [1671:1676] = wiki (5 docs)
docs [1676:1678] = wiki_like (2 docs)
docs [1678:1679] = wiki (1 docs)
docs [1679:1680] = wiki_like (1 docs)
docs [1680:1683] = wiki (3 docs)
docs [1683:1684] = wiki_like (1 docs)
docs [1684:1685] = wiki (1 docs)
docs [1685:1689] = wiki_like (4 docs)
docs [1689:1690] = wiki (1 docs)
docs [1690:1694] = wiki_like (4 docs)
docs [1694:1695] = wiki (1 docs)
docs [1695:1696] = wiki_like (1 docs)
docs [1696:1699] = wiki (3 docs)
docs [1699:1700] = wiki_like (1 docs)
docs [1700:1701] = wiki (1 docs)
docs [1701:1703] = wiki_like (2 docs)
docs [1703:1706] = wiki (3 docs)
docs [1706:1713] = wiki_like (7 docs)
docs [1713:2346] = web_news (633 docs)
docs [2346:2705] = techQA (359 docs)
[metadata] {"noOutputExpected": false}The dev set has a clean 3-block structure by tokens: Wikipedia/encyclopedic ~25% (docs 0–1712), web+news ~50% (docs 1713–2345), technical Q&A ~25% (docs 2346–2704). This matches "equal parts" of 4 registers (web+news = 2 registers = 50%).
My strategy: register-balanced domain-matching. Build a per-register Naive-Bayes log-count-ratio classifier (positives = that register's decoded dev docs, negatives = random pool), score every pool doc, then interleave the best docs per register so the consumed 12M tokens match the target's 25/50/25 mix. This both filters junk (target-vs-pool) and matches the eval distribution (minimizes cross-entropy).
Let me write curate.py.
args
{
"taskId": "1",
"status": "completed"
}TaskUpdate result
Updated task #1 status
[metadata] {"success": true, "taskId": "1", "updatedFields": ["status"], "statusChange": {"from": "in_progress", "to": "completed"}}args
{
"taskId": "2",
"status": "in_progress"
}TaskUpdate result
Updated task #2 status
[metadata] {"success": true, "taskId": "2", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}encrypted chain-of-thought — recorded, not exposed
args
{
"file_path": "/workspace/submission/curate.py",
"content": "#!/usr/bin/env python3\n\"\"\"Curate a raw web pool into a training selection for a small LM.\n\nSTATED CRITERION\n----------------\nThe scoring target is a BROAD high-quality English mix in three registers\n(measured on the disclosed dev target `multi_dev.npy`):\n ~25% encyclopedic (Wikipedia) ~50% general web + news prose ~25% technical Q&A\n\nWe therefore select pool documents by *domain match to that target*, register by\nregister, and assemble them so that the first 12M tokens the trainer consumes\nreproduce the target's register mix. Concretely:\n\n 1. Decode `multi_dev.npy` (the disclosed target, GPT-2 tokens) back to text and\n split it into its three contiguous register blocks -> these are POSITIVES.\n 2. Take a random sample of the pool as the background/NEGATIVE class.\n 3. For each register, fit a Naive-Bayes log-count-ratio linear scorer over hashed\n word uni+bigram features (NBSVM-style). A pool doc's score under register R is\n the length-normalised sum of NB weights of its features = evidence that the\n doc reads like register R rather than like average web.\n 4. Rank pool docs within each register; drop obvious junk (too short / low\n alpha-ratio / duplicate). Interleave the three ranked lists by *token* quota\n (25/50/25) so the consumed 12M-token budget matches the target mix, best\n documents first. Append the remaining ranked docs as backfill.\n\nOutput: /workspace/submission/selection.json (pool ids, priority order).\n\nThis is a pure function of the pool + the disclosed dev target; nothing is\nhand-picked. Re-running reproduces the same selection (deterministic hashing +\nfixed seeds).\n\"\"\"\nimport json, re, sys, zlib, time, os\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEVNPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/workspace/artifacts\"\nos.makedirs(CACHE, exist_ok=True)\n\nBUDGET = 12_000_000 # trainer stops here\nEOS = 50256\nV = 1 << 20 # hashed feature dimension\nNEG_SAMPLE = 20_000 # background docs\nMIN_TOKENS = 32 # drop tiny fragments\nSEED = 0\n# target register token mix (from dev): wiki / web+news / techQA\nMIX = {\"wiki\": 0.25, \"webnews\": 0.50, \"techqa\": 0.25}\n# dev register block boundaries (doc index into the EOS-split dev docs)\nDEV_BLOCKS = {\"wiki\": (0, 1713), \"webnews\": (1713, 2346), \"techqa\": (2346, 100000)}\nTARGET_TOK = 16_000_000 # build a balanced list up to here, then backfill\nBACKFILL_TOK = 24_000_000 # total tokens of ids to emit (margin over budget)\n\nWORD = re.compile(r\"[a-z0-9]+\")\n\ndef norm(text):\n \"\"\"Lowercase; undo WikiText artifacts so encyclopedic positives read like raw\n web (content, not formatting, should drive the match).\"\"\"\n t = text.lower()\n t = t.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\n return t\n\ndef feats(text, cap=400):\n \"\"\"Hashed uni+bigram feature ids (deterministic crc32), unique per doc.\n `cap` words keeps scoring O(1) per doc and length-normalises implicitly.\"\"\"\n w = WORD.findall(norm(text))[:cap]\n if not w:\n return np.empty(0, dtype=np.int64)\n ids = [zlib.crc32(x.encode()) & (V - 1) for x in w]\n for i in range(len(w) - 1):\n ids.append(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (V - 1))\n return np.unique(np.asarray(ids, dtype=np.int64))\n\ndef load_pool():\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n ids.append(r[\"id\"]); texts.append(r[\"text\"])\n return ids, texts\n\ndef dev_register_texts():\n a = np.load(DEVNPY).astype(np.int64)\n docs, cur = [], []\n for t in a:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else:\n cur.append(t)\n if cur: docs.append(cur)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n out = {}\n for reg, (lo, hi) in DEV_BLOCKS.items():\n out[reg] = [tok.decode(d) for d in docs[lo:hi]]\n return out\n\ndef nb_weights(pos_texts, neg_texts, alpha=1.0):\n \"\"\"NBSVM log-count-ratio r over the hashed feature space.\"\"\"\n p = np.zeros(V, dtype=np.float64)\n q = np.zeros(V, dtype=np.float64)\n for t in pos_texts:\n np.add.at(p, feats(t), 1.0)\n for t in neg_texts:\n np.add.at(q, feats(t), 1.0)\n p = (p + alpha) / (p.sum() + alpha * V)\n q = (q + alpha) / (q.sum() + alpha * V)\n return np.log(p) - np.log(q)\n\ndef main():\n t0 = time.time()\n print(\"loading pool ...\", flush=True)\n ids, texts = load_pool()\n N = len(ids)\n id_arr = np.asarray(ids)\n print(f\" {N} docs\", flush=True)\n\n # ---- exact GPT-2 token counts (same tokenizer the packer uses) ----\n cnt_path = f\"{CACHE}/pool_counts.npy\"\n if os.path.exists(cnt_path):\n counts = np.load(cnt_path)\n else:\n os.environ[\"TOKENIZERS_PARALLELISM\"] = \"true\"\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n counts = np.zeros(N, dtype=np.int64)\n B = 4000\n for s in range(0, N, B):\n enc = tok(texts[s:s+B], add_special_tokens=False)[\"input_ids\"]\n counts[s:s+len(enc)] = [len(e) for e in enc]\n np.save(cnt_path, counts)\n counts_eos = counts + 1 # packer appends an EOS per doc\n print(f\" tokenized ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- feature cache for the whole pool (unique hashed ids per doc) ----\n fz_path = f\"{CACHE}/pool_feats.npz\"\n if os.path.exists(fz_path):\n z = np.load(fz_path, allow_pickle=True)\n flat, off = z[\"flat\"], z[\"off\"]\n else:\n flat_list, off = [], np.zeros(N + 1, dtype=np.int64)\n for i, t in enumerate(texts):\n f = feats(t)\n flat_list.append(f); off[i+1] = off[i] + len(f)\n flat = np.concatenate(flat_list) if flat_list else np.empty(0, np.int64)\n np.savez(fz_path, flat=flat, off=off)\n print(f\" featurized ({time.time()-t0:.0f}s)\", flush=True)\n\n def doc_feats(i):\n return flat[off[i]:off[i+1]]\n\n # ---- negatives: random background sample of the pool ----\n rng = np.random.default_rng(SEED)\n neg_idx = rng.choice(N, size=min(NEG_SAMPLE, N), replace=False)\n neg_texts = [texts[i] for i in neg_idx]\n\n # ---- per-register NB scorers ----\n dev_txt = dev_register_texts()\n scores = {}\n for reg in MIX:\n r = nb_weights(dev_txt[reg], neg_texts)\n s = np.empty(N, dtype=np.float32)\n for i in range(N):\n f = doc_feats(i)\n s[i] = r[f].mean() if len(f) else -1e9 # length-normalised evidence\n scores[reg] = s\n print(f\" scored register {reg} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- assign each kept doc to its best-matching register, rank within ----\n stack = np.vstack([scores[r] for r in MIX]) # [3, N]\n reg_names = list(MIX)\n best_reg = stack.argmax(0)\n queues = {}\n for ri, reg in enumerate(reg_names):\n cand = np.where(ok & (best_reg == ri))[0]\n cand = cand[np.argsort(-scores[reg][cand])] # best first\n queues[reg] = list(cand)\n for reg in reg_names:\n print(f\" {reg}: {len(queues[reg])} docs\", flush=True)\n\n # ---- interleave by token quota to match MIX in the consumed budget ----\n pos = {r: 0 for r in reg_names}\n tok_taken = {r: 0 for r in reg_names}\n order, used = [], set()\n total = 0\n while total < TARGET_TOK:\n # pick the register furthest below its token quota that still has docs\n cand = [r for r in reg_names if pos[r] < len(queues[r])]\n if not cand: break\n r = min(cand, key=lambda r: tok_taken[r] / MIX[r])\n di = queues[r][pos[r]]; pos[r] += 1\n if di in used: continue\n used.add(di); order.append(int(id_arr[di]))\n tok_taken[r] += int(counts_eos[di]); total += int(counts_eos[di])\n\n print(\" balanced-portion tokens by register:\",\n {r: tok_taken[r] for r in reg_names},\n \"sum\", total, flush=True)\n\n # ---- backfill: remaining kept docs, globally best target-likeness first ----\n best_score = stack.max(0)\n rest = [i for i in np.argsort(-best_score) if ok[i] and int(id_arr[i]) not in used]\n bt = total\n for i in rest:\n if bt >= BACKFILL_TOK: break\n order.append(int(id_arr[i])); bt += int(counts_eos[i])\n\n json.dump(order, open(OUT, \"w\"))\n print(f\"wrote {len(order)} ids (~{bt/1e6:.1f}M tokens) -> {OUT} ({time.time()-t0:.0f}s)\")\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
File created successfully at: /workspace/submission/curate.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/submission/curate.py", "content": "#!/usr/bin/env python3\n\"\"\"Curate a raw web pool into a training selection for a small LM.\n\nSTATED CRITERION\n----------------\nThe scoring target is a BROAD high-quality English mix in three registers\n(measured on the disclosed dev target `multi_dev.npy`):\n ~25% encyclopedic (Wikipedia) ~50% general web + news prose ~25% technical Q&A\n\nWe therefore select pool documents by *domain match to that target*, register by\nregister, and assemble them so that the first 12M tokens the trainer consumes\nreproduce the target's register mix. Concretely:\n\n 1. Decode `multi_dev.npy` (the disclosed target, GPT-2 tokens) back to text and\n split it into its three contiguous register blocks -> these are POSITIVES.\n 2. Take a random sample of the pool as the background/NEGATIVE class.\n 3. For each register, fit a Naive-Bayes log-count-ratio linear scorer over hashed\n word uni+bigram features (NBSVM-style). A pool doc's score under register R is\n the length-normalised sum of NB weights of its features = evidence that the\n doc reads like register R rather than like average web.\n 4. Rank pool docs within each register; drop obvious junk (too short / low\n alpha-ratio / duplicate). Interleave the three ranked lists by *token* quota\n (25/50/25) so the consumed 12M-token budget matches the target mix, best\n documents first. Append the remaining ranked docs as backfill.\n\nOutput: /workspace/submission/selection.json (pool ids, priority order).\n\nThis is a pure function of the pool + the disclosed dev target; nothing is\nhand-picked. Re-running reproduces the same selection (deterministic hashing +\nfixed seeds).\n\"\"\"\nimport json, re, sys, zlib, time, os\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEVNPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/workspace/artifacts\"\nos.makedirs(CACHE, exist_ok=True)\n\nBUDGET = 12_000_000 # trainer stops here\nEOS = 50256\nV = 1 << 20 # hashed feature dimension\nNEG_SAMPLE = 20_000 # background docs\nMIN_TOKENS = 32 # drop tiny fragments\nSEED = 0\n# target register token mix (from dev): wiki / web+news / techQA\nMIX = {\"wiki\": 0.25, \"webnews\": 0.50, \"techqa\": 0.25}\n# dev register block boundaries (doc index into the EOS-split dev docs)\nDEV_BLOCKS = {\"wiki\": (0, 1713), \"webnews\": (1713, 2346), \"techqa\": (2346, 100000)}\nTARGET_TOK = 16_000_000 # build a balanced list up to here, then backfill\nBACKFILL_TOK = 24_000_000 # total tokens of ids to emit (margin over budget)\n\nWORD = re.compile(r\"[a-z0-9]+\")\n\ndef norm(text):\n \"\"\"Lowercase; undo WikiText artifacts so encyclopedic positives read like raw\n web (content, not formatting, should drive the match).\"\"\"\n t = text.lower()\n t = t.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\n return t\n\ndef feats(text, cap=400):\n \"\"\"Hashed uni+bigram feature ids (deterministic crc32), unique per doc.\n `cap` words keeps scoring O(1) per doc and length-normalises implicitly.\"\"\"\n w = WORD.findall(norm(text))[:cap]\n if not w:\n return np.empty(0, dtype=np.int64)\n ids = [zlib.crc32(x.encode()) & (V - 1) for x in w]\n for i in range(len(w) - 1):\n ids.append(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (V - 1))\n return np.unique(np.asarray(ids, dtype=np.int64))\n\ndef load_pool():\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n ids.append(r[\"id\"]); texts.append(r[\"text\"])\n return ids, texts\n\ndef dev_register_texts():\n a = np.load(DEVNPY).astype(np.int64)\n docs, cur = [], []\n for t in a:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else:\n cur.append(t)\n if cur: docs.append(cur)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n out = {}\n for reg, (lo, hi) in DEV_BLOCKS.items():\n out[reg] = [tok.decode(d) for d in docs[lo:hi]]\n return out\n\ndef nb_weights(pos_texts, neg_texts, alpha=1.0):\n \"\"\"NBSVM log-count-ratio r over the hashed feature space.\"\"\"\n p = np.zeros(V, dtype=np.float64)\n q = np.zeros(V, dtype=np.float64)\n for t in pos_texts:\n np.add.at(p, feats(t), 1.0)\n for t in neg_texts:\n np.add.at(q, feats(t), 1.0)\n p = (p + alpha) / (p.sum() + alpha * V)\n q = (q + alpha) / (q.sum() + alpha * V)\n return np.log(p) - np.log(q)\n\ndef main():\n t0 = time.time()\n print(\"loading pool ...\", flush=True)\n ids, texts = load_pool()\n N = len(ids)\n id_arr = np.asarray(ids)\n print(f\" {N} docs\", flush=True)\n\n # ---- exact GPT-2 token counts (same tokenizer the packer uses) ----\n cnt_path = f\"{CACHE}/pool_counts.npy\"\n if os.path.exists(cnt_path):\n counts = np.load(cnt_path)\n else:\n os.environ[\"TOKENIZERS_PARALLELISM\"] = \"true\"\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n counts = np.zeros(N, dtype=np.int64)\n B = 4000\n for s in range(0, N, B):\n enc = tok(texts[s:s+B], add_special_tokens=False)[\"input_ids\"]\n counts[s:s+len(enc)] = [len(e) for e in enc]\n np.save(cnt_path, counts)\n counts_eos = counts + 1 # packer appends an EOS per doc\n print(f\" tokenized ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- feature cache for the whole pool (unique hashed ids per doc) ----\n fz_path = f\"{CACHE}/pool_feats.npz\"\n if os.path.exists(fz_path):\n z = np.load(fz_path, allow_pickle=True)\n flat, off = z[\"flat\"], z[\"off\"]\n else:\n flat_list, off = [], np.zeros(N + 1, dtype=np.int64)\n for i, t in enumerate(texts):\n f = feats(t)\n flat_list.append(f); off[i+1] = off[i] + len(f)\n flat = np.concatenate(flat_list) if flat_list else np.empty(0, np.int64)\n np.savez(fz_path, flat=flat, off=off)\n print(f\" featurized ({time.time()-t0:.0f}s)\", flush=True)\n\n def doc_feats(i):\n return flat[off[i]:off[i+1]]\n\n # ---- negatives: random background sample of the pool ----\n rng = np.random.default_rng(SEED)\n neg_idx = rng.choice(N, size=min(NEG_SAMPLE, N), replace=False)\n neg_texts = [texts[i] for i in neg_idx]\n\n # ---- per-register NB scorers ----\n dev_txt = dev_register_texts()\n scores = {}\n for reg in MIX:\n r = nb_weights(dev_txt[reg], neg_texts)\n s = np.empty(N, dtype=np.float32)\n for i in range(N):\n f = doc_feats(i)\n s[i] = r[f].mean() if len(f) else -1e9 # length-normalised evidence\n scores[reg] = s\n print(f\" scored register {reg} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- assign each kept doc to its best-matching register, rank within ----\n stack = np.vstack([scores[r] for r in MIX]) # [3, N]\n reg_names = list(MIX)\n best_reg = stack.argmax(0)\n queues = {}\n for ri, reg in enumerate(reg_names):\n cand = np.where(ok & (best_reg == ri))[0]\n cand = cand[np.argsort(-scores[reg][cand])] # best first\n queues[reg] = list(cand)\n for reg in reg_names:\n print(f\" {reg}: {len(queues[reg])} docs\", flush=True)\n\n # ---- interleave by token quota to match MIX in the consumed budget ----\n pos = {r: 0 for r in reg_names}\n tok_taken = {r: 0 for r in reg_names}\n order, used = [], set()\n total = 0\n while total < TARGET_TOK:\n # pick the register furthest below its token quota that still has docs\n cand = [r for r in reg_names if pos[r] < len(queues[r])]\n if not cand: break\n r = min(cand, key=lambda r: tok_taken[r] / MIX[r])\n di = queues[r][pos[r]]; pos[r] += 1\n if di in used: continue\n used.add(di); order.append(int(id_arr[di]))\n tok_taken[r] += int(counts_eos[di]); total += int(counts_eos[di])\n\n print(\" balanced-portion tokens by register:\",\n {r: tok_taken[r] for r in reg_names},\n \"sum\", total, flush=True)\n\n # ---- backfill: remaining kept docs, globally best target-likeness first ----\n best_score = stack.max(0)\n rest = [i for i in np.argsort(-best_score) if ok[i] and int(id_arr[i]) not in used]\n bt = total\n for i in rest:\n if bt >= BACKFILL_TOK: break\n order.append(int(id_arr[i])); bt += int(counts_eos[i])\n\n json.dump(order, open(OUT, \"w\"))\n print(f\"wrote {len(order)} ids (~{bt/1e6:.1f}M tokens) -> {OUT} ({time.time()-t0:.0f}s)\")\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [], "originalFile": null, "userModified": false}Now run curate.py (first run builds the tokenization + feature caches).
args
{
"command": "export TOKENIZERS_PARALLELISM=true && timeout 560 python3 submission/curate.py 2>&1 | grep -v \"longer than the specified\\|indexing errors\"",
"description": "Run curation pipeline (builds caches)"
}Bash result
Command did not complete within its 120s timeout and was moved to the background (ID: blcz0vi4u). Output is being written to: /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/blcz0vi4u.output. You will be notified when it completes. To check interim output, use Read on that file path.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "blcz0vi4u", "timedOutAfterMs": 120000}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 45; grep -v \"longer than\\|indexing errors\" /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/blcz0vi4u.output 2>/dev/null; echo \"---timer---\"; cat /workspace/.timer/remaining_secs",
"description": "Check curate progress"
}Bash result
loading pool ...
182016 docs
tokenized (89s)
featurized (137s)
scored register wiki (144s)
scored register webnews (151s)
---timer---
4789
[stdout]
loading pool ...
182016 docs
tokenized (89s)
featurized (137s)
scored register wiki (144s)
scored register webnews (151s)
---timer---
4789
[metadata] {"noOutputExpected": false}Good — caches built (tokenize 89s, featurize 48s), scoring is fast (~7s/register). Let me wait for it to finish.
args
{
"command": "sleep 30; grep -v \"longer than\\|indexing errors\" /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/blcz0vi4u.output 2>/dev/null; echo \"---exists?---\"; ls -la /workspace/submission/selection.json 2>/dev/null && python3 -c \"import json;s=json.load(open('/workspace/submission/selection.json'));print('ids',len(s),'first10',s[:10])\"",
"description": "Check curate completion and selection"
}Bash result
loading pool ...
182016 docs
tokenized (89s)
featurized (137s)
scored register wiki (144s)
scored register webnews (151s)
scored register techqa (157s)
quality gate keeps 168181/182016 (172s)
wiki: 24197 docs
webnews: 140319 docs
techqa: 3665 docs
balanced-portion tokens by register: {'wiki': 3995339, 'webnews': 7986522, 'techqa': 4019370} sum 16001231
wrote 8379 ids (~24.0M tokens) -> /workspace/submission/selection.json (172s)
---exists?---
-rw-r--r-- 1 root root 63185 Aug 4 04:55 /workspace/submission/selection.json
ids 8379 first10 [177415, 178539, 165616, 125654, 116421, 163066, 169616, 135468, 8221, 123405]
[stdout]
loading pool ...
182016 docs
tokenized (89s)
featurized (137s)
scored register wiki (144s)
scored register webnews (151s)
scored register techqa (157s)
quality gate keeps 168181/182016 (172s)
wiki: 24197 docs
webnews: 140319 docs
techqa: 3665 docs
balanced-portion tokens by register: {'wiki': 3995339, 'webnews': 7986522, 'techqa': 4019370} sum 16001231
wrote 8379 ids (~24.0M tokens) -> /workspace/submission/selection.json (172s)
---exists?---
-rw-r--r-- 1 root root 63185 Aug 4 04:55 /workspace/submission/selection.json
ids 8379 first10 [177415, 178539, 165616, 125654, 116421, 163066, 169616, 135468, 8221, 123405]
[metadata] {"noOutputExpected": false}Curation succeeded. Register queues: wiki 24k docs, webnews 140k, techQA only 3,665 docs (code/HTML is scarce in the pool) — but it still supplied 4M balanced tokens. Balanced portion is a clean 25/50/25 mix. Selection = 8,379 ids (~24M tokens).
Now the real signal: run dev_eval on my selection, and in parallel prepare a random baseline for comparison. Let me start the curated eval.
args
{
"command": "cd /workspace; cat > /workspace/artifacts/make_random.py <<'PY'\nimport json, numpy as np\nrng=np.random.default_rng(123)\nids=[]\nwith open('/workspace/data/pool.jsonl') as f:\n for line in f:\n ids.append(json.loads(line)['id'])\nsel=list(rng.permutation(ids)[:30000]) # ~random, plenty to cover 12M tokens\njson.dump([int(x) for x in sel], open('/workspace/artifacts/random_sel.json','w'))\nprint('random selection', len(sel), 'ids')\nPY\npython3 /workspace/artifacts/make_random.py",
"description": "Create random baseline selection"
}Bash result
random selection 30000 ids
[stdout]
random selection 30000 ids
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; echo \"=== DEV EVAL: curated selection ===\"; bash dev_eval.sh /workspace/submission/selection.json 2>&1 | grep -v \"longer than\\|indexing errors\"",
"description": "Dev-eval curated selection (pack + train ~2.5min)",
"run_in_background": true
}Bash result
Command running in background with ID: bm535fqpp. Output is being written to: /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/bm535fqpp.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bm535fqpp", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}Curated eval is running (~2.5 min). Let me monitor it.
args
{
"command": "sleep 100; grep -v \"longer than\\|indexing errors\" /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/bm535fqpp.output 2>/dev/null; echo \"---timer---\"; cat /workspace/.timer/remaining_secs",
"description": "Check curated eval progress"
}Bash result
=== DEV EVAL: curated selection ===
packed 12000000 tokens from 8379 ids -> /tmp/dev_train.npy
HELDOUT_PPL 741.871 loss 6.6092 train_s 54 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
---timer---
4609
[stdout]
=== DEV EVAL: curated selection ===
packed 12000000 tokens from 8379 ids -> /tmp/dev_train.npy
HELDOUT_PPL 741.871 loss 6.6092 train_s 54 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
---timer---
4609
[metadata] {"noOutputExpected": false}Curated selection: held-out PPL = 741.87. Now I need the random baseline to know if that's good. Let me run it.
args
{
"taskId": "2",
"status": "completed"
}TaskUpdate result
Updated task #2 status
[metadata] {"success": true, "taskId": "2", "updatedFields": ["status"], "statusChange": {"from": "in_progress", "to": "completed"}}args
{
"taskId": "3",
"status": "in_progress"
}TaskUpdate result
Updated task #3 status
[metadata] {"success": true, "taskId": "3", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; echo \"=== DEV EVAL: RANDOM baseline ===\"; bash dev_eval.sh /workspace/artifacts/random_sel.json 2>&1 | grep -v \"longer than\\|indexing errors\"",
"description": "Dev-eval random baseline",
"run_in_background": true
}Bash result
Command running in background with ID: bl0p70iqn. Output is being written to: /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/bl0p70iqn.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bl0p70iqn", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}While the baseline runs, let me inspect what my selection actually picked (quality sanity-check) — reading a few docs from each part of the priority list.
args
{
"command": "cd /workspace; python3 -c \"\nimport json\nsel=json.load(open('/workspace/submission/selection.json'))\ntext={}\nimport sys\nwant=set(sel[:6]) | set(sel[100:103]) | set(sel[4000:4003])\nwith open('/workspace/data/pool.jsonl') as f:\n for line in f:\n r=json.loads(line)\n if r['id'] in want:\n text[r['id']]=r['text']\n if len(text)==len(want): break\nfor i in list(sel[:6])+list(sel[100:103])+list(sel[4000:4003]):\n t=text.get(i,'')\n print(f'--- id {i} (len {len(t)} chars) ---')\n print(t[:260].replace(chr(10),' | '))\n print()\n\"",
"description": "Inspect sample of selected docs"
}Bash result
--- id 177415 (len 28149 chars) ---
isWiki, The Collaborative Numismatics Project | Recent Additions | Price Reductions | Gold | Greek | Greek Gold | Archaic Origins | Classical Fine Art | Persian Empire | Celtic & Tribal | Geographic - All Periods | Britain | Gaul | Iberia | Italy | Sicily | Akragas | Gela | Himera | Kamarina | Katan
--- id 178539 (len 15572 chars) ---
| No | Yes<|endoftext|>AdsApp.SitelinkIterator | Google Ads scripts | Google Developers | Google Ads scripts | list | 所有产品 | 首页 | 指南 | 参考网页 | 示例 | 支持 | Solutions | 首页 | 指南 | 参考网页 | 示例 | 支持 | Solutions | 所有产品 | What's New | AdsApp | 概览 | Ad customizers | Items | AdCustomizerItem | AdCustomizerItemBuilder | AdC
--- id 165616 (len 11710 chars) ---
. | POWERED by | X<|endoftext|>精品四声母域名出售_—1224域名站 | 1224域名站 | 您正在访问的这些域名都可以转让! 好域名,是一道亮丽的风景! | 首页 | 单词 | 三声母 | 三数字 | 四数字 | 五数字 | 三杂 | 四声母 | 双拼 | 其他 | 已售 | wwww.cn | 四维,旺旺旺旺(豹子) | 询价 | Since | cnmq.com | 中国民企,cn开 | 询价 | Since 2009 | cnlx.com | 中国留学,中国旅行,cn开 | 询价 | Since 2009 | kpck.cn | 询价 | Since 2012 | wbym.cn | 询价 | Since
--- id 125654 (len 11487 chars) ---
post.com Silverton Arts / Entertainment Jobs > Silverton > Silverton Arts / Entertainment Jobs,free,TX,Texas,classified ad,classified ads | Silverton Arts / Entertainment Jobs @ Adpost.com Classifieds | Adpost.com Classifieds | USA | Search | Within Employment: | | | ---
--- id 116421 (len 16772 chars) ---
%}<|endoftext|>place match football Fcgb/ol Stade matmut atlantique a bordeaux dimanche 28 avril 2019 réf. 1066484 | recherche billet pas cher | Recherches populaires : MARNE LA VALLEE vers BORDEAUX, PARIS vers MARSEILLE, Hellfest, Muse, Ed Sheeran, Blanche Gardin
--- id 163066 (len 9964 chars) ---
Merritt Birds for Sale, Adoption, Buy, Sell > Merritt > Merritt Birds for Sale, Adoption, Buy, Sell,free,NC,North Carolina,classified ad,classified ads | Merritt Birds for Sale, Adoption, Buy, Sell @ Adpost.com Classifieds | Adpost.com Classifieds | USA | Search | With
--- id 125209 (len 9205 chars) ---
de Trinidad y Tobago | $("#ZoomMenos").after(' | '); //PlayPause.png // | // | $("#FotoAnterior").after(' | '); $("#FotoAnterior").after(' | 5 | '); $('#FotoSiguiente').css("left","130px"); $('#ImagenMostrando').smartZoom({'containerClass':'zoomableContainer','maxScale':1,
--- id 167006 (len 14965 chars) ---
more details. (13)<|endoftext|>Toyota Proace Compact 2.0 D 120 | Yhteystiedot | Valikko | Autot / Proace Compact 2.0 D 120 | Yhteystiedot | Varaa huolto | Varaa koeajo | Laske rahoitus | Ota yhteyttä | Varaa huolto | Varaa koeajo | Laske rahoitus | Ota yhteyttä | Uudet autot | Mallisto |
--- id 157735 (len 6295 chars) ---
�پ کاران | /// System.Int32 /// System.String /// /// /// return this._invoke(this._get_path(), 'VotingY',false,{Id:Id,Val:Val,Ip:Ip},succeededCallback,failedCallback,userContext); }, VotingN:function(Id,Val,Ip,succeededCallback, failedCallback, userContext) { /
--- id 164536 (len 8742 chars) ---
eu<|endoftext|>813-651: Tripp's Jewelery | Home | Existing Customers | Frequently Asked Questions | Welcome to Tripps.com! | SEARCH | Checkout | My Cart : 0 item(s) / $0.00 | Rings | For Faceted Gems | Rounds | Ovals | Emerald Cut | Marquise Cut | Pears | Hearts | Trillions | Squares | Tourmalin
--- id 180260 (len 3834 chars) ---
aji adopted by news channel | About Us | Copyright Policy | Disclaimer | Authors | Contact Us | Assamese | | Sunday, April 21, 2019 | HOME | NORTH EAST | Assam | Meghalaya | Tripura | Mizoram | Manipur | Nagaland | Arunachal | Sikkim | Neighbours | Nepal | Bhutan | China | Myanmar | Bangladesh | Business | En
--- id 142462 (len 6316 chars) ---
for Apple iPad Air 2 (A1567) | EverythingTitleCompatibilityProduct CodeDescription | Search | Advanced Search | Welcome to S&F Electronics! It's 05:00 Sunday 21 April 2019. There are 71 clients connected. | Mobile Phone Spare Parts | Console Spare Parts | Spare Parts Table
[stdout]
--- id 177415 (len 28149 chars) ---
isWiki, The Collaborative Numismatics Project | Recent Additions | Price Reductions | Gold | Greek | Greek Gold | Archaic Origins | Classical Fine Art | Persian Empire | Celtic & Tribal | Geographic - All Periods | Britain | Gaul | Iberia | Italy | Sicily | Akragas | Gela | Himera | Kamarina | Katan
--- id 178539 (len 15572 chars) ---
| No | Yes<|endoftext|>AdsApp.SitelinkIterator | Google Ads scripts | Google Developers | Google Ads scripts | list | 所有产品 | 首页 | 指南 | 参考网页 | 示例 | 支持 | Solutions | 首页 | 指南 | 参考网页 | 示例 | 支持 | Solutions | 所有产品 | What's New | AdsApp | 概览 | Ad customizers | Items | AdCustomizerItem | AdCustomizerItemBuilder | AdC
--- id 165616 (len 11710 chars) ---
. | POWERED by | X<|endoftext|>精品四声母域名出售_—1224域名站 | 1224域名站 | 您正在访问的这些域名都可以转让! 好域名,是一道亮丽的风景! | 首页 | 单词 | 三声母 | 三数字 | 四数字 | 五数字 | 三杂 | 四声母 | 双拼 | 其他 | 已售 | wwww.cn | 四维,旺旺旺旺(豹子) | 询价 | Since | cnmq.com | 中国民企,cn开 | 询价 | Since 2009 | cnlx.com | 中国留学,中国旅行,cn开 | 询价 | Since 2009 | kpck.cn | 询价 | Since 2012 | wbym.cn | 询价 | Since
--- id 125654 (len 11487 chars) ---
post.com Silverton Arts / Entertainment Jobs > Silverton > Silverton Arts / Entertainment Jobs,free,TX,Texas,classified ad,classified ads | Silverton Arts / Entertainment Jobs @ Adpost.com Classifieds | Adpost.com Classifieds | USA | Search | Within Employment: | | | ---
--- id 116421 (len 16772 chars) ---
%}<|endoftext|>place match football Fcgb/ol Stade matmut atlantique a bordeaux dimanche 28 avril 2019 réf. 1066484 | recherche billet pas cher | Recherches populaires : MARNE LA VALLEE vers BORDEAUX, PARIS vers MARSEILLE, Hellfest, Muse, Ed Sheeran, Blanche Gardin
--- id 163066 (len 9964 chars) ---
Merritt Birds for Sale, Adoption, Buy, Sell > Merritt > Merritt Birds for Sale, Adoption, Buy, Sell,free,NC,North Carolina,classified ad,classified ads | Merritt Birds for Sale, Adoption, Buy, Sell @ Adpost.com Classifieds | Adpost.com Classifieds | USA | Search | With
--- id 125209 (len 9205 chars) ---
de Trinidad y Tobago | $("#ZoomMenos").after(' | '); //PlayPause.png // | // | $("#FotoAnterior").after(' | '); $("#FotoAnterior").after(' | 5 | '); $('#FotoSiguiente').css("left","130px"); $('#ImagenMostrando').smartZoom({'containerClass':'zoomableContainer','maxScale':1,
--- id 167006 (len 14965 chars) ---
more details. (13)<|endoftext|>Toyota Proace Compact 2.0 D 120 | Yhteystiedot | Valikko | Autot / Proace Compact 2.0 D 120 | Yhteystiedot | Varaa huolto | Varaa koeajo | Laske rahoitus | Ota yhteyttä | Varaa huolto | Varaa koeajo | Laske rahoitus | Ota yhteyttä | Uudet autot | Mallisto |
--- id 157735 (len 6295 chars) ---
�پ کاران | /// System.Int32 /// System.String /// /// /// return this._invoke(this._get_path(), 'VotingY',false,{Id:Id,Val:Val,Ip:Ip},succeededCallback,failedCallback,userContext); }, VotingN:function(Id,Val,Ip,succeededCallback, failedCallback, userContext) { /
--- id 164536 (len 8742 chars) ---
eu<|endoftext|>813-651: Tripp's Jewelery | Home | Existing Customers | Frequently Asked Questions | Welcome to Tripps.com! | SEARCH | Checkout | My Cart : 0 item(s) / $0.00 | Rings | For Faceted Gems | Rounds | Ovals | Emerald Cut | Marquise Cut | Pears | Hearts | Trillions | Squares | Tourmalin
--- id 180260 (len 3834 chars) ---
aji adopted by news channel | About Us | Copyright Policy | Disclaimer | Authors | Contact Us | Assamese | | Sunday, April 21, 2019 | HOME | NORTH EAST | Assam | Meghalaya | Tripura | Mizoram | Manipur | Nagaland | Arunachal | Sikkim | Neighbours | Nepal | Bhutan | China | Myanmar | Bangladesh | Business | En
--- id 142462 (len 6316 chars) ---
for Apple iPad Air 2 (A1567) | EverythingTitleCompatibilityProduct CodeDescription | Search | Advanced Search | Welcome to S&F Electronics! It's 05:00 Sunday 21 April 2019. There are 71 clients connected. | Mobile Phone Spare Parts | Console Spare Parts | Spare Parts Table
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "grep -v \"longer than\\|indexing errors\" /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/bl0p70iqn.output 2>/dev/null; echo \"---timer---\"; cat /workspace/.timer/remaining_secs",
"description": "Check random baseline result"
}Bash result
=== DEV EVAL: RANDOM baseline ===
---timer---
4509
[stdout]
=== DEV EVAL: RANDOM baseline ===
---timer---
4509
[metadata] {"noOutputExpected": false}Baseline still packing. Meanwhile I found a critical problem: my top-priority docs are junk — navigation menus, classified ads, CJK/foreign text (ids 177415, 165616 etc.). The NB classifier alone ranks word-dense boilerplate highly. I need a real English-prose quality gate. Let me rewrite that part of curate.py (all cached tokenization/features are reused, so re-run is fast and CPU-only — won't contend with the GPU baseline).
args
{
"file_path": "/workspace/submission/curate.py",
"offset": 168,
"limit": 45
}Read result
168
169 # ---- quality gate + dedup ----
170 alpha_ratio = np.array([
171 (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts
172 ])
173 ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)
174 seen, dup = set(), np.zeros(N, dtype=bool)
175 for i in range(N):
176 h = zlib.crc32(norm(texts[i])[:512].encode())
177 if h in seen: dup[i] = True
178 else: seen.add(h)
179 ok &= ~dup
180 print(f" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)", flush=True)
181
182 # ---- assign each kept doc to its best-matching register, rank within ----
183 stack = np.vstack([scores[r] for r in MIX]) # [3, N]
184 reg_names = list(MIX)
185 best_reg = stack.argmax(0)
186 queues = {}
187 for ri, reg in enumerate(reg_names):
188 cand = np.where(ok & (best_reg == ri))[0]
189 cand = cand[np.argsort(-scores[reg][cand])] # best first
190 queues[reg] = list(cand)
191 for reg in reg_names:
192 print(f" {reg}: {len(queues[reg])} docs", flush=True)
193
194 # ---- interleave by token quota to match MIX in the consumed budget ----
195 pos = {r: 0 for r in reg_names}
196 tok_taken = {r: 0 for r in reg_names}
197 order, used = [], set()
198 total = 0
199 while total < TARGET_TOK:
200 # pick the register furthest below its token quota that still has docs
201 cand = [r for r in reg_names if pos[r] < len(queues[r])]
202 if not cand: break
203 r = min(cand, key=lambda r: tok_taken[r] / MIX[r])
204 di = queues[r][pos[r]]; pos[r] += 1
205 if di in used: continue
206 used.add(di); order.append(int(id_arr[di]))
207 tok_taken[r] += int(counts_eos[di]); total += int(counts_eos[di])
208
209 print(" balanced-portion tokens by register:",
210 {r: tok_taken[r] for r in reg_names},
211 "sum", total, flush=True)
212
[metadata] {"type": "text", "file": {"filePath": "/workspace/submission/curate.py", "content": "\n # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- assign each kept doc to its best-matching register, rank within ----\n stack = np.vstack([scores[r] for r in MIX]) # [3, N]\n reg_names = list(MIX)\n best_reg = stack.argmax(0)\n queues = {}\n for ri, reg in enumerate(reg_names):\n cand = np.where(ok & (best_reg == ri))[0]\n cand = cand[np.argsort(-scores[reg][cand])] # best first\n queues[reg] = list(cand)\n for reg in reg_names:\n print(f\" {reg}: {len(queues[reg])} docs\", flush=True)\n\n # ---- interleave by token quota to match MIX in the consumed budget ----\n pos = {r: 0 for r in reg_names}\n tok_taken = {r: 0 for r in reg_names}\n order, used = [], set()\n total = 0\n while total < TARGET_TOK:\n # pick the register furthest below its token quota that still has docs\n cand = [r for r in reg_names if pos[r] < len(queues[r])]\n if not cand: break\n r = min(cand, key=lambda r: tok_taken[r] / MIX[r])\n di = queues[r][pos[r]]; pos[r] += 1\n if di in used: continue\n used.add(di); order.append(int(id_arr[di]))\n tok_taken[r] += int(counts_eos[di]); total += int(counts_eos[di])\n\n print(\" balanced-portion tokens by register:\",\n {r: tok_taken[r] for r in reg_names},\n \"sum\", total, flush=True)\n", "numLines": 45, "startLine": 168, "totalLines": 226}}The baseline finished. Let me get its number, then rewrite the quality gate.
args
{
"command": "grep HELDOUT /tmp/claude-0/-workspace/25296ac2-340c-4280-bfe7-333d9d814974/tasks/bl0p70iqn.output",
"description": "Get baseline PPL"
}Bash result
HELDOUT_PPL 474.471 loss 6.1622 train_s 54 /tmp/dev_train.npy
[stdout]
HELDOUT_PPL 474.471 loss 6.1622 train_s 54 /tmp/dev_train.npy
[metadata] {"noOutputExpected": false}*Critical finding: random baseline = 474.47, my curated = 741.87 — my selection is worse than random.* The NB scorer + weak gate actively surfaced word-dense junk (foreign text, nav menus, classified ads). The quality gate is the decisive lever. Let me replace it with a real English-prose gate (register-aware so techQA code/HTML survives), and cache the signals.
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": " # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)",
"new_string": " # ---- quality gate + dedup ----\n # Random web is full of boilerplate (nav menus, classified ads, link farms) and\n # non-English pages; a random draw already contains enough real prose that naive\n # domain scoring *underperforms* random by selecting long word-dense junk. So we\n # first require each doc to read like fluent English prose (or, for the technical\n # register, like real HTML/code Q&A), THEN rank by domain match.\n qs_path = f\"{CACHE}/pool_qsig.npz\"\n if os.path.exists(qs_path):\n z = np.load(qs_path)\n nw, stop_frac, ascii_ratio, frac_long, is_html = (\n z[\"nw\"], z[\"stop_frac\"], z[\"ascii_ratio\"], z[\"frac_long\"], z[\"is_html\"])\n else:\n nw = np.zeros(N); stop_frac = np.zeros(N); ascii_ratio = np.zeros(N)\n frac_long = np.zeros(N); is_html = np.zeros(N, dtype=bool)\n seg_re = re.compile(r\"[\\n|•·\\t]\")\n for i, t in enumerate(texts):\n head = t[:4000]\n low = head.lower()\n w = WORD.findall(low); nwi = len(w); nw[i] = nwi\n if nwi:\n stop_frac[i] = sum(x in STOP for x in w) / nwi\n letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)\n segs = seg_re.split(head)\n sc = [len(WORD.findall(s)) for s in segs]; tot = sum(sc)\n if tot:\n frac_long[i] = sum(c for c in sc if c >= 8) / tot\n is_html[i] = (\"<p>\" in low or \"<pre\" in low or \"<code\" in low\n or \"</p>\" in low or \"<\" in low or \">\" in low)\n np.savez(qs_path, nw=nw, stop_frac=stop_frac, ascii_ratio=ascii_ratio,\n frac_long=frac_long, is_html=is_html)\n\n prose_ok = (nw >= 50) & (ascii_ratio >= 0.92) & (stop_frac >= 0.22) & (frac_long >= 0.50)\n code_ok = is_html & (nw >= 20) & (ascii_ratio >= 0.85)\n ok = (counts >= MIN_TOKENS) & (prose_ok | code_ok)\n\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} \"\n f\"(prose {int((prose_ok&~dup).sum())}, code {int((code_ok&~dup).sum())}) \"\n f\"({time.time()-t0:.0f}s)\", flush=True)"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": " # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)", "newString": " # ---- quality gate + dedup ----\n # Random web is full of boilerplate (nav menus, classified ads, link farms) and\n # non-English pages; a random draw already contains enough real prose that naive\n # domain scoring *underperforms* random by selecting long word-dense junk. So we\n # first require each doc to read like fluent English prose (or, for the technical\n # register, like real HTML/code Q&A), THEN rank by domain match.\n qs_path = f\"{CACHE}/pool_qsig.npz\"\n if os.path.exists(qs_path):\n z = np.load(qs_path)\n nw, stop_frac, ascii_ratio, frac_long, is_html = (\n z[\"nw\"], z[\"stop_frac\"], z[\"ascii_ratio\"], z[\"frac_long\"], z[\"is_html\"])\n else:\n nw = np.zeros(N); stop_frac = np.zeros(N); ascii_ratio = np.zeros(N)\n frac_long = np.zeros(N); is_html = np.zeros(N, dtype=bool)\n seg_re = re.compile(r\"[\\n|•·\\t]\")\n for i, t in enumerate(texts):\n head = t[:4000]\n low = head.lower()\n w = WORD.findall(low); nwi = len(w); nw[i] = nwi\n if nwi:\n stop_frac[i] = sum(x in STOP for x in w) / nwi\n letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)\n segs = seg_re.split(head)\n sc = [len(WORD.findall(s)) for s in segs]; tot = sum(sc)\n if tot:\n frac_long[i] = sum(c for c in sc if c >= 8) / tot\n is_html[i] = (\"<p>\" in low or \"<pre\" in low or \"<code\" in low\n or \"</p>\" in low or \"<\" in low or \">\" in low)\n np.savez(qs_path, nw=nw, stop_frac=stop_frac, ascii_ratio=ascii_ratio,\n frac_long=frac_long, is_html=is_html)\n\n prose_ok = (nw >= 50) & (ascii_ratio >= 0.92) & (stop_frac >= 0.22) & (frac_long >= 0.50)\n code_ok = is_html & (nw >= 20) & (ascii_ratio >= 0.85)\n ok = (counts >= MIN_TOKENS) & (prose_ok | code_ok)\n\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} \"\n f\"(prose {int((prose_ok&~dup).sum())}, code {int((code_ok&~dup).sum())}) \"\n f\"({time.time()-t0:.0f}s)\", flush=True)", "originalFile": "#!/usr/bin/env python3\n\"\"\"Curate a raw web pool into a training selection for a small LM.\n\nSTATED CRITERION\n----------------\nThe scoring target is a BROAD high-quality English mix in three registers\n(measured on the disclosed dev target `multi_dev.npy`):\n ~25% encyclopedic (Wikipedia) ~50% general web + news prose ~25% technical Q&A\n\nWe therefore select pool documents by *domain match to that target*, register by\nregister, and assemble them so that the first 12M tokens the trainer consumes\nreproduce the target's register mix. Concretely:\n\n 1. Decode `multi_dev.npy` (the disclosed target, GPT-2 tokens) back to text and\n split it into its three contiguous register blocks -> these are POSITIVES.\n 2. Take a random sample of the pool as the background/NEGATIVE class.\n 3. For each register, fit a Naive-Bayes log-count-ratio linear scorer over hashed\n word uni+bigram features (NBSVM-style). A pool doc's score under register R is\n the length-normalised sum of NB weights of its features = evidence that the\n doc reads like register R rather than like average web.\n 4. Rank pool docs within each register; drop obvious junk (too short / low\n alpha-ratio / duplicate). Interleave the three ranked lists by *token* quota\n (25/50/25) so the consumed 12M-token budget matches the target mix, best\n documents first. Append the remaining ranked docs as backfill.\n\nOutput: /workspace/submission/selection.json (pool ids, priority order).\n\nThis is a pure function of the pool + the disclosed dev target; nothing is\nhand-picked. Re-running reproduces the same selection (deterministic hashing +\nfixed seeds).\n\"\"\"\nimport json, re, sys, zlib, time, os\nimport numpy as np\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEVNPY = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/workspace/artifacts\"\nos.makedirs(CACHE, exist_ok=True)\n\nBUDGET = 12_000_000 # trainer stops here\nEOS = 50256\nV = 1 << 20 # hashed feature dimension\nNEG_SAMPLE = 20_000 # background docs\nMIN_TOKENS = 32 # drop tiny fragments\nSEED = 0\n# target register token mix (from dev): wiki / web+news / techQA\nMIX = {\"wiki\": 0.25, \"webnews\": 0.50, \"techqa\": 0.25}\n# dev register block boundaries (doc index into the EOS-split dev docs)\nDEV_BLOCKS = {\"wiki\": (0, 1713), \"webnews\": (1713, 2346), \"techqa\": (2346, 100000)}\nTARGET_TOK = 16_000_000 # build a balanced list up to here, then backfill\nBACKFILL_TOK = 24_000_000 # total tokens of ids to emit (margin over budget)\n\nWORD = re.compile(r\"[a-z0-9]+\")\n\ndef norm(text):\n \"\"\"Lowercase; undo WikiText artifacts so encyclopedic positives read like raw\n web (content, not formatting, should drive the match).\"\"\"\n t = text.lower()\n t = t.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\n return t\n\ndef feats(text, cap=400):\n \"\"\"Hashed uni+bigram feature ids (deterministic crc32), unique per doc.\n `cap` words keeps scoring O(1) per doc and length-normalises implicitly.\"\"\"\n w = WORD.findall(norm(text))[:cap]\n if not w:\n return np.empty(0, dtype=np.int64)\n ids = [zlib.crc32(x.encode()) & (V - 1) for x in w]\n for i in range(len(w) - 1):\n ids.append(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (V - 1))\n return np.unique(np.asarray(ids, dtype=np.int64))\n\ndef load_pool():\n ids, texts = [], []\n with open(POOL) as f:\n for line in f:\n r = json.loads(line)\n ids.append(r[\"id\"]); texts.append(r[\"text\"])\n return ids, texts\n\ndef dev_register_texts():\n a = np.load(DEVNPY).astype(np.int64)\n docs, cur = [], []\n for t in a:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else:\n cur.append(t)\n if cur: docs.append(cur)\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n out = {}\n for reg, (lo, hi) in DEV_BLOCKS.items():\n out[reg] = [tok.decode(d) for d in docs[lo:hi]]\n return out\n\ndef nb_weights(pos_texts, neg_texts, alpha=1.0):\n \"\"\"NBSVM log-count-ratio r over the hashed feature space.\"\"\"\n p = np.zeros(V, dtype=np.float64)\n q = np.zeros(V, dtype=np.float64)\n for t in pos_texts:\n np.add.at(p, feats(t), 1.0)\n for t in neg_texts:\n np.add.at(q, feats(t), 1.0)\n p = (p + alpha) / (p.sum() + alpha * V)\n q = (q + alpha) / (q.sum() + alpha * V)\n return np.log(p) - np.log(q)\n\ndef main():\n t0 = time.time()\n print(\"loading pool ...\", flush=True)\n ids, texts = load_pool()\n N = len(ids)\n id_arr = np.asarray(ids)\n print(f\" {N} docs\", flush=True)\n\n # ---- exact GPT-2 token counts (same tokenizer the packer uses) ----\n cnt_path = f\"{CACHE}/pool_counts.npy\"\n if os.path.exists(cnt_path):\n counts = np.load(cnt_path)\n else:\n os.environ[\"TOKENIZERS_PARALLELISM\"] = \"true\"\n from transformers import AutoTokenizer\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n counts = np.zeros(N, dtype=np.int64)\n B = 4000\n for s in range(0, N, B):\n enc = tok(texts[s:s+B], add_special_tokens=False)[\"input_ids\"]\n counts[s:s+len(enc)] = [len(e) for e in enc]\n np.save(cnt_path, counts)\n counts_eos = counts + 1 # packer appends an EOS per doc\n print(f\" tokenized ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- feature cache for the whole pool (unique hashed ids per doc) ----\n fz_path = f\"{CACHE}/pool_feats.npz\"\n if os.path.exists(fz_path):\n z = np.load(fz_path, allow_pickle=True)\n flat, off = z[\"flat\"], z[\"off\"]\n else:\n flat_list, off = [], np.zeros(N + 1, dtype=np.int64)\n for i, t in enumerate(texts):\n f = feats(t)\n flat_list.append(f); off[i+1] = off[i] + len(f)\n flat = np.concatenate(flat_list) if flat_list else np.empty(0, np.int64)\n np.savez(fz_path, flat=flat, off=off)\n print(f\" featurized ({time.time()-t0:.0f}s)\", flush=True)\n\n def doc_feats(i):\n return flat[off[i]:off[i+1]]\n\n # ---- negatives: random background sample of the pool ----\n rng = np.random.default_rng(SEED)\n neg_idx = rng.choice(N, size=min(NEG_SAMPLE, N), replace=False)\n neg_texts = [texts[i] for i in neg_idx]\n\n # ---- per-register NB scorers ----\n dev_txt = dev_register_texts()\n scores = {}\n for reg in MIX:\n r = nb_weights(dev_txt[reg], neg_texts)\n s = np.empty(N, dtype=np.float32)\n for i in range(N):\n f = doc_feats(i)\n s[i] = r[f].mean() if len(f) else -1e9 # length-normalised evidence\n scores[reg] = s\n print(f\" scored register {reg} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- quality gate + dedup ----\n alpha_ratio = np.array([\n (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts\n ])\n ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)\n seen, dup = set(), np.zeros(N, dtype=bool)\n for i in range(N):\n h = zlib.crc32(norm(texts[i])[:512].encode())\n if h in seen: dup[i] = True\n else: seen.add(h)\n ok &= ~dup\n print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)\n\n # ---- assign each kept doc to its best-matching register, rank within ----\n stack = np.vstack([scores[r] for r in MIX]) # [3, N]\n reg_names = list(MIX)\n best_reg = stack.argmax(0)\n queues = {}\n for ri, reg in enumerate(reg_names):\n cand = np.where(ok & (best_reg == ri))[0]\n cand = cand[np.argsort(-scores[reg][cand])] # best first\n queues[reg] = list(cand)\n for reg in reg_names:\n print(f\" {reg}: {len(queues[reg])} docs\", flush=True)\n\n # ---- interleave by token quota to match MIX in the consumed budget ----\n pos = {r: 0 for r in reg_names}\n tok_taken = {r: 0 for r in reg_names}\n order, used = [], set()\n total = 0\n while total < TARGET_TOK:\n # pick the register furthest below its token quota that still has docs\n cand = [r for r in reg_names if pos[r] < len(queues[r])]\n if not cand: break\n r = min(cand, key=lambda r: tok_taken[r] / MIX[r])\n di = queues[r][pos[r]]; pos[r] += 1\n if di in used: continue\n used.add(di); order.append(int(id_arr[di]))\n tok_taken[r] += int(counts_eos[di]); total += int(counts_eos[di])\n\n print(\" balanced-portion tokens by register:\",\n {r: tok_taken[r] for r in reg_names},\n \"sum\", total, flush=True)\n\n # ---- backfill: remaining kept docs, globally best target-likeness first ----\n best_score = stack.max(0)\n rest = [i for i in np.argsort(-best_score) if ok[i] and int(id_arr[i]) not in used]\n bt = total\n for i in rest:\n if bt >= BACKFILL_TOK: break\n order.append(int(id_arr[i])); bt += int(counts_eos[i])\n\n json.dump(order, open(OUT, \"w\"))\n print(f\"wrote {len(order)} ids (~{bt/1e6:.1f}M tokens) -> {OUT} ({time.time()-t0:.0f}s)\")\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [{"oldStart": 167, "oldLines": 17, "newStart": 167, "newLines": 51, "lines": [" print(f\" scored register {reg} ({time.time()-t0:.0f}s)\", flush=True)", " ", " # ---- quality gate + dedup ----", "- alpha_ratio = np.array([", "- (sum(c.isalpha() for c in t[:2000]) / max(1, len(t[:2000]))) for t in texts", "- ])", "- ok = (counts >= MIN_TOKENS) & (alpha_ratio >= 0.55)", "+ # Random web is full of boilerplate (nav menus, classified ads, link farms) and", "+ # non-English pages; a random draw already contains enough real prose that naive", "+ # domain scoring *underperforms* random by selecting long word-dense junk. So we", "+ # first require each doc to read like fluent English prose (or, for the technical", "+ # register, like real HTML/code Q&A), THEN rank by domain match.", "+ qs_path = f\"{CACHE}/pool_qsig.npz\"", "+ if os.path.exists(qs_path):", "+ z = np.load(qs_path)", "+ nw, stop_frac, ascii_ratio, frac_long, is_html = (", "+ z[\"nw\"], z[\"stop_frac\"], z[\"ascii_ratio\"], z[\"frac_long\"], z[\"is_html\"])", "+ else:", "+ nw = np.zeros(N); stop_frac = np.zeros(N); ascii_ratio = np.zeros(N)", "+ frac_long = np.zeros(N); is_html = np.zeros(N, dtype=bool)", "+ seg_re = re.compile(r\"[\\n|•·\\t]\")", "+ for i, t in enumerate(texts):", "+ head = t[:4000]", "+ low = head.lower()", "+ w = WORD.findall(low); nwi = len(w); nw[i] = nwi", "+ if nwi:", "+ stop_frac[i] = sum(x in STOP for x in w) / nwi", "+ letters = [c for c in head if c.isalpha()]", "+ if letters:", "+ ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)", "+ segs = seg_re.split(head)", "+ sc = [len(WORD.findall(s)) for s in segs]; tot = sum(sc)", "+ if tot:", "+ frac_long[i] = sum(c for c in sc if c >= 8) / tot", "+ is_html[i] = (\"<p>\" in low or \"<pre\" in low or \"<code\" in low", "+ or \"</p>\" in low or \"<\" in low or \">\" in low)", "+ np.savez(qs_path, nw=nw, stop_frac=stop_frac, ascii_ratio=ascii_ratio,", "+ frac_long=frac_long, is_html=is_html)", "+", "+ prose_ok = (nw >= 50) & (ascii_ratio >= 0.92) & (stop_frac >= 0.22) & (frac_long >= 0.50)", "+ code_ok = is_html & (nw >= 20) & (ascii_ratio >= 0.85)", "+ ok = (counts >= MIN_TOKENS) & (prose_ok | code_ok)", "+", " seen, dup = set(), np.zeros(N, dtype=bool)", " for i in range(N):", " h = zlib.crc32(norm(texts[i])[:512].encode())", " if h in seen: dup[i] = True", " else: seen.add(h)", " ok &= ~dup", "- print(f\" quality gate keeps {int(ok.sum())}/{N} ({time.time()-t0:.0f}s)\", flush=True)", "+ print(f\" quality gate keeps {int(ok.sum())}/{N} \"", "+ f\"(prose {int((prose_ok&~dup).sum())}, code {int((code_ok&~dup).sum())}) \"", "+ f\"({time.time()-t0:.0f}s)\", flush=True)", " ", " # ---- assign each kept doc to its best-matching register, rank within ----", " stack = np.vstack([scores[r] for r in MIX]) # [3, N]"]}], "userModified": false, "replaceAll": false}Now add the STOP word set near the top.
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": "WORD = re.compile(r\"[a-z0-9]+\")",
"new_string": "WORD = re.compile(r\"[a-z0-9]+\")\n# common English function words -> fluent prose has high density; keyword/menu\n# boilerplate and non-English text do not.\nSTOP = set(\"the of and a to in is was that for it as with his he be on by at are \"\n \"this or from an but not they which you were her had has have their one \"\n \"all we been would there so what when who will more if out up into than \"\n \"them then its also may i she do no these your can other time about\".split())"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": "WORD = re.compile(r\"[a-z0-9]+\")", "newString": "WORD = re.compile(r\"[a-z0-9]+\")\n# common English function words -> fluent prose has high density; keyword/menu\n# boilerplate and non-English text do not.\nSTOP = set(\"the of and a to in is was that for it as with his he be on by at are \"\n \"this or from an but not they which you were her had has have their one \"\n \"all we been would there so what when who will more if out up into than \"\n \"them then its also may i she do no these your can other time about\".split())", "originalFile": null, "structuredPatch": [{"oldStart": 52, "oldLines": 6, "newStart": 52, "newLines": 12, "lines": [" BACKFILL_TOK = 24_000_000 # total tokens of ids to emit (margin over budget)", " ", " WORD = re.compile(r\"[a-z0-9]+\")", "+# common English function words -> fluent prose has high density; keyword/menu", "+# boilerplate and non-English text do not.", "+STOP = set(\"the of and a to in is was that for it as with his he be on by at are \"", "+ \"this or from an but not they which you were her had has have their one \"", "+ \"all we been would there so what when who will more if out up into than \"", "+ \"them then its also may i she do no these your can other time about\".split())", " ", " def norm(text):", " \"\"\"Lowercase; undo WikiText artifacts so encyclopedic positives read like raw"]}], "userModified": false, "replaceAll": false}Now re-run curation with the strong gate (reuses caches; computes quality signals once).
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=true; timeout 300 python3 submission/curate.py 2>&1 | grep -v \"longer than\\|indexing errors\"; echo \"---timer---\"; cat /workspace/.timer/remaining_secs",
"description": "Re-run curation with strong quality gate"
}Bash result
loading pool ...
182016 docs
tokenized (6s)
featurized (7s)
scored register wiki (18s)
scored register webnews (24s)
scored register techqa (31s)
Traceback (most recent call last):
File "/workspace/submission/curate.py", line 265, in <module>
main()
File "/workspace/submission/curate.py", line 198, in main
ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)
File "/workspace/submission/curate.py", line 198, in <genexpr>
ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)
TypeError: ord() expected a character, but string of length 2 found
---timer---
4328
[stdout]
loading pool ...
182016 docs
tokenized (6s)
featurized (7s)
scored register wiki (18s)
scored register webnews (24s)
scored register techqa (31s)
Traceback (most recent call last):
File "/workspace/submission/curate.py", line 265, in <module>
main()
File "/workspace/submission/curate.py", line 198, in main
ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)
File "/workspace/submission/curate.py", line 198, in <genexpr>
ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)
TypeError: ord() expected a character, but string of length 2 found
---timer---
4328
[metadata] {"noOutputExpected": false}c.lower() can return 2 chars (e.g. Turkish İ). Fix: test the raw char's codepoint against A–Z/a–z ranges.
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": " letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)",
"new_string": " letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(\n 65 <= ord(c) <= 90 or 97 <= ord(c) <= 122 for c in letters\n ) / len(letters)"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": " letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)", "newString": " letters = [c for c in head if c.isalpha()]\n if letters:\n ascii_ratio[i] = sum(\n 65 <= ord(c) <= 90 or 97 <= ord(c) <= 122 for c in letters\n ) / len(letters)", "originalFile": null, "structuredPatch": [{"oldStart": 195, "oldLines": 7, "newStart": 195, "newLines": 9, "lines": [" stop_frac[i] = sum(x in STOP for x in w) / nwi", " letters = [c for c in head if c.isalpha()]", " if letters:", "- ascii_ratio[i] = sum(97 <= ord(c.lower()) <= 122 for c in letters) / len(letters)", "+ ascii_ratio[i] = sum(", "+ 65 <= ord(c) <= 90 or 97 <= ord(c) <= 122 for c in letters", "+ ) / len(letters)", " segs = seg_re.split(head)", " sc = [len(WORD.findall(s)) for s in segs]; tot = sum(sc)", " if tot:"]}], "userModified": false, "replaceAll": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=true; timeout 300 python3 submission/curate.py 2>&1 | grep -v \"longer than\\|indexing errors\"; echo \"---timer---\"; cat /workspace/.timer/remaining_secs",
"description": "Re-run curation with fixed ascii check"
}Bash result
loading pool ...
182016 docs
tokenized (6s)
featurized (7s)
scored register wiki (18s)
scored register webnews (24s)
scored register techqa (30s)
quality gate keeps 129695/182016 (prose 129619, code 174) (105s)
wiki: 17209 docs
webnews: 110400 docs
techqa: 2086 docs
balanced-portion tokens by register: {'wiki': 4890518, 'webnews': 9781669, 'techqa': 1328215} sum 16000402
wrote 22106 ids (~24.0M tokens) -> /workspace/submission/selection.json (105s)
---timer---
4218
[stdout]
loading pool ...
182016 docs
tokenized (6s)
featurized (7s)
scored register wiki (18s)
scored register webnews (24s)
scored register techqa (30s)
quality gate keeps 129695/182016 (prose 129619, code 174) (105s)
wiki: 17209 docs
webnews: 110400 docs
techqa: 2086 docs
balanced-portion tokens by register: {'wiki': 4890518, 'webnews': 9781669, 'techqa': 1328215} sum 16000402
wrote 22106 ids (~24.0M tokens) -> /workspace/submission/selection.json (105s)
---timer---
4218
[metadata] {"noOutputExpected": false}