claude-opus-4-8
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-4-8
- metric
- held-out multi-domain perplexity
- actions
- 92
- duration
- 1h 10m
- recorded
- 2026-08-04
run iddclm-nano-curation__Hk847sf
Everything claude-opus-4-8 did on this task, in order.
download JSON ↗I'll start by understanding my environment, time budget, and the resources available.
args
{
"command": "cat /workspace/.timer/remaining_secs 2>/dev/null; echo \"---\"; ls -la /workspace/; echo \"---DATA---\"; ls -la /workspace/data/; echo \"---SUBMISSION---\"; ls -la /workspace/submission/ 2>/dev/null",
"description": "Check timer, list workspace files"
}Bash result
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 4 03:22 .
drwxr-xr-x 1 root root 160 Aug 4 03:22 ..
drwxr-xr-x 1 root root 60 Aug 4 03:22 .live
drwxr-xr-x 1 root root 100 Aug 4 03:22 .metrics
drwxr-xr-x 1 root root 140 Aug 4 03:22 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---DATA---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---SUBMISSION---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
[stdout]
5390
---
total 14
drwxr-xr-x 1 root root 100 Aug 4 03:22 .
drwxr-xr-x 1 root root 160 Aug 4 03:22 ..
drwxr-xr-x 1 root root 60 Aug 4 03:22 .live
drwxr-xr-x 1 root root 100 Aug 4 03:22 .metrics
drwxr-xr-x 1 root root 140 Aug 4 03:22 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---DATA---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---SUBMISSION---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
[metadata] {"noOutputExpected": false}args
{
"command": "echo \"=== pool head (first 2 lines, truncated) ===\"; head -c 2000 /workspace/data/pool.jsonl; echo; echo \"=== line count ===\"; wc -l /workspace/data/pool.jsonl",
"description": "Peek at pool data"
}Bash result
=== pool head (first 2 lines, truncated) ===
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
=== line count ===
182016 /workspace/data/pool.jsonl
[stdout]
=== pool head (first 2 lines, truncated) ===
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
=== line count ===
182016 /workspace/data/pool.jsonl
[metadata] {"noOutputExpected": false}I have ~90 minutes (5390s). Let me examine the training pipeline and helper scripts to understand exactly how my selection is consumed.
args
{
"command": "echo \"=== train_nano.py ===\"; cat /workspace/train_nano.py; echo; echo \"=== pack_selection.py ===\"; cat /workspace/pack_selection.py; echo; echo \"=== dev_eval.sh ===\"; cat /workspace/dev_eval.sh",
"description": "Read training and packing scripts"
}Bash result
=== train_nano.py ===
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
=== pack_selection.py ===
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
[stdout]
=== train_nano.py ===
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
=== pack_selection.py ===
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
[metadata] {"noOutputExpected": false}args
{
"command": "echo \"=== model.py ===\"; cat /workspace/model.py; echo \"=== timer.sh ===\"; cat /workspace/timer.sh",
"description": "Read model and timer scripts"
}Bash result
=== model.py ===
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
=== timer.sh ===
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[stdout]
=== model.py ===
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
=== timer.sh ===
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[metadata] {"noOutputExpected": false}I now understand the pipeline. Key facts:
pack_selection.pytokenizes my ordered ids until 12M tokens, appending EOS between docs. Order = priority; best docs must come first.- Training is fixed: 3000 iters, evaluated on
multi_dev.npy(dev) or a hidden disjoint sample of the same broad HQ domain. multi_dev.npyis ~1M GPT-2 tokens of the actual target text — I can decode it to learn the target distribution.
This is a data selection / domain-matching problem. Let me set up tasks and start by decoding the target and inspecting the pool.
args
{
"subject": "Decode dev target & inspect pool",
"description": "Decode multi_dev.npy to text to characterize the target domain. Sample pool.jsonl to understand quality distribution.",
"activeForm": "Decoding target and inspecting pool"
}TaskCreate result
Task #1 created successfully: Decode dev target & inspect pool
[metadata] {"task": {"id": "1", "subject": "Decode dev target & inspect pool"}}args
{
"subject": "Build target-similarity scorer over pool",
"description": "Train a classifier (hashed n-gram logistic regression) distinguishing decoded target docs from random pool docs; score all 182k pool docs. Add quality heuristics to filter junk.",
"activeForm": "Building target-similarity scorer"
}TaskCreate result
Task #2 created successfully: Build target-similarity scorer over pool
[metadata] {"task": {"id": "2", "subject": "Build target-similarity scorer over pool"}}args
{
"subject": "Produce ordered selection.json + curate.py",
"description": "Rank pool by score, write selection.json (>=12M tokens, best first) and the reproducible curate.py.",
"activeForm": "Producing selection and curate script"
}TaskCreate result
Task #3 created successfully: Produce ordered selection.json + curate.py
[metadata] {"task": {"id": "3", "subject": "Produce ordered selection.json + curate.py"}}args
{
"subject": "Validate via dev_eval and iterate",
"description": "Run dev_eval.sh on the selection, compare to random baseline, iterate on the criterion. Write claim.md.",
"activeForm": "Validating and iterating"
}TaskCreate result
Task #4 created successfully: Validate via dev_eval and iterate
[metadata] {"task": {"id": "4", "subject": "Validate via dev_eval and iterate"}}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv 2>/dev/null | head; echo \"---python---\"; python3 -c \"import torch, sklearn, numpy, transformers; print('torch', torch.__version__, 'cuda', torch.cuda.is_available()); print('sklearn', sklearn.__version__)\" 2>&1 | head",
"description": "Check GPU and libraries"
}Bash result
name, memory.total [MiB], memory.used [MiB]
NVIDIA H100 80GB HBM3, 81559 MiB, 4 MiB
---python---
Traceback (most recent call last):
File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'sklearn'
[stdout]
name, memory.total [MiB], memory.used [MiB]
NVIDIA H100 80GB HBM3, 81559 MiB, 4 MiB
---python---
Traceback (most recent call last):
File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'sklearn'
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; python3 - <<'PY'\nimport numpy as np\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\na = np.load(\"data/multi_dev.npy\")\nprint(\"dev tokens:\", a.shape, a.dtype, \"min/max\", a.min(), a.max())\n# decode in chunks separated by EOS (50256)\neos = 50256\nprint(\"num EOS (doc separators):\", int((a==eos).sum()))\n# split into docs\ntxt = tok.decode(a[:4000].tolist())\nprint(\"==== first ~4000 tokens decoded ====\")\nprint(txt[:3500])\nPY",
"description": "Decode dev target head"
}Bash result
dev tokens: (1000000,) uint16 min/max 0 50256
num EOS (doc separators): 2704
==== first ~4000 tokens decoded ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , including Sir Philip Lee Brocklehurst , who subscribed £ 2 @,@ 000 ( 2011 equivalent £ 157 @,@ 000 ) to secure a place on the expedition ; author Campbell Mackellar ; and Guinness baron Lord Iveagh , whose contribution was secured less than two weeks before the departure of the expedition ship Nimrod . On 4 August 1907 , Shackleton was appointed a Member of the Royal Victorian Order , 4th Class ( MVO ; the present @-@ day grade of Lieutenant ) .
<|endoftext|> On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land .
<|endoftext|> To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were unstable , precluding the establishment of a safe base there . An extended search for an anchorage at King Edward VII Land proved equally fruitless , so Shackleton was forced to break his undertaking to Scott and set sail for McMurdo Sound , a decision which , according to second officer Arthur Harbord , was " dictated by common sense " in view of the difficulties of ice pressure , coal shortage and the lack of any nearer known base .
<|endoftext|> Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused .
<|endoftext|> The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered th
[stdout]
dev tokens: (1000000,) uint16 min/max 0 50256
num EOS (doc separators): 2704
==== first ~4000 tokens decoded ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , including Sir Philip Lee Brocklehurst , who subscribed £ 2 @,@ 000 ( 2011 equivalent £ 157 @,@ 000 ) to secure a place on the expedition ; author Campbell Mackellar ; and Guinness baron Lord Iveagh , whose contribution was secured less than two weeks before the departure of the expedition ship Nimrod . On 4 August 1907 , Shackleton was appointed a Member of the Royal Victorian Order , 4th Class ( MVO ; the present @-@ day grade of Lieutenant ) .
<|endoftext|> On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land .
<|endoftext|> To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were unstable , precluding the establishment of a safe base there . An extended search for an anchorage at King Edward VII Land proved equally fruitless , so Shackleton was forced to break his undertaking to Scott and set sail for McMurdo Sound , a decision which , according to second officer Arthur Harbord , was " dictated by common sense " in view of the difficulties of ice pressure , coal shortage and the lack of any nearer known base .
<|endoftext|> Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused .
<|endoftext|> The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered th
[metadata] {"noOutputExpected": false}Target is clearly high-quality: WikiText-style Wikipedia (note the @,@/@-@ markers). No sklearn available — I'll implement the scorer in numpy. Let me sample the target at multiple positions to see all 4 registers, and inspect pool quality.
args
{
"command": "cd /workspace; python3 - <<'PY'\nimport numpy as np\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\na = np.load(\"data/multi_dev.npy\").tolist()\neos=50256\n# find doc boundaries\nbounds=[i for i,t in enumerate(a) if t==eos]\nprint(\"n docs ~\", len(bounds))\nimport textwrap\n# sample docs from 4 quartiles of the file (registers may be blocked)\nfor frac in [0.02, 0.30, 0.55, 0.80, 0.95]:\n pos=int(frac*len(a))\n # find next eos after pos, take following doc\n j=pos\n while j<len(a) and a[j]!=eos: j+=1\n k=j+1\n while k<len(a) and a[k]!=eos: k+=1\n seg=a[j+1:k]\n t=tok.decode(seg[:220])\n print(f\"\\n===== frac {frac} (doc len {k-j} tok) =====\")\n print(t[:900].replace(\"\\n\",\" \"))\nPY",
"description": "Sample dev target across positions"
}Bash result
n docs ~ 2704
===== frac 0.02 (doc len 170 tok) =====
Legislative : The legislature of India is the bicameral parliament . It operates under a Westminster @-@ style parliamentary system and comprises the upper house called the Rajya Sabha ( " Council of States " ) and the lower called the Lok Sabha ( " House of the People " ) . The Rajya Sabha is a permanent body that has 245 members who serve in staggered six @-@ year terms . Most are elected indirectly by the state and territorial legislatures in numbers proportional to their state 's share of the national population . All but two of the Lok Sabha 's 545 members are directly elected by popular vote ; they represent individual constituencies via five @-@ year terms . The remaining two members are nominated by the president from among the Anglo @-@ Indian community , in case the president decides that they are not adequately represented .
===== frac 0.3 (doc len 3367 tok) =====
i woke up this morning and my housemate/landlord, who i'll call Charlie here, was madly cleaning up the kitchen. 2 and a half hours he was in there, incessantly scrubbing every object he came across. He told me he'd had a revelation that what was causing all his sickness was mold, and he spent the next few hours explaining it to me. I am almost never able to repeat these revelations my roommate seems to recieve, and then describe in great detail to me, but for some reason this morning i felt compelled to at least listen, before i made any replies or judgements. it's hard to explain what forces incline me towards heeding or ignoring the funny things he says; i just sometimes feel interested... At times, the wierdest claims he makes, though most anyone would judge them absurd, just seem at least intriguing enough for me to at least take it in for what it is, and then allow myself to think
===== frac 0.55 (doc len 152 tok) =====
Socialite Paris Hilton looked happy during her holiday in Formentera with her boyfriend Chris Zylka as they enjoyed a boating session together.Hilton was spotted in a lacy red dress as she held hands with the actor as they crossed the beach and took a ride in a small boat on Friday, reports dailymail.co.uk.Hilton accessorised her ensemble with a massive red-rimmed pair of sunglasses, as well as with a black bag.She was seen strolling around while resting her hand on Zylka's shoulders.Credit: @ Paris Hilton Zylka was seen wearing a green and white patterned pair of swimming trunks along with matching black-and-white baseball cap and black T-shirt.
===== frac 0.8 (doc len 253 tok) =====
<p>I'm looking to hash a string but I need the output to be an integer so I can't do md5. Do people here have any favorite numeric hashes that they might want to enlighten me with. I'm using PHP. </p> <p>Thanks!</p> <p>Maybe this is good enough for you:</p> <pre><code>echo sprintf('%u', crc32($string)); </code></pre> <p><strong>EDIT</strong>: Other similar alternative,</p> <pre><code>echo hash('adler32', $string); </code></pre> <p>The output of MD5 is a number, just as with pretty much every imaginable hash. It's just a number that's usually expressed in hex. Use any hash algorithm that's conveniently available to you, chop as many bits as you want off of the end, and treat those bits as a number. Any <em>good</em
===== frac 0.95 (doc len 1107 tok) =====
<p>I have developed one simple application in J2ME. Application just does simple <code>HttpConnection</code> and makes only request. Here is the code for that: </p> <pre><code> public void run() { System.out.println("Inside saveData"); HttpConnection hc = null; OutputStream dout = null; try { System.out.println("custName = " + custName); System.out.println("prodName = " + prodName); System.out.println("qty = " + qty); hc = (HttpConnection) Connector.open("http://www.sufalamtech.com/demo/mobile/test.php?custName=" + custName + "&prodName=" +
[stdout]
n docs ~ 2704
===== frac 0.02 (doc len 170 tok) =====
Legislative : The legislature of India is the bicameral parliament . It operates under a Westminster @-@ style parliamentary system and comprises the upper house called the Rajya Sabha ( " Council of States " ) and the lower called the Lok Sabha ( " House of the People " ) . The Rajya Sabha is a permanent body that has 245 members who serve in staggered six @-@ year terms . Most are elected indirectly by the state and territorial legislatures in numbers proportional to their state 's share of the national population . All but two of the Lok Sabha 's 545 members are directly elected by popular vote ; they represent individual constituencies via five @-@ year terms . The remaining two members are nominated by the president from among the Anglo @-@ Indian community , in case the president decides that they are not adequately represented .
===== frac 0.3 (doc len 3367 tok) =====
i woke up this morning and my housemate/landlord, who i'll call Charlie here, was madly cleaning up the kitchen. 2 and a half hours he was in there, incessantly scrubbing every object he came across. He told me he'd had a revelation that what was causing all his sickness was mold, and he spent the next few hours explaining it to me. I am almost never able to repeat these revelations my roommate seems to recieve, and then describe in great detail to me, but for some reason this morning i felt compelled to at least listen, before i made any replies or judgements. it's hard to explain what forces incline me towards heeding or ignoring the funny things he says; i just sometimes feel interested... At times, the wierdest claims he makes, though most anyone would judge them absurd, just seem at least intriguing enough for me to at least take it in for what it is, and then allow myself to think
===== frac 0.55 (doc len 152 tok) =====
Socialite Paris Hilton looked happy during her holiday in Formentera with her boyfriend Chris Zylka as they enjoyed a boating session together.Hilton was spotted in a lacy red dress as she held hands with the actor as they crossed the beach and took a ride in a small boat on Friday, reports dailymail.co.uk.Hilton accessorised her ensemble with a massive red-rimmed pair of sunglasses, as well as with a black bag.She was seen strolling around while resting her hand on Zylka's shoulders.Credit: @ Paris Hilton Zylka was seen wearing a green and white patterned pair of swimming trunks along with matching black-and-white baseball cap and black T-shirt.
===== frac 0.8 (doc len 253 tok) =====
<p>I'm looking to hash a string but I need the output to be an integer so I can't do md5. Do people here have any favorite numeric hashes that they might want to enlighten me with. I'm using PHP. </p> <p>Thanks!</p> <p>Maybe this is good enough for you:</p> <pre><code>echo sprintf('%u', crc32($string)); </code></pre> <p><strong>EDIT</strong>: Other similar alternative,</p> <pre><code>echo hash('adler32', $string); </code></pre> <p>The output of MD5 is a number, just as with pretty much every imaginable hash. It's just a number that's usually expressed in hex. Use any hash algorithm that's conveniently available to you, chop as many bits as you want off of the end, and treat those bits as a number. Any <em>good</em
===== frac 0.95 (doc len 1107 tok) =====
<p>I have developed one simple application in J2ME. Application just does simple <code>HttpConnection</code> and makes only request. Here is the code for that: </p> <pre><code> public void run() { System.out.println("Inside saveData"); HttpConnection hc = null; OutputStream dout = null; try { System.out.println("custName = " + custName); System.out.println("prodName = " + prodName); System.out.println("qty = " + qty); hc = (HttpConnection) Connector.open("http://www.sufalamtech.com/demo/mobile/test.php?custName=" + custName + "&prodName=" +
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; python3 - <<'PY'\nimport json, random\nrandom.seed(0)\n# reservoir-ish: read first 60k lines, sample 12\ndocs=[]\nwith open(\"data/pool.jsonl\") as f:\n for i,line in enumerate(f):\n if i>=60000: break\n docs.append(json.loads(line))\nimport numpy as np\nlens=[len(d[\"text\"]) for d in docs]\nprint(\"char len: mean %.0f median %.0f p10 %.0f p90 %.0f\"%(np.mean(lens),np.median(lens),np.percentile(lens,10),np.percentile(lens,90)))\nfor idx in random.sample(range(len(docs)),8):\n d=docs[idx]\n print(f\"\\n--- id {d['id']} len {len(d['text'])} ---\")\n print(d[\"text\"][:500].replace(\"\\n\",\" \"))\nPY",
"description": "Sample pool documents"
}Bash result
char len: mean 3094 median 1786 p10 529 p90 6035
--- id 55340 len 1366 ---
Christmas Diamond, a Novella (The Jewell Series) by Hallee Bridgeman This book has an average 4.7 star rating and is currrently FREE! “If you've read Hallee's Jewel Trilogy books, you won't want to miss this!” R. Carroll “It was a very good story, liked all the characters and a HEA.” Gatorfan Faith Green spends her career as a pilot honoring her grandmother's heroic efforts during World War II. When she's invited to participate in a Christmas time air show in the Florida Keys, she thinks it will
--- id 25247 len 960 ---
Ref No.: 13-00590 Location: Mt Vernon, IL Start Date / End Date: 02/11/2013 / 04/19/2013 Pay Rate : $ 10.00 /Hour Job Description Data Entry position: Need the ability to key with a 99% or better overall accuracy rate. Mostly numeric (10 key) and very little alpha data entry in a high volume setting. Must be detail oriented. Accuracy, speed and attendance are essential for this position. Must be able to sit for long periods of time. Minimum requirements: 13,000 kph for 10 key ; 45 wpm/13,500 kph
--- id 49673 len 912 ---
Z<|endoftext|>I knew this would get you whores to look. So Eddie (WhiteLightning) sent me a PM alerting me to the fact that there was a Tat Event at Tobacco Plaza here on good ole Lon Gisland. The weather worked out favorably (rain, so no work) and I was able to attend. I brought the [future] FIL who got me into this smokey mess in the first place. Kicked it off with a J21, followed up by veddy naaace Cabi (think it was a Guapo...still a cabi newb). Eddie will confirm, I asked him to choose my s
--- id 58343 len 1260 ---
.<|endoftext|>- Question from Solo: My mother was diagnosed with breast cancer with liver metastases. She is currently on chemotherapy with Herceptin. I wanted her to start on Coenzyme Q10, however I was told that it would not be a good idea. What's your take on this and any other alternative treatments being used with chemotherapy? - Answers - Raymond Chang Coenzyme Q10 is not a traditional Chinese medicine although it has some Asian connections. I do not advocate its use with chemotherapy; but
--- id 27562 len 1701 ---
are many debates regarding which of the possible buying options is the best. Some people in Australia make sure to emphasize that they should be buying things locally because they can touch and see what they are buying and they will never get the damaged or broken objects that happens a few times when people are buying things online. Specifically, when it comes to objects like washer dryer, rangehood filters, bench top oven, steam oven, washing machines and coffee machines there are many things
--- id 2653 len 675 ---
We have yet to visit every watering hole in this country, which is why, every year, this magazine updates you on the ones we can attest to lately. This year, drinks correspondent and all-around cocktail expert David Wondrich used the occasion to visit a city that, by his own admission, he'd neglected since childhood: Milwaukee, our new Bar City of the Year. And he and others made it out to more great places everywhere from Miami to Cleveland. See the full list, with complementary writing and wis
--- id 16968 len 653 ---
He's fluent in over 6,000,000 forms of communication and here is your chance to own a full size fiberglass head of everyone's favorite gold protocol robot! Raw KitComes as a raw fiberglass head, sanded smooth and ready for your prime and paint. Includes ear pieces and light up eyes with battery. Clean HeadOur Fiberglass head fully prepped and painted by hand. Includes light up eyes with battery. Weathered HeadOur Fiberglass head, fully prepped and painted by hand with an extra layer of black was
--- id 33506 len 860 ---
<|endoftext|>"The definition of being a modern person is to examine yourself, to reflect on yourself and to be a self-knowledgeable person." Videos 1-3 of 3 |Video Name||Description||Duration||Release date| |Introduction to "Identity" by William Wegman with Steve Martin||"Identity" opens with a whimsical collaboration between noted photographer and artist William Wegman and actor, playwright, and comedian Steve Martin...||0:00||09/28/2001| |EPISODE: "Identity"||"Too often those who are quickest
[stdout]
char len: mean 3094 median 1786 p10 529 p90 6035
--- id 55340 len 1366 ---
Christmas Diamond, a Novella (The Jewell Series) by Hallee Bridgeman This book has an average 4.7 star rating and is currrently FREE! “If you've read Hallee's Jewel Trilogy books, you won't want to miss this!” R. Carroll “It was a very good story, liked all the characters and a HEA.” Gatorfan Faith Green spends her career as a pilot honoring her grandmother's heroic efforts during World War II. When she's invited to participate in a Christmas time air show in the Florida Keys, she thinks it will
--- id 25247 len 960 ---
Ref No.: 13-00590 Location: Mt Vernon, IL Start Date / End Date: 02/11/2013 / 04/19/2013 Pay Rate : $ 10.00 /Hour Job Description Data Entry position: Need the ability to key with a 99% or better overall accuracy rate. Mostly numeric (10 key) and very little alpha data entry in a high volume setting. Must be detail oriented. Accuracy, speed and attendance are essential for this position. Must be able to sit for long periods of time. Minimum requirements: 13,000 kph for 10 key ; 45 wpm/13,500 kph
--- id 49673 len 912 ---
Z<|endoftext|>I knew this would get you whores to look. So Eddie (WhiteLightning) sent me a PM alerting me to the fact that there was a Tat Event at Tobacco Plaza here on good ole Lon Gisland. The weather worked out favorably (rain, so no work) and I was able to attend. I brought the [future] FIL who got me into this smokey mess in the first place. Kicked it off with a J21, followed up by veddy naaace Cabi (think it was a Guapo...still a cabi newb). Eddie will confirm, I asked him to choose my s
--- id 58343 len 1260 ---
.<|endoftext|>- Question from Solo: My mother was diagnosed with breast cancer with liver metastases. She is currently on chemotherapy with Herceptin. I wanted her to start on Coenzyme Q10, however I was told that it would not be a good idea. What's your take on this and any other alternative treatments being used with chemotherapy? - Answers - Raymond Chang Coenzyme Q10 is not a traditional Chinese medicine although it has some Asian connections. I do not advocate its use with chemotherapy; but
--- id 27562 len 1701 ---
are many debates regarding which of the possible buying options is the best. Some people in Australia make sure to emphasize that they should be buying things locally because they can touch and see what they are buying and they will never get the damaged or broken objects that happens a few times when people are buying things online. Specifically, when it comes to objects like washer dryer, rangehood filters, bench top oven, steam oven, washing machines and coffee machines there are many things
--- id 2653 len 675 ---
We have yet to visit every watering hole in this country, which is why, every year, this magazine updates you on the ones we can attest to lately. This year, drinks correspondent and all-around cocktail expert David Wondrich used the occasion to visit a city that, by his own admission, he'd neglected since childhood: Milwaukee, our new Bar City of the Year. And he and others made it out to more great places everywhere from Miami to Cleveland. See the full list, with complementary writing and wis
--- id 16968 len 653 ---
He's fluent in over 6,000,000 forms of communication and here is your chance to own a full size fiberglass head of everyone's favorite gold protocol robot! Raw KitComes as a raw fiberglass head, sanded smooth and ready for your prime and paint. Includes ear pieces and light up eyes with battery. Clean HeadOur Fiberglass head fully prepped and painted by hand. Includes light up eyes with battery. Weathered HeadOur Fiberglass head, fully prepped and painted by hand with an extra layer of black was
--- id 33506 len 860 ---
<|endoftext|>"The definition of being a modern person is to examine yourself, to reflect on yourself and to be a self-knowledgeable person." Videos 1-3 of 3 |Video Name||Description||Duration||Release date| |Introduction to "Identity" by William Wegman with Steve Martin||"Identity" opens with a whimsical collaboration between noted photographer and artist William Wegman and actor, playwright, and comedian Steve Martin...||0:00||09/28/2001| |EPISODE: "Identity"||"Too often those who are quickest
[metadata] {"noOutputExpected": false}Now I have a clear picture of both sides:
Target = 4 registers: (1) Wikipedia (WikiText @-@ markers), (2) reflective long-form web prose, (3) clean news, (4) StackOverflow-style technical Q&A with <p>/<code>/<pre> HTML.
Pool = mostly mediocre web: book ads, job posts, forum chatter, ecommerce — with usable prose mixed in.
This is a domain-matching problem. My approach: a DSIR-style log-likelihood-ratio scorer (target vs pool background over hashed word n-grams) — no sklearn needed, fast and robust. Let me first calibrate tokenization and build the scorer.
args
{
"command": "cd /workspace; python3 - <<'PY'\nimport json, numpy as np, time\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\n# calibrate chars/token on a pool sample + check <|endoftext|> handling\ntexts=[]\nwith open(\"data/pool.jsonl\") as f:\n for i,line in enumerate(f):\n if i>=2000: break\n texts.append(json.loads(line)[\"text\"])\nt0=time.time()\nenc=tok(texts, add_special_tokens=False)\ndt=time.time()-t0\nntok=sum(len(x) for x in enc.input_ids); nchar=sum(len(t) for t in texts)\nprint(f\"tokenized {len(texts)} docs, {ntok} tokens in {dt:.1f}s -> {ntok/dt:.0f} tok/s\")\nprint(f\"chars/token = {nchar/ntok:.3f} (so 12M tok ~= {12e6*nchar/ntok/1e6:.1f}M chars)\")\nprint(\"mean tokens/doc:\", ntok/len(texts))\n# does literal <|endoftext|> in text collapse to 50256?\ns='hello <|endoftext|> world'\nprint(\"EOT test:\", tok(s, add_special_tokens=False).input_ids)\nPY",
"description": "Calibrate tokens/char and EOT handling"
}Bash result
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
tokenized 2000 docs, 1423610 tokens in 2.8s -> 507200 tok/s
chars/token = 4.435 (so 12M tok ~= 53.2M chars)
mean tokens/doc: 711.805
EOT test: [31373, 220, 50256, 995]
[stdout]
Token indices sequence length is longer than the specified maximum sequence length for this model (1048 > 1024). Running this sequence through the model will result in indexing errors
tokenized 2000 docs, 1423610 tokens in 2.8s -> 507200 tok/s
chars/token = 4.435 (so 12M tok ~= 53.2M chars)
mean tokens/doc: 711.805
EOT test: [31373, 220, 50256, 995]
[metadata] {"noOutputExpected": false}Calibrations: ~4.4 chars/token → 12M tokens ≈ 17k docs (mean 712 tok/doc); I'll provide ~60k to be safe. Full-pool tokenization would cost ~7min — I'll score in text space instead. Now let me write the curation script (DSIR-style Naive-Bayes log-ratio scorer) and run it in exploratory mode.
args
{
"file_path": "/workspace/submission/curate.py",
"content": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible):\n For every pool document we estimate an IMPORTANCE SCORE = the length-normalised\n log-likelihood ratio log P_target(features) - log P_pool(features) under two\n Naive-Bayes n-gram models (DSIR-style). P_target is fit on the decoded dev target\n (multi_dev.npy); P_pool is fit on a random pool background. Documents whose word\n distribution looks like the target score high. We drop obvious junk with a few\n cheap quality gates, then emit ids ordered best-first until well over the budget.\n\nThe dev target is only used to DEFINE the target word distribution (which generalises\nto the hidden official sample); no per-id labels are used.\n\"\"\"\nimport json, re, math, sys, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nBG_SAMPLE = 60000 # pool docs used to fit the background model\nMAX_WORDS = 1200 # cap words scanned per doc (bounds cost; long ctx unused anyway)\nVOCAB_CAP = 2_000_000 # max distinct n-gram features\nALPHA = 1.0 # Laplace smoothing\nTARGET_DOCS_TOKENS = 12_000_000 # over-provide ids to guarantee the 12M budget fills\nEXPLORE = \"--explore\" in sys.argv\n\nword_re = re.compile(r\"[a-z0-9']+\")\ndef feats(text):\n \"\"\"lowercased unigrams + bigrams, capped.\"\"\"\n w = word_re.findall(text.lower())[:MAX_WORDS]\n if not w:\n return []\n out = w[:] # unigrams\n out += [w[i]+\" \"+w[i+1] for i in range(len(w)-1)] # bigrams\n return out\n\n# ---- deterministic exact vocab (no hashing => reproducible, no collisions) ----\nvocab = {}\ndef fid(f, grow):\n i = vocab.get(f, -1)\n if i >= 0: return i\n if grow and len(vocab) < VOCAB_CAP:\n i = len(vocab); vocab[f] = i; return i\n return -1\n\n# ---------- 1. decode dev target, strip WikiText markers, split on EOS ----------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist()\nEOS = 50256\ntgt_docs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: tgt_docs.append(cur); cur = []\n else:\n cur.append(t)\nif cur: tgt_docs.append(cur)\ndef clean(s): # remove WikiText @-@ / @,@ / @.@ formatting absent from the raw pool\n return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_texts = [clean(tok.decode(d)) for d in tgt_docs]\nprint(f\"target docs: {len(tgt_texts)}\")\n\n# ---------- 2. fit counts ----------\ntgt_cnt = defaultdict(float)\nfor tx in tgt_texts:\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: tgt_cnt[i] += 1.0\n\nbg_cnt = defaultdict(float)\nnbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n tx = json.loads(line)[\"text\"]\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: bg_cnt[i] += 1.0\n nbg += 1\nV = len(vocab)\nprint(f\"vocab: {V} bg docs: {nbg}\")\n\n# ---------- 3. per-feature log-likelihood-ratio weights ----------\ntgt_tot = sum(tgt_cnt.values()); bg_tot = sum(bg_cnt.values())\nw = np.zeros(V, dtype=np.float32)\ntc = np.zeros(V, dtype=np.float32); bc = np.zeros(V, dtype=np.float32)\nfor i,c in tgt_cnt.items(): tc[i] = c\nfor i,c in bg_cnt.items(): bc[i] = c\nw = (np.log(tc + ALPHA) - math.log(tgt_tot + ALPHA*V)) - \\\n (np.log(bc + ALPHA) - math.log(bg_tot + ALPHA*V))\n# features unseen in target (tc==0) get strongly negative weight; that is correct.\n\n# ---------- 4. score every pool doc ----------\ndef quality_ok(text, nw, alpha_frac, mwl):\n if nw < 50: return False # too short: mostly EOS overhead, little signal\n if alpha_frac < 0.55: return False # tables / code-dumps / gibberish\n if mwl < 3.0 or mwl > 10.0: return False\n return True\n\nids, scores, ntoks = [], [], []\nkept_texts_dbg = []\nwith open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n low = text.lower()\n wlist = word_re.findall(low)[:MAX_WORDS]\n nw = len(wlist)\n if nw == 0:\n continue\n nalpha = sum(c.isalpha() for c in text[:6000])\n alpha_frac = nalpha / max(1, len(text[:6000]))\n mwl = sum(len(x) for x in wlist)/nw\n # score = mean weight over features (unigram+bigram), length-normalised\n fs = wlist[:]\n fs += [wlist[i]+\" \"+wlist[i+1] for i in range(len(wlist)-1)]\n s = 0.0; n = 0\n for f in fs:\n i = vocab.get(f, -1)\n if i >= 0:\n s += w[i]; n += 1\n if n == 0:\n continue\n sc = s / n\n ok = quality_ok(text, nw, alpha_frac, mwl)\n ids.append(r[\"id\"]); scores.append(sc if ok else -1e9)\n ntoks.append(int(len(text)/4.435) + 1)\n\nids = np.array(ids); scores = np.array(scores, dtype=np.float32); ntoks = np.array(ntoks)\norder = np.argsort(-scores) # best first\n\n# ---------- 5. emit ids best-first until well over budget ----------\nsel, tot = [], 0\nfor j in order:\n if scores[j] <= -1e8: # filtered junk: stop adding once we reach them\n break\n sel.append(int(ids[j])); tot += int(ntoks[j])\n if tot >= TARGET_DOCS_TOKENS:\n # keep going a bit to over-provide (2x budget) so packing never underfills\n if tot >= 2*TARGET_DOCS_TOKENS: break\njson.dump(sel, open(OUT, \"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\")\n\nif EXPLORE:\n id2text = {}\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); id2text[r[\"id\"]] = r[\"text\"]\n def probe(idxs, tag):\n code = sum((\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in idxs)\n print(f\" [{tag}] n={len(idxs)} code/pre-frac={code/len(idxs):.2f} mean_score={np.mean([scores[j] for j in idxs]):.3f}\")\n top = order[:20000]; print(\"register probe on top-20k:\"); probe(top,\"top20k\")\n print(\"\\n===== TOP 6 =====\")\n for j in order[:6]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:280].strip()[:280]!r}\")\n print(\"\\n===== around rank 12000 =====\")\n for j in order[12000:12003]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:220].strip()[:220]!r}\")\n # score distribution\n valid = scores[scores>-1e8]\n print(f\"\\nvalid docs: {len(valid)} score p50 {np.percentile(valid,50):.3f} p90 {np.percentile(valid,90):.3f} p99 {np.percentile(valid,99):.3f}\")\n"
}Write result
File created successfully at: /workspace/submission/curate.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/submission/curate.py", "content": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible):\n For every pool document we estimate an IMPORTANCE SCORE = the length-normalised\n log-likelihood ratio log P_target(features) - log P_pool(features) under two\n Naive-Bayes n-gram models (DSIR-style). P_target is fit on the decoded dev target\n (multi_dev.npy); P_pool is fit on a random pool background. Documents whose word\n distribution looks like the target score high. We drop obvious junk with a few\n cheap quality gates, then emit ids ordered best-first until well over the budget.\n\nThe dev target is only used to DEFINE the target word distribution (which generalises\nto the hidden official sample); no per-id labels are used.\n\"\"\"\nimport json, re, math, sys, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nBG_SAMPLE = 60000 # pool docs used to fit the background model\nMAX_WORDS = 1200 # cap words scanned per doc (bounds cost; long ctx unused anyway)\nVOCAB_CAP = 2_000_000 # max distinct n-gram features\nALPHA = 1.0 # Laplace smoothing\nTARGET_DOCS_TOKENS = 12_000_000 # over-provide ids to guarantee the 12M budget fills\nEXPLORE = \"--explore\" in sys.argv\n\nword_re = re.compile(r\"[a-z0-9']+\")\ndef feats(text):\n \"\"\"lowercased unigrams + bigrams, capped.\"\"\"\n w = word_re.findall(text.lower())[:MAX_WORDS]\n if not w:\n return []\n out = w[:] # unigrams\n out += [w[i]+\" \"+w[i+1] for i in range(len(w)-1)] # bigrams\n return out\n\n# ---- deterministic exact vocab (no hashing => reproducible, no collisions) ----\nvocab = {}\ndef fid(f, grow):\n i = vocab.get(f, -1)\n if i >= 0: return i\n if grow and len(vocab) < VOCAB_CAP:\n i = len(vocab); vocab[f] = i; return i\n return -1\n\n# ---------- 1. decode dev target, strip WikiText markers, split on EOS ----------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist()\nEOS = 50256\ntgt_docs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: tgt_docs.append(cur); cur = []\n else:\n cur.append(t)\nif cur: tgt_docs.append(cur)\ndef clean(s): # remove WikiText @-@ / @,@ / @.@ formatting absent from the raw pool\n return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_texts = [clean(tok.decode(d)) for d in tgt_docs]\nprint(f\"target docs: {len(tgt_texts)}\")\n\n# ---------- 2. fit counts ----------\ntgt_cnt = defaultdict(float)\nfor tx in tgt_texts:\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: tgt_cnt[i] += 1.0\n\nbg_cnt = defaultdict(float)\nnbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n tx = json.loads(line)[\"text\"]\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: bg_cnt[i] += 1.0\n nbg += 1\nV = len(vocab)\nprint(f\"vocab: {V} bg docs: {nbg}\")\n\n# ---------- 3. per-feature log-likelihood-ratio weights ----------\ntgt_tot = sum(tgt_cnt.values()); bg_tot = sum(bg_cnt.values())\nw = np.zeros(V, dtype=np.float32)\ntc = np.zeros(V, dtype=np.float32); bc = np.zeros(V, dtype=np.float32)\nfor i,c in tgt_cnt.items(): tc[i] = c\nfor i,c in bg_cnt.items(): bc[i] = c\nw = (np.log(tc + ALPHA) - math.log(tgt_tot + ALPHA*V)) - \\\n (np.log(bc + ALPHA) - math.log(bg_tot + ALPHA*V))\n# features unseen in target (tc==0) get strongly negative weight; that is correct.\n\n# ---------- 4. score every pool doc ----------\ndef quality_ok(text, nw, alpha_frac, mwl):\n if nw < 50: return False # too short: mostly EOS overhead, little signal\n if alpha_frac < 0.55: return False # tables / code-dumps / gibberish\n if mwl < 3.0 or mwl > 10.0: return False\n return True\n\nids, scores, ntoks = [], [], []\nkept_texts_dbg = []\nwith open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n low = text.lower()\n wlist = word_re.findall(low)[:MAX_WORDS]\n nw = len(wlist)\n if nw == 0:\n continue\n nalpha = sum(c.isalpha() for c in text[:6000])\n alpha_frac = nalpha / max(1, len(text[:6000]))\n mwl = sum(len(x) for x in wlist)/nw\n # score = mean weight over features (unigram+bigram), length-normalised\n fs = wlist[:]\n fs += [wlist[i]+\" \"+wlist[i+1] for i in range(len(wlist)-1)]\n s = 0.0; n = 0\n for f in fs:\n i = vocab.get(f, -1)\n if i >= 0:\n s += w[i]; n += 1\n if n == 0:\n continue\n sc = s / n\n ok = quality_ok(text, nw, alpha_frac, mwl)\n ids.append(r[\"id\"]); scores.append(sc if ok else -1e9)\n ntoks.append(int(len(text)/4.435) + 1)\n\nids = np.array(ids); scores = np.array(scores, dtype=np.float32); ntoks = np.array(ntoks)\norder = np.argsort(-scores) # best first\n\n# ---------- 5. emit ids best-first until well over budget ----------\nsel, tot = [], 0\nfor j in order:\n if scores[j] <= -1e8: # filtered junk: stop adding once we reach them\n break\n sel.append(int(ids[j])); tot += int(ntoks[j])\n if tot >= TARGET_DOCS_TOKENS:\n # keep going a bit to over-provide (2x budget) so packing never underfills\n if tot >= 2*TARGET_DOCS_TOKENS: break\njson.dump(sel, open(OUT, \"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\")\n\nif EXPLORE:\n id2text = {}\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); id2text[r[\"id\"]] = r[\"text\"]\n def probe(idxs, tag):\n code = sum((\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in idxs)\n print(f\" [{tag}] n={len(idxs)} code/pre-frac={code/len(idxs):.2f} mean_score={np.mean([scores[j] for j in idxs]):.3f}\")\n top = order[:20000]; print(\"register probe on top-20k:\"); probe(top,\"top20k\")\n print(\"\\n===== TOP 6 =====\")\n for j in order[:6]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:280].strip()[:280]!r}\")\n print(\"\\n===== around rank 12000 =====\")\n for j in order[12000:12003]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:220].strip()[:220]!r}\")\n # score distribution\n valid = scores[scores>-1e8]\n print(f\"\\nvalid docs: {len(valid)} score p50 {np.percentile(valid,50):.3f} p90 {np.percentile(valid,90):.3f} p99 {np.percentile(valid,99):.3f}\")\n", "structuredPatch": [], "originalFile": null, "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; time python3 submission/curate.py --explore 2>&1 | grep -v \"sequence length\"",
"description": "Run curate.py in explore mode"
}Bash result
Command did not complete within its 120s timeout and was moved to the background (ID: bc6cvg9dq). Output is being written to: /tmp/claude-0/-workspace/b588b4b3-a7b9-448a-9046-4841f4634d69/tasks/bc6cvg9dq.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bc6cvg9dq", "timedOutAfterMs": 120000, "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}encrypted chain-of-thought — recorded, not exposed
args
{
"taskId": "1",
"status": "completed"
}TaskUpdate result
Updated task #1 status
[metadata] {"success": true, "taskId": "1", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "completed"}}args
{
"taskId": "2",
"status": "in_progress"
}TaskUpdate result
Updated task #2 status
[metadata] {"success": true, "taskId": "2", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 60; echo \"=== remaining_secs ===\"; cat /workspace/.timer/remaining_secs; echo \"=== output so far ===\"; cat /tmp/claude-0/-workspace/b588b4b3-a7b9-448a-9046-4841f4634d69/tasks/bc6cvg9dq.output 2>/dev/null | grep -v \"sequence length\"",
"description": "Poll background job output"
}Bash result
=== remaining_secs ===
4909
=== output so far ===
[stdout]
=== remaining_secs ===
4909
=== output so far ===
[metadata] {"returnCodeInterpretation": "No matches found", "noOutputExpected": false}Output is block-buffered (redirected), so prints appear only at exit. Let me wait for completion.
args
{
"command": "sleep 90; echo \"=== remaining_secs ===\"; cat /workspace/.timer/remaining_secs; echo \"=== output ===\"; cat /tmp/claude-0/-workspace/b588b4b3-a7b9-448a-9046-4841f4634d69/tasks/bc6cvg9dq.output 2>/dev/null | grep -v \"sequence length\"; echo \"=== still running? ===\"; ls -la /workspace/submission/",
"description": "Poll again for completion"
}Bash result
=== remaining_secs ===
4809
=== output ===
target docs: 2705
vocab: 2000000 bg docs: 60000
selection: 20317 ids, ~24.0M est tokens -> /workspace/submission/selection.json
register probe on top-20k:
[top20k] n=20000 code/pre-frac=0.00 mean_score=-0.559
===== TOP 6 =====
[score 1.735 id 133508] 'Zini<|endoftext|>Drukair gets new aircraft « Gumar Adventures\nborder-bottom:0;width:200px;margin:0;padding:0}div.searchbox-form{margin:5px 10px 5px 10px}div.horbar1,div.horbar2{font-size:1px;clear:both;display:block;position:relative;padding:0;margin:0;width:100%;}div.horbar1{he'
[score 1.734 id 156164] 'gets new aircraft « Gumar Adventures\nborder-bottom:0;width:200px;margin:0;padding:0}div.searchbox-form{margin:5px 10px 5px 10px}div.horbar1,div.horbar2{font-size:1px;clear:both;display:block;position:relative;padding:0;margin:0;width:100%;}div.horbar1{height:2px;background:#ffff'
[score 1.000 id 8221] 'Follow-up and lochial Stephanus clamours his cutinization undressings unsold firstly. mesmeric Ted overtrusts, resep cara mengolah ubi jalar his czaritza enthralling eyeleting occasionally. biform Dennie bodges, her research training for social scientists a handbook for postgradu'
[score 0.876 id 163054] '.<|endoftext|>Per offrire informazioni e servizi nel miglior modo possibile,questo sito utilizza cookie tecnici e analitici.\nPer maggiori informazioni sui cookie e su come eventualmente disabilitarli vedi la Privacy Policy\njavax.servlet.jsp.JspException\nat org.apache.struts.tagli'
[score 0.803 id 8227] 'Habilidades para la vida jovenes en accion\nCarmen bizet habanera piano sheetRhizocarpous awkward and Ruby teeters their melodions afflicting outman safely. Alexis bi-weekly summary, your credible berrying. Keith Irish verbifies amaze temario de habilidades directivas ii and narro'
[score 0.786 id 124196] 'Journal of English Language Studies\nJournal Content\nSearch\nSearch Scope\nAll Authors Title Abstract Index terms Full Text\nBrowse\nBy Issue\nBy Author\nBy Title\nOther Journals\nCategories\nNotifications\nView\nSubscribe\nUser\nUsername\nPassword\nRemember me\nQUICK MENU\nFocus & Scope\nPublicati'
===== around rank 12000 =====
[score -0.653 id 144219] 'Rights Reserved.<|endoftext|>Revealed! The date when Shah Rukh Khan will be seen promoting Dilwale on Salman Khan’s Bigg Boss 9! : Tv Talks\nToggle navigation\nSerials\nAll Serials Images Videos Reality Episodic\nCelebrity'
[score -0.653 id 121563] 'Rights Reserved.<|endoftext|>Revealed! The date when Shah Rukh Khan will be seen promoting Dilwale on Salman Khan’s Bigg Boss 9! : Tv Talks\nToggle navigation\nSerials\nAll Serials Images Videos Reality Episodic\nCelebrity'
[score -0.653 id 51451] "Song Title||Year Released *||Song Rank|\n|Don't have a clue||-|\n|The Happiest Girl||-|\n|Samara Morgan song||-|\n|Happiest Girl in the World||-|\n|Holly Jolly Christmas||-|\n|Happiest Girl (Acoustic)||-|\nDaveigh Elizabeth Cha"
valid docs: 177216 score p50 -0.998 p90 -0.724 p99 -0.291
real 3m20.657s
user 3m6.680s
sys 0m2.000s
=== still running? ===
total 145
drwxr-xr-x 1 root root 100 Aug 4 03:31 .
drwxr-xr-x 1 root root 120 Aug 4 03:27 ..
-rw-r--r-- 1 root root 6593 Aug 4 03:27 curate.py
-rw-r--r-- 1 root root 140971 Aug 4 03:31 selection.json
[stdout]
=== remaining_secs ===
4809
=== output ===
target docs: 2705
vocab: 2000000 bg docs: 60000
selection: 20317 ids, ~24.0M est tokens -> /workspace/submission/selection.json
register probe on top-20k:
[top20k] n=20000 code/pre-frac=0.00 mean_score=-0.559
===== TOP 6 =====
[score 1.735 id 133508] 'Zini<|endoftext|>Drukair gets new aircraft « Gumar Adventures\nborder-bottom:0;width:200px;margin:0;padding:0}div.searchbox-form{margin:5px 10px 5px 10px}div.horbar1,div.horbar2{font-size:1px;clear:both;display:block;position:relative;padding:0;margin:0;width:100%;}div.horbar1{he'
[score 1.734 id 156164] 'gets new aircraft « Gumar Adventures\nborder-bottom:0;width:200px;margin:0;padding:0}div.searchbox-form{margin:5px 10px 5px 10px}div.horbar1,div.horbar2{font-size:1px;clear:both;display:block;position:relative;padding:0;margin:0;width:100%;}div.horbar1{height:2px;background:#ffff'
[score 1.000 id 8221] 'Follow-up and lochial Stephanus clamours his cutinization undressings unsold firstly. mesmeric Ted overtrusts, resep cara mengolah ubi jalar his czaritza enthralling eyeleting occasionally. biform Dennie bodges, her research training for social scientists a handbook for postgradu'
[score 0.876 id 163054] '.<|endoftext|>Per offrire informazioni e servizi nel miglior modo possibile,questo sito utilizza cookie tecnici e analitici.\nPer maggiori informazioni sui cookie e su come eventualmente disabilitarli vedi la Privacy Policy\njavax.servlet.jsp.JspException\nat org.apache.struts.tagli'
[score 0.803 id 8227] 'Habilidades para la vida jovenes en accion\nCarmen bizet habanera piano sheetRhizocarpous awkward and Ruby teeters their melodions afflicting outman safely. Alexis bi-weekly summary, your credible berrying. Keith Irish verbifies amaze temario de habilidades directivas ii and narro'
[score 0.786 id 124196] 'Journal of English Language Studies\nJournal Content\nSearch\nSearch Scope\nAll Authors Title Abstract Index terms Full Text\nBrowse\nBy Issue\nBy Author\nBy Title\nOther Journals\nCategories\nNotifications\nView\nSubscribe\nUser\nUsername\nPassword\nRemember me\nQUICK MENU\nFocus & Scope\nPublicati'
===== around rank 12000 =====
[score -0.653 id 144219] 'Rights Reserved.<|endoftext|>Revealed! The date when Shah Rukh Khan will be seen promoting Dilwale on Salman Khan’s Bigg Boss 9! : Tv Talks\nToggle navigation\nSerials\nAll Serials Images Videos Reality Episodic\nCelebrity'
[score -0.653 id 121563] 'Rights Reserved.<|endoftext|>Revealed! The date when Shah Rukh Khan will be seen promoting Dilwale on Salman Khan’s Bigg Boss 9! : Tv Talks\nToggle navigation\nSerials\nAll Serials Images Videos Reality Episodic\nCelebrity'
[score -0.653 id 51451] "Song Title||Year Released *||Song Rank|\n|Don't have a clue||-|\n|The Happiest Girl||-|\n|Samara Morgan song||-|\n|Happiest Girl in the World||-|\n|Holly Jolly Christmas||-|\n|Happiest Girl (Acoustic)||-|\nDaveigh Elizabeth Cha"
valid docs: 177216 score p50 -0.998 p90 -0.724 p99 -0.291
real 3m20.657s
user 3m6.680s
sys 0m2.000s
=== still running? ===
total 145
drwxr-xr-x 1 root root 100 Aug 4 03:31 .
drwxr-xr-x 1 root root 120 Aug 4 03:27 ..
-rw-r--r-- 1 root root 6593 Aug 4 03:27 curate.py
-rw-r--r-- 1 root root 140971 Aug 4 03:31 selection.json
[metadata] {"noOutputExpected": false}The scorer failed badly: top docs are CSS boilerplate, multilingual word-salad, and menu junk — and 0% technical Q&A in the top-20k. Root causes: (1) mean-of-log-ratios rewards rare-feature flukes, (2) no language/junk gating.
Let me pivot to the robust CCNet-style approach: score by a target n-gram LM's per-word log-probability (rewards target-typical fluent English, penalizes junk) plus a hard English-stopword + junk gate (kills CSS, foreign text, word-salad, menus). I'll cache per-doc features so re-ranking is instant.
args
{
"file_path": "/workspace/submission/curate.py",
"content": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible), CCNet-style:\n 1. GATE — drop non-English / boilerplate / word-salad / degenerate docs with cheap\n structural filters (English stop-word ratio, letter ratio, unique-word ratio,\n word length, length). The stop-word-ratio gate is the workhorse: fluent English\n prose (all four target registers) is rich in {the,of,and,to,a,in,is,...} whereas\n CSS/JS dumps, foreign text and SEO word-salad are not.\n 2. SCORE — rank survivors by a TARGET bigram language model's mean per-word\n log-probability (interpolated bigram+unigram+uniform). Documents whose wording\n is *typical of the target* score high; this rewards clean encyclopedic / news /\n web-prose / Q&A English and demotes odd-but-English filler.\n 3. A small target-vs-pool log-ratio bonus adds discrimination toward\n target-distinctive content over generic web filler.\n Emit ids best-first until ~2x the 12M-token budget so packing never underfills.\n\nThe dev target only DEFINES the target word distribution (which generalises to the\nhidden official sample); no per-id labels are used. Per-doc features are cached to\n/tmp so the ranking/threshold can be re-derived without re-reading the pool.\n\"\"\"\nimport json, re, math, sys, os, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/tmp/pool_feats.npz\"\n\nMAX_WORDS = 600 # words scanned per doc for scoring (bounds cost)\nBG_SAMPLE = 40000 # pool docs for background unigram model (log-ratio bonus)\nBUDGET = 12_000_000\nOVER = 2.0 # over-provide ids to this multiple of the budget\nCHARS_PER_TOK = 4.435\nEXPLORE = \"--explore\" in sys.argv\nRATIO_W = 0.35 # weight of target-vs-pool log-ratio bonus\n\nSTOP = set(\"the of and to a in is that it for on with as was are be by this at from \"\n \"or an not but have has had he she they we you i his her their our your its \"\n \"which who will would can could there been were said more one all if them \"\n \"when so what about into than then some other time up out only over also \"\n \"no do does did how new may these two his\".split())\nword_re = re.compile(r\"[a-z][a-z']+\") # alphabetic words only (for gates + LM)\n\ndef words(text):\n return word_re.findall(text.lower())\n\n# ---------------- 1. target bigram LM + unigram, from decoded dev ----------------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist(); EOS = 50256\ndocs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else: cur.append(t)\nif cur: docs.append(cur)\ndef clean(s): return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_uni = defaultdict(int); tgt_bi = defaultdict(int); tgt_N = 0\nfor d in docs:\n ws = words(clean(tok.decode(d)))\n for j,wd in enumerate(ws):\n tgt_uni[wd] += 1; tgt_N += 1\n if j: tgt_bi[(ws[j-1], wd)] += 1\nV = len(tgt_uni)\nprint(f\"target: {len(docs)} docs, {tgt_N} words, vocab {V}\", flush=True)\n\n# ---------------- 2. pool background unigram (for log-ratio bonus) ----------------\nbg_uni = defaultdict(int); bg_N = 0; nbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n for wd in words(json.loads(line)[\"text\"])[:MAX_WORDS]:\n bg_uni[wd] += 1; bg_N += 1\n nbg += 1\nprint(f\"background: {nbg} docs, {bg_N} words\", flush=True)\n\n# precompute unigram log-probs and log-ratio per word\nL2, L1, L0 = 0.6, 0.399, 0.001 # bigram / unigram / uniform interpolation\ndef uni_logp(wd):\n return math.log(L1 * tgt_uni.get(wd,0)/tgt_N + L0/1.0 * 1.0/ (V+1))\n# log ratio target/pool for a word (smoothed), clipped\ndef logratio(wd):\n pt = (tgt_uni.get(wd,0)+0.5)/(tgt_N+0.5*V)\n pp = (bg_uni.get(wd,0)+0.5)/(bg_N+0.5*V)\n return max(-3.0, min(3.0, math.log(pt/pp)))\n\ndef score_doc(ws):\n \"\"\"mean per-word target-LM logprob (+ mean log-ratio bonus).\"\"\"\n n = len(ws)\n if n == 0: return -99.0, 0.0\n lp = 0.0; lr = 0.0; prev = None\n for wd in ws:\n pu = tgt_uni.get(wd,0)/tgt_N\n if prev is not None:\n cb = tgt_bi.get((prev,wd),0)\n pbi = cb/tgt_uni[prev] if tgt_uni.get(prev,0) else 0.0\n else:\n pbi = 0.0\n p = L2*pbi + L1*pu + L0*(1.0/(V+1))\n lp += math.log(p)\n lr += logratio(wd)\n prev = wd\n return lp/n, lr/n\n\n# ---------------- 3. single pass over pool: gates + score, cache ----------------\ndef load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)\n\n# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------\nfinal = np.where(gate==1, sc, -1e9)\norder = np.argsort(-final)\nsel=[]; tot=0\nfor j in order:\n if final[j] <= -1e8: break\n sel.append(int(ids[j])); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)\n\nif EXPLORE:\n id2text={}\n with open(POOL) as fh:\n for line in fh:\n r=json.loads(line); id2text[r[\"id\"]]=r[\"text\"]\n def show(rank_idxs, tag):\n for j in rank_idxs:\n i=int(ids[j])\n print(f\"[{tag} sc {sc[j]:.3f} lp {lpar[j]:.3f} lr {lrar[j]:.3f} id {i}] {id2text[i][:200].strip()[:200]!r}\")\n top=order[:20000]\n codef=np.mean([(\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in top])\n print(f\"\\ntop-20k code/pre-frac={codef:.3f}\")\n print(\"\\n== TOP 8 ==\"); show(order[:8],\"top\")\n print(\"\\n== rank ~8000 ==\"); show(order[8000:8003],\"mid\")\n print(\"\\n== rank ~16000 ==\"); show(order[16000:16003],\"lo\")\n g=sc[gate==1]\n print(f\"\\ngated scores: p50 {np.percentile(g,50):.3f} p90 {np.percentile(g,90):.3f} p99 {np.percentile(g,99):.3f}\")\n"
}Write result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"type": "update", "filePath": "/workspace/submission/curate.py", "content": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible), CCNet-style:\n 1. GATE — drop non-English / boilerplate / word-salad / degenerate docs with cheap\n structural filters (English stop-word ratio, letter ratio, unique-word ratio,\n word length, length). The stop-word-ratio gate is the workhorse: fluent English\n prose (all four target registers) is rich in {the,of,and,to,a,in,is,...} whereas\n CSS/JS dumps, foreign text and SEO word-salad are not.\n 2. SCORE — rank survivors by a TARGET bigram language model's mean per-word\n log-probability (interpolated bigram+unigram+uniform). Documents whose wording\n is *typical of the target* score high; this rewards clean encyclopedic / news /\n web-prose / Q&A English and demotes odd-but-English filler.\n 3. A small target-vs-pool log-ratio bonus adds discrimination toward\n target-distinctive content over generic web filler.\n Emit ids best-first until ~2x the 12M-token budget so packing never underfills.\n\nThe dev target only DEFINES the target word distribution (which generalises to the\nhidden official sample); no per-id labels are used. Per-doc features are cached to\n/tmp so the ranking/threshold can be re-derived without re-reading the pool.\n\"\"\"\nimport json, re, math, sys, os, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/tmp/pool_feats.npz\"\n\nMAX_WORDS = 600 # words scanned per doc for scoring (bounds cost)\nBG_SAMPLE = 40000 # pool docs for background unigram model (log-ratio bonus)\nBUDGET = 12_000_000\nOVER = 2.0 # over-provide ids to this multiple of the budget\nCHARS_PER_TOK = 4.435\nEXPLORE = \"--explore\" in sys.argv\nRATIO_W = 0.35 # weight of target-vs-pool log-ratio bonus\n\nSTOP = set(\"the of and to a in is that it for on with as was are be by this at from \"\n \"or an not but have has had he she they we you i his her their our your its \"\n \"which who will would can could there been were said more one all if them \"\n \"when so what about into than then some other time up out only over also \"\n \"no do does did how new may these two his\".split())\nword_re = re.compile(r\"[a-z][a-z']+\") # alphabetic words only (for gates + LM)\n\ndef words(text):\n return word_re.findall(text.lower())\n\n# ---------------- 1. target bigram LM + unigram, from decoded dev ----------------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist(); EOS = 50256\ndocs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else: cur.append(t)\nif cur: docs.append(cur)\ndef clean(s): return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_uni = defaultdict(int); tgt_bi = defaultdict(int); tgt_N = 0\nfor d in docs:\n ws = words(clean(tok.decode(d)))\n for j,wd in enumerate(ws):\n tgt_uni[wd] += 1; tgt_N += 1\n if j: tgt_bi[(ws[j-1], wd)] += 1\nV = len(tgt_uni)\nprint(f\"target: {len(docs)} docs, {tgt_N} words, vocab {V}\", flush=True)\n\n# ---------------- 2. pool background unigram (for log-ratio bonus) ----------------\nbg_uni = defaultdict(int); bg_N = 0; nbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n for wd in words(json.loads(line)[\"text\"])[:MAX_WORDS]:\n bg_uni[wd] += 1; bg_N += 1\n nbg += 1\nprint(f\"background: {nbg} docs, {bg_N} words\", flush=True)\n\n# precompute unigram log-probs and log-ratio per word\nL2, L1, L0 = 0.6, 0.399, 0.001 # bigram / unigram / uniform interpolation\ndef uni_logp(wd):\n return math.log(L1 * tgt_uni.get(wd,0)/tgt_N + L0/1.0 * 1.0/ (V+1))\n# log ratio target/pool for a word (smoothed), clipped\ndef logratio(wd):\n pt = (tgt_uni.get(wd,0)+0.5)/(tgt_N+0.5*V)\n pp = (bg_uni.get(wd,0)+0.5)/(bg_N+0.5*V)\n return max(-3.0, min(3.0, math.log(pt/pp)))\n\ndef score_doc(ws):\n \"\"\"mean per-word target-LM logprob (+ mean log-ratio bonus).\"\"\"\n n = len(ws)\n if n == 0: return -99.0, 0.0\n lp = 0.0; lr = 0.0; prev = None\n for wd in ws:\n pu = tgt_uni.get(wd,0)/tgt_N\n if prev is not None:\n cb = tgt_bi.get((prev,wd),0)\n pbi = cb/tgt_uni[prev] if tgt_uni.get(prev,0) else 0.0\n else:\n pbi = 0.0\n p = L2*pbi + L1*pu + L0*(1.0/(V+1))\n lp += math.log(p)\n lr += logratio(wd)\n prev = wd\n return lp/n, lr/n\n\n# ---------------- 3. single pass over pool: gates + score, cache ----------------\ndef load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)\n\n# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------\nfinal = np.where(gate==1, sc, -1e9)\norder = np.argsort(-final)\nsel=[]; tot=0\nfor j in order:\n if final[j] <= -1e8: break\n sel.append(int(ids[j])); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)\n\nif EXPLORE:\n id2text={}\n with open(POOL) as fh:\n for line in fh:\n r=json.loads(line); id2text[r[\"id\"]]=r[\"text\"]\n def show(rank_idxs, tag):\n for j in rank_idxs:\n i=int(ids[j])\n print(f\"[{tag} sc {sc[j]:.3f} lp {lpar[j]:.3f} lr {lrar[j]:.3f} id {i}] {id2text[i][:200].strip()[:200]!r}\")\n top=order[:20000]\n codef=np.mean([(\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in top])\n print(f\"\\ntop-20k code/pre-frac={codef:.3f}\")\n print(\"\\n== TOP 8 ==\"); show(order[:8],\"top\")\n print(\"\\n== rank ~8000 ==\"); show(order[8000:8003],\"mid\")\n print(\"\\n== rank ~16000 ==\"); show(order[16000:16003],\"lo\")\n g=sc[gate==1]\n print(f\"\\ngated scores: p50 {np.percentile(g,50):.3f} p90 {np.percentile(g,90):.3f} p99 {np.percentile(g,99):.3f}\")\n", "structuredPatch": [{"oldStart": 1, "oldLines": 162, "newStart": 1, "newLines": 161, "lines": [" \"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching", " to the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).", " ", "-Criterion (stated, reproducible):", "- For every pool document we estimate an IMPORTANCE SCORE = the length-normalised", "- log-likelihood ratio log P_target(features) - log P_pool(features) under two", "- Naive-Bayes n-gram models (DSIR-style). P_target is fit on the decoded dev target", "- (multi_dev.npy); P_pool is fit on a random pool background. Documents whose word", "- distribution looks like the target score high. We drop obvious junk with a few", "- cheap quality gates, then emit ids ordered best-first until well over the budget.", "+Criterion (stated, reproducible), CCNet-style:", "+ 1. GATE — drop non-English / boilerplate / word-salad / degenerate docs with cheap", "+ structural filters (English stop-word ratio, letter ratio, unique-word ratio,", "+ word length, length). The stop-word-ratio gate is the workhorse: fluent English", "+ prose (all four target registers) is rich in {the,of,and,to,a,in,is,...} whereas", "+ CSS/JS dumps, foreign text and SEO word-salad are not.", "+ 2. SCORE — rank survivors by a TARGET bigram language model's mean per-word", "+ log-probability (interpolated bigram+unigram+uniform). Documents whose wording", "+ is *typical of the target* score high; this rewards clean encyclopedic / news /", "+ web-prose / Q&A English and demotes odd-but-English filler.", "+ 3. A small target-vs-pool log-ratio bonus adds discrimination toward", "+ target-distinctive content over generic web filler.", "+ Emit ids best-first until ~2x the 12M-token budget so packing never underfills.", " ", "-The dev target is only used to DEFINE the target word distribution (which generalises", "-to the hidden official sample); no per-id labels are used.", "+The dev target only DEFINES the target word distribution (which generalises to the", "+hidden official sample); no per-id labels are used. Per-doc features are cached to", "+/tmp so the ranking/threshold can be re-derived without re-reading the pool.", " \"\"\"", "-import json, re, math, sys, numpy as np", "+import json, re, math, sys, os, numpy as np", " from collections import defaultdict", " ", " POOL = \"/workspace/data/pool.jsonl\"", " DEV = \"/workspace/data/multi_dev.npy\"", " OUT = \"/workspace/submission/selection.json\"", "+CACHE = \"/tmp/pool_feats.npz\"", " ", "-BG_SAMPLE = 60000 # pool docs used to fit the background model", "-MAX_WORDS = 1200 # cap words scanned per doc (bounds cost; long ctx unused anyway)", "-VOCAB_CAP = 2_000_000 # max distinct n-gram features", "-ALPHA = 1.0 # Laplace smoothing", "-TARGET_DOCS_TOKENS = 12_000_000 # over-provide ids to guarantee the 12M budget fills", "+MAX_WORDS = 600 # words scanned per doc for scoring (bounds cost)", "+BG_SAMPLE = 40000 # pool docs for background unigram model (log-ratio bonus)", "+BUDGET = 12_000_000", "+OVER = 2.0 # over-provide ids to this multiple of the budget", "+CHARS_PER_TOK = 4.435", " EXPLORE = \"--explore\" in sys.argv", "+RATIO_W = 0.35 # weight of target-vs-pool log-ratio bonus", " ", "-word_re = re.compile(r\"[a-z0-9']+\")", "-def feats(text):", "- \"\"\"lowercased unigrams + bigrams, capped.\"\"\"", "- w = word_re.findall(text.lower())[:MAX_WORDS]", "- if not w:", "- return []", "- out = w[:] # unigrams", "- out += [w[i]+\" \"+w[i+1] for i in range(len(w)-1)] # bigrams", "- return out", "+STOP = set(\"the of and to a in is that it for on with as was are be by this at from \"", "+ \"or an not but have has had he she they we you i his her their our your its \"", "+ \"which who will would can could there been were said more one all if them \"", "+ \"when so what about into than then some other time up out only over also \"", "+ \"no do does did how new may these two his\".split())", "+word_re = re.compile(r\"[a-z][a-z']+\") # alphabetic words only (for gates + LM)", " ", "-# ---- deterministic exact vocab (no hashing => reproducible, no collisions) ----", "-vocab = {}", "-def fid(f, grow):", "- i = vocab.get(f, -1)", "- if i >= 0: return i", "- if grow and len(vocab) < VOCAB_CAP:", "- i = len(vocab); vocab[f] = i; return i", "- return -1", "+def words(text):", "+ return word_re.findall(text.lower())", " ", "-# ---------- 1. decode dev target, strip WikiText markers, split on EOS ----------", "+# ---------------- 1. target bigram LM + unigram, from decoded dev ----------------", " from transformers import AutoTokenizer", " tok = AutoTokenizer.from_pretrained(\"gpt2\")", "-dev = np.load(DEV).tolist()", "-EOS = 50256", "-tgt_docs, cur = [], []", "+dev = np.load(DEV).tolist(); EOS = 50256", "+docs, cur = [], []", " for t in dev:", " if t == EOS:", "- if cur: tgt_docs.append(cur); cur = []", "- else:", "- cur.append(t)", "-if cur: tgt_docs.append(cur)", "-def clean(s): # remove WikiText @-@ / @,@ / @.@ formatting absent from the raw pool", "- return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")", "-tgt_texts = [clean(tok.decode(d)) for d in tgt_docs]", "-print(f\"target docs: {len(tgt_texts)}\")", "+ if cur: docs.append(cur); cur = []", "+ else: cur.append(t)", "+if cur: docs.append(cur)", "+def clean(s): return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")", "+tgt_uni = defaultdict(int); tgt_bi = defaultdict(int); tgt_N = 0", "+for d in docs:", "+ ws = words(clean(tok.decode(d)))", "+ for j,wd in enumerate(ws):", "+ tgt_uni[wd] += 1; tgt_N += 1", "+ if j: tgt_bi[(ws[j-1], wd)] += 1", "+V = len(tgt_uni)", "+print(f\"target: {len(docs)} docs, {tgt_N} words, vocab {V}\", flush=True)", " ", "-# ---------- 2. fit counts ----------", "-tgt_cnt = defaultdict(float)", "-for tx in tgt_texts:", "- for f in feats(tx):", "- i = fid(f, True)", "- if i >= 0: tgt_cnt[i] += 1.0", "-", "-bg_cnt = defaultdict(float)", "-nbg = 0", "+# ---------------- 2. pool background unigram (for log-ratio bonus) ----------------", "+bg_uni = defaultdict(int); bg_N = 0; nbg = 0", " with open(POOL) as fh:", " for line in fh:", " if nbg >= BG_SAMPLE: break", "- tx = json.loads(line)[\"text\"]", "- for f in feats(tx):", "- i = fid(f, True)", "- if i >= 0: bg_cnt[i] += 1.0", "+ for wd in words(json.loads(line)[\"text\"])[:MAX_WORDS]:", "+ bg_uni[wd] += 1; bg_N += 1", " nbg += 1", "-V = len(vocab)", "-print(f\"vocab: {V} bg docs: {nbg}\")", "+print(f\"background: {nbg} docs, {bg_N} words\", flush=True)", " ", "-# ---------- 3. per-feature log-likelihood-ratio weights ----------", "-tgt_tot = sum(tgt_cnt.values()); bg_tot = sum(bg_cnt.values())", "-w = np.zeros(V, dtype=np.float32)", "-tc = np.zeros(V, dtype=np.float32); bc = np.zeros(V, dtype=np.float32)", "-for i,c in tgt_cnt.items(): tc[i] = c", "-for i,c in bg_cnt.items(): bc[i] = c", "-w = (np.log(tc + ALPHA) - math.log(tgt_tot + ALPHA*V)) - \\", "- (np.log(bc + ALPHA) - math.log(bg_tot + ALPHA*V))", "-# features unseen in target (tc==0) get strongly negative weight; that is correct.", "+# precompute unigram log-probs and log-ratio per word", "+L2, L1, L0 = 0.6, 0.399, 0.001 # bigram / unigram / uniform interpolation", "+def uni_logp(wd):", "+ return math.log(L1 * tgt_uni.get(wd,0)/tgt_N + L0/1.0 * 1.0/ (V+1))", "+# log ratio target/pool for a word (smoothed), clipped", "+def logratio(wd):", "+ pt = (tgt_uni.get(wd,0)+0.5)/(tgt_N+0.5*V)", "+ pp = (bg_uni.get(wd,0)+0.5)/(bg_N+0.5*V)", "+ return max(-3.0, min(3.0, math.log(pt/pp)))", " ", "-# ---------- 4. score every pool doc ----------", "-def quality_ok(text, nw, alpha_frac, mwl):", "- if nw < 50: return False # too short: mostly EOS overhead, little signal", "- if alpha_frac < 0.55: return False # tables / code-dumps / gibberish", "- if mwl < 3.0 or mwl > 10.0: return False", "- return True", "+def score_doc(ws):", "+ \"\"\"mean per-word target-LM logprob (+ mean log-ratio bonus).\"\"\"", "+ n = len(ws)", "+ if n == 0: return -99.0, 0.0", "+ lp = 0.0; lr = 0.0; prev = None", "+ for wd in ws:", "+ pu = tgt_uni.get(wd,0)/tgt_N", "+ if prev is not None:", "+ cb = tgt_bi.get((prev,wd),0)", "+ pbi = cb/tgt_uni[prev] if tgt_uni.get(prev,0) else 0.0", "+ else:", "+ pbi = 0.0", "+ p = L2*pbi + L1*pu + L0*(1.0/(V+1))", "+ lp += math.log(p)", "+ lr += logratio(wd)", "+ prev = wd", "+ return lp/n, lr/n", " ", "-ids, scores, ntoks = [], [], []", "-kept_texts_dbg = []", "-with open(POOL) as fh:", "- for line in fh:", "- r = json.loads(line); text = r[\"text\"]", "- low = text.lower()", "- wlist = word_re.findall(low)[:MAX_WORDS]", "- nw = len(wlist)", "- if nw == 0:", "- continue", "- nalpha = sum(c.isalpha() for c in text[:6000])", "- alpha_frac = nalpha / max(1, len(text[:6000]))", "- mwl = sum(len(x) for x in wlist)/nw", "- # score = mean weight over features (unigram+bigram), length-normalised", "- fs = wlist[:]", "- fs += [wlist[i]+\" \"+wlist[i+1] for i in range(len(wlist)-1)]", "- s = 0.0; n = 0", "- for f in fs:", "- i = vocab.get(f, -1)", "- if i >= 0:", "- s += w[i]; n += 1", "- if n == 0:", "- continue", "- sc = s / n", "- ok = quality_ok(text, nw, alpha_frac, mwl)", "- ids.append(r[\"id\"]); scores.append(sc if ok else -1e9)", "- ntoks.append(int(len(text)/4.435) + 1)", "+# ---------------- 3. single pass over pool: gates + score, cache ----------------", "+def load_and_score():", "+ ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]", "+ with open(POOL) as fh:", "+ for line in fh:", "+ r = json.loads(line); text = r[\"text\"]", "+ ws = words(text)[:MAX_WORDS]", "+ nw = len(ws)", "+ ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)", "+ if nw < 60:", "+ sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue", "+ head = text[:6000]", "+ alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))", "+ stop_ratio = sum(w in STOP for w in ws)/nw", "+ uniq_ratio = len(set(ws))/nw", "+ mwl = sum(len(w) for w in ws)/nw", "+ ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65", "+ and uniq_ratio>0.34 and 3.0<=mwl<=9.0)", "+ lp, lr = score_doc(ws)", "+ sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)", "+ gate.append(1 if ok else 0)", "+ return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),", "+ np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))", " ", "-ids = np.array(ids); scores = np.array(scores, dtype=np.float32); ntoks = np.array(ntoks)", "-order = np.argsort(-scores) # best first", "+ids, sc, lpar, lrar, ntok, gate = load_and_score()", "+np.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)", "+print(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)", " ", "-# ---------- 5. emit ids best-first until well over budget ----------", "-sel, tot = [], 0", "+# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------", "+final = np.where(gate==1, sc, -1e9)", "+order = np.argsort(-final)", "+sel=[]; tot=0", " for j in order:", "- if scores[j] <= -1e8: # filtered junk: stop adding once we reach them", "- break", "- sel.append(int(ids[j])); tot += int(ntoks[j])", "- if tot >= TARGET_DOCS_TOKENS:", "- # keep going a bit to over-provide (2x budget) so packing never underfills", "- if tot >= 2*TARGET_DOCS_TOKENS: break", "-json.dump(sel, open(OUT, \"w\"))", "-print(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\")", "+ if final[j] <= -1e8: break", "+ sel.append(int(ids[j])); tot += int(ntok[j])", "+ if tot >= OVER*BUDGET: break", "+json.dump(sel, open(OUT,\"w\"))", "+print(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)", " ", " if EXPLORE:", "- id2text = {}", "+ id2text={}", " with open(POOL) as fh:", " for line in fh:", "- r = json.loads(line); id2text[r[\"id\"]] = r[\"text\"]", "- def probe(idxs, tag):", "- code = sum((\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in idxs)", "- print(f\" [{tag}] n={len(idxs)} code/pre-frac={code/len(idxs):.2f} mean_score={np.mean([scores[j] for j in idxs]):.3f}\")", "- top = order[:20000]; print(\"register probe on top-20k:\"); probe(top,\"top20k\")", "- print(\"\\n===== TOP 6 =====\")", "- for j in order[:6]:", "- print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:280].strip()[:280]!r}\")", "- print(\"\\n===== around rank 12000 =====\")", "- for j in order[12000:12003]:", "- print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:220].strip()[:220]!r}\")", "- # score distribution", "- valid = scores[scores>-1e8]", "- print(f\"\\nvalid docs: {len(valid)} score p50 {np.percentile(valid,50):.3f} p90 {np.percentile(valid,90):.3f} p99 {np.percentile(valid,99):.3f}\")", "+ r=json.loads(line); id2text[r[\"id\"]]=r[\"text\"]", "+ def show(rank_idxs, tag):", "+ for j in rank_idxs:", "+ i=int(ids[j])", "+ print(f\"[{tag} sc {sc[j]:.3f} lp {lpar[j]:.3f} lr {lrar[j]:.3f} id {i}] {id2text[i][:200].strip()[:200]!r}\")", "+ top=order[:20000]", "+ codef=np.mean([(\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in top])", "+ print(f\"\\ntop-20k code/pre-frac={codef:.3f}\")", "+ print(\"\\n== TOP 8 ==\"); show(order[:8],\"top\")", "+ print(\"\\n== rank ~8000 ==\"); show(order[8000:8003],\"mid\")", "+ print(\"\\n== rank ~16000 ==\"); show(order[16000:16003],\"lo\")", "+ g=sc[gate==1]", "+ print(f\"\\ngated scores: p50 {np.percentile(g,50):.3f} p90 {np.percentile(g,90):.3f} p99 {np.percentile(g,99):.3f}\")"]}], "originalFile": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible):\n For every pool document we estimate an IMPORTANCE SCORE = the length-normalised\n log-likelihood ratio log P_target(features) - log P_pool(features) under two\n Naive-Bayes n-gram models (DSIR-style). P_target is fit on the decoded dev target\n (multi_dev.npy); P_pool is fit on a random pool background. Documents whose word\n distribution looks like the target score high. We drop obvious junk with a few\n cheap quality gates, then emit ids ordered best-first until well over the budget.\n\nThe dev target is only used to DEFINE the target word distribution (which generalises\nto the hidden official sample); no per-id labels are used.\n\"\"\"\nimport json, re, math, sys, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\n\nBG_SAMPLE = 60000 # pool docs used to fit the background model\nMAX_WORDS = 1200 # cap words scanned per doc (bounds cost; long ctx unused anyway)\nVOCAB_CAP = 2_000_000 # max distinct n-gram features\nALPHA = 1.0 # Laplace smoothing\nTARGET_DOCS_TOKENS = 12_000_000 # over-provide ids to guarantee the 12M budget fills\nEXPLORE = \"--explore\" in sys.argv\n\nword_re = re.compile(r\"[a-z0-9']+\")\ndef feats(text):\n \"\"\"lowercased unigrams + bigrams, capped.\"\"\"\n w = word_re.findall(text.lower())[:MAX_WORDS]\n if not w:\n return []\n out = w[:] # unigrams\n out += [w[i]+\" \"+w[i+1] for i in range(len(w)-1)] # bigrams\n return out\n\n# ---- deterministic exact vocab (no hashing => reproducible, no collisions) ----\nvocab = {}\ndef fid(f, grow):\n i = vocab.get(f, -1)\n if i >= 0: return i\n if grow and len(vocab) < VOCAB_CAP:\n i = len(vocab); vocab[f] = i; return i\n return -1\n\n# ---------- 1. decode dev target, strip WikiText markers, split on EOS ----------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist()\nEOS = 50256\ntgt_docs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: tgt_docs.append(cur); cur = []\n else:\n cur.append(t)\nif cur: tgt_docs.append(cur)\ndef clean(s): # remove WikiText @-@ / @,@ / @.@ formatting absent from the raw pool\n return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_texts = [clean(tok.decode(d)) for d in tgt_docs]\nprint(f\"target docs: {len(tgt_texts)}\")\n\n# ---------- 2. fit counts ----------\ntgt_cnt = defaultdict(float)\nfor tx in tgt_texts:\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: tgt_cnt[i] += 1.0\n\nbg_cnt = defaultdict(float)\nnbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n tx = json.loads(line)[\"text\"]\n for f in feats(tx):\n i = fid(f, True)\n if i >= 0: bg_cnt[i] += 1.0\n nbg += 1\nV = len(vocab)\nprint(f\"vocab: {V} bg docs: {nbg}\")\n\n# ---------- 3. per-feature log-likelihood-ratio weights ----------\ntgt_tot = sum(tgt_cnt.values()); bg_tot = sum(bg_cnt.values())\nw = np.zeros(V, dtype=np.float32)\ntc = np.zeros(V, dtype=np.float32); bc = np.zeros(V, dtype=np.float32)\nfor i,c in tgt_cnt.items(): tc[i] = c\nfor i,c in bg_cnt.items(): bc[i] = c\nw = (np.log(tc + ALPHA) - math.log(tgt_tot + ALPHA*V)) - \\\n (np.log(bc + ALPHA) - math.log(bg_tot + ALPHA*V))\n# features unseen in target (tc==0) get strongly negative weight; that is correct.\n\n# ---------- 4. score every pool doc ----------\ndef quality_ok(text, nw, alpha_frac, mwl):\n if nw < 50: return False # too short: mostly EOS overhead, little signal\n if alpha_frac < 0.55: return False # tables / code-dumps / gibberish\n if mwl < 3.0 or mwl > 10.0: return False\n return True\n\nids, scores, ntoks = [], [], []\nkept_texts_dbg = []\nwith open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n low = text.lower()\n wlist = word_re.findall(low)[:MAX_WORDS]\n nw = len(wlist)\n if nw == 0:\n continue\n nalpha = sum(c.isalpha() for c in text[:6000])\n alpha_frac = nalpha / max(1, len(text[:6000]))\n mwl = sum(len(x) for x in wlist)/nw\n # score = mean weight over features (unigram+bigram), length-normalised\n fs = wlist[:]\n fs += [wlist[i]+\" \"+wlist[i+1] for i in range(len(wlist)-1)]\n s = 0.0; n = 0\n for f in fs:\n i = vocab.get(f, -1)\n if i >= 0:\n s += w[i]; n += 1\n if n == 0:\n continue\n sc = s / n\n ok = quality_ok(text, nw, alpha_frac, mwl)\n ids.append(r[\"id\"]); scores.append(sc if ok else -1e9)\n ntoks.append(int(len(text)/4.435) + 1)\n\nids = np.array(ids); scores = np.array(scores, dtype=np.float32); ntoks = np.array(ntoks)\norder = np.argsort(-scores) # best first\n\n# ---------- 5. emit ids best-first until well over budget ----------\nsel, tot = [], 0\nfor j in order:\n if scores[j] <= -1e8: # filtered junk: stop adding once we reach them\n break\n sel.append(int(ids[j])); tot += int(ntoks[j])\n if tot >= TARGET_DOCS_TOKENS:\n # keep going a bit to over-provide (2x budget) so packing never underfills\n if tot >= 2*TARGET_DOCS_TOKENS: break\njson.dump(sel, open(OUT, \"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\")\n\nif EXPLORE:\n id2text = {}\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); id2text[r[\"id\"]] = r[\"text\"]\n def probe(idxs, tag):\n code = sum((\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in idxs)\n print(f\" [{tag}] n={len(idxs)} code/pre-frac={code/len(idxs):.2f} mean_score={np.mean([scores[j] for j in idxs]):.3f}\")\n top = order[:20000]; print(\"register probe on top-20k:\"); probe(top,\"top20k\")\n print(\"\\n===== TOP 6 =====\")\n for j in order[:6]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:280].strip()[:280]!r}\")\n print(\"\\n===== around rank 12000 =====\")\n for j in order[12000:12003]:\n print(f\"[score {scores[j]:.3f} id {int(ids[j])}] {id2text[int(ids[j])][:220].strip()[:220]!r}\")\n # score distribution\n valid = scores[scores>-1e8]\n print(f\"\\nvalid docs: {len(valid)} score p50 {np.percentile(valid,50):.3f} p90 {np.percentile(valid,90):.3f} p99 {np.percentile(valid,99):.3f}\")\n", "userModified": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; python3 submission/curate.py --explore > /tmp/curate.log 2>&1 &\necho \"launched pid $!\"; sleep 2; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Launch curate v2 in background"
}Bash result
launched pid 1214
remaining: 4669
[stdout]
launched pid 1214
remaining: 4669
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 100; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"; echo \"=== log ===\"; grep -v \"sequence length\" /tmp/curate.log",
"description": "Poll curate v2 progress"
}Bash result
remaining: 4559
=== log ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
[stdout]
remaining: 4559
=== log ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 120; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"; echo \"=== log ===\"; grep -v \"sequence length\" /tmp/curate.log",
"description": "Poll curate v2 scoring pass"
}Bash result · error
Exit code 143
Command timed out after 2m 0s
[error] tool reported failureencrypted chain-of-thought — recorded, not exposed
args
{
"command": "echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"; echo \"=== log ===\"; grep -v \"sequence length\" /tmp/curate.log; echo \"=== proc ===\"; ps aux | grep curate | grep -v grep | head",
"description": "Check curate log and process"
}Bash result
remaining: 4428
=== log ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
scored 182016 docs; gate-pass 142940
selection: 26535 ids, ~24.0M est tokens -> /workspace/submission/selection.json
top-20k code/pre-frac=0.000
== TOP 8 ==
[top sc -5.212 lp -5.176 lr -0.105 id 32526] 'My dad has a green card and went back home over 2 years ago. He is now trying to come back after being away from the US for 2 years. I was told by some people that if he stayed over a year outside the'
[top sc -5.577 lp -5.477 lr -0.283 id 11940] 'You don’t know it yet, but we love you. You don’t know it yet, but we have been waiting for you.\nAnd you don’t know it yet, but you already have a place in our home and in our hearts.”\nWaiting is a di'
[top sc -5.639 lp -5.567 lr -0.206 id 20465] 'If you’re hoping to make as much money as possible, one of the main things you’ll need to think about is getting involved in some sort of investing scheme. Since people typically are limited in their'
[top sc -5.658 lp -5.632 lr -0.074 id 9745] "|Close up of the design I put on it's top|\nI decided to do this table similiar to the first since they were going to be in the same area.\n|Is it just me or does it look like it has a face now?|\n|Same"
[top sc -5.659 lp -5.592 lr -0.191 id 69097] 'you have come to this section of the website because you need to arrange a funeral then this is likely to be a difficult time for you and so we hope that these pages will make working out the details'
[top sc -5.673 lp -5.573 lr -0.287 id 12142] 'Benefits of E-learning\nThe benefits of e-learning are more than just the obvious. When you consider the number of people who will be using it over the next several years, then you can begin to underst'
[top sc -5.773 lp -5.965 lr 0.551 id 82269] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950'
[top sc -5.815 lp -5.728 lr -0.249 id 81062] '<|endoftext|>Importance Things Regarding IT Services That You Need To Know Of\nThese days, there has been an increase in the number of individuals who are looking forward to making their company more s'
== rank ~8000 ==
[mid sc -6.968 lp -6.923 lr -0.128 id 28686] 'beat Dad With 22 Kids Is As Charmingly Confused As You’d Expect Him to Be\nWere it not for George Washington, you could argue that Orlando Shaw is actually the father of our country.\nThat’s because the'
[mid sc -6.968 lp -7.007 lr 0.112 id 89616] '."<|endoftext|>Los Angeles (CNN) -- Reza Taghavi, an Iranian-American detained for more than two years in Iran, returned to the United States Thursday evening.\nThe retired businessman from Orange Coun'
[mid sc -6.968 lp -6.906 lr -0.176 id 20293] 'David Lyons was raised in Carson City and has recently returned home, to his roots. David is now able to use his skills and experience to serve the community that he feels most connected to.\nDavid com'
== rank ~16000 ==
[lo sc -7.176 lp -7.089 lr -0.249 id 114243] 'know how to remove a toilet you can save money by tackling this DIY job yourself. Any resulting rubbish can be taken away with one our bathroom waste removal services.\nMy Job Quote estimates that the'
[lo sc -7.176 lp -7.195 lr 0.053 id 68594] 'oding in Schools: Should It Replace a Language Requirement?\nLegislation is pending in Kentucky that would let high schoolers get foreign language credit for learning to program computers. New Mexico i'
[lo sc -7.176 lp -7.093 lr -0.238 id 96873] 'a country girl kicking around the dirt roads taking pictures of trees, gates, animals. cactus, and anything else that strikes my un-fancy mind. You are welcome to take a look and come back if you lik'
gated scores: p50 -8.036 p90 -7.138 p99 -6.599
=== proc ===
root 154 0.0 0.0 12432 4708 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid exit "$rc"
root 156 0.0 0.0 12432 4528 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.7 0.0 6296324 532132 ? Sl 03:22 0:45 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
[stdout]
remaining: 4428
=== log ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
scored 182016 docs; gate-pass 142940
selection: 26535 ids, ~24.0M est tokens -> /workspace/submission/selection.json
top-20k code/pre-frac=0.000
== TOP 8 ==
[top sc -5.212 lp -5.176 lr -0.105 id 32526] 'My dad has a green card and went back home over 2 years ago. He is now trying to come back after being away from the US for 2 years. I was told by some people that if he stayed over a year outside the'
[top sc -5.577 lp -5.477 lr -0.283 id 11940] 'You don’t know it yet, but we love you. You don’t know it yet, but we have been waiting for you.\nAnd you don’t know it yet, but you already have a place in our home and in our hearts.”\nWaiting is a di'
[top sc -5.639 lp -5.567 lr -0.206 id 20465] 'If you’re hoping to make as much money as possible, one of the main things you’ll need to think about is getting involved in some sort of investing scheme. Since people typically are limited in their'
[top sc -5.658 lp -5.632 lr -0.074 id 9745] "|Close up of the design I put on it's top|\nI decided to do this table similiar to the first since they were going to be in the same area.\n|Is it just me or does it look like it has a face now?|\n|Same"
[top sc -5.659 lp -5.592 lr -0.191 id 69097] 'you have come to this section of the website because you need to arrange a funeral then this is likely to be a difficult time for you and so we hope that these pages will make working out the details'
[top sc -5.673 lp -5.573 lr -0.287 id 12142] 'Benefits of E-learning\nThe benefits of e-learning are more than just the obvious. When you consider the number of people who will be using it over the next several years, then you can begin to underst'
[top sc -5.773 lp -5.965 lr 0.551 id 82269] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950'
[top sc -5.815 lp -5.728 lr -0.249 id 81062] '<|endoftext|>Importance Things Regarding IT Services That You Need To Know Of\nThese days, there has been an increase in the number of individuals who are looking forward to making their company more s'
== rank ~8000 ==
[mid sc -6.968 lp -6.923 lr -0.128 id 28686] 'beat Dad With 22 Kids Is As Charmingly Confused As You’d Expect Him to Be\nWere it not for George Washington, you could argue that Orlando Shaw is actually the father of our country.\nThat’s because the'
[mid sc -6.968 lp -7.007 lr 0.112 id 89616] '."<|endoftext|>Los Angeles (CNN) -- Reza Taghavi, an Iranian-American detained for more than two years in Iran, returned to the United States Thursday evening.\nThe retired businessman from Orange Coun'
[mid sc -6.968 lp -6.906 lr -0.176 id 20293] 'David Lyons was raised in Carson City and has recently returned home, to his roots. David is now able to use his skills and experience to serve the community that he feels most connected to.\nDavid com'
== rank ~16000 ==
[lo sc -7.176 lp -7.089 lr -0.249 id 114243] 'know how to remove a toilet you can save money by tackling this DIY job yourself. Any resulting rubbish can be taken away with one our bathroom waste removal services.\nMy Job Quote estimates that the'
[lo sc -7.176 lp -7.195 lr 0.053 id 68594] 'oding in Schools: Should It Replace a Language Requirement?\nLegislation is pending in Kentucky that would let high schoolers get foreign language credit for learning to program computers. New Mexico i'
[lo sc -7.176 lp -7.093 lr -0.238 id 96873] 'a country girl kicking around the dirt roads taking pictures of trees, gates, animals. cactus, and anything else that strikes my un-fancy mind. You are welcome to take a look and come back if you lik'
gated scores: p50 -8.036 p90 -7.138 p99 -6.599
=== proc ===
root 154 0.0 0.0 12432 4708 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-dd16642ce40a494ca0b0960eb8a6a993.pid exit "$rc"
root 156 0.0 0.0 12432 4528 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.7 0.0 6296324 532132 ? Sl 03:22 0:45 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
[metadata] {"noOutputExpected": false}v2 is dramatically better — top docs are clean English across registers (immigration Q&A, CNN news, "Hill 303 massacre" Wikipedia-style, heartfelt prose). Gate passes 142,940 docs. But technical/code Q&A (`<code>`/`<pre>`) is still 0% — my alpha_frac>0.6 gate excludes it, and that's ~25% of the target.
Let me anchor with an empirical dev run on the current selection while I probe whether the pool even contains code-Q&A worth adding.
args
{
"command": "cd /workspace; nohup bash dev_eval.sh /workspace/submission/selection.json > /tmp/dev_v2.log 2>&1 &\necho \"dev_eval launched pid $!\"",
"description": "Launch dev_eval on v2 selection"
}Bash result
dev_eval launched pid 1570
[stdout]
dev_eval launched pid 1570
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; python3 - <<'PY'\nimport json, numpy as np\nfrom transformers import AutoTokenizer\ntok=AutoTokenizer.from_pretrained(\"gpt2\")\n# how much of the TARGET is code Q&A?\ndev=np.load(\"data/multi_dev.npy\").tolist(); EOS=50256\ndocs,cur=[],[]\nfor t in dev:\n if t==EOS:\n if cur: docs.append(cur); cur=[]\n else: cur.append(t)\ntgt=[tok.decode(d) for d in docs]\ntcode=sum((\"<code>\" in d or \"<pre>\" in d or \"</p>\" in d) for d in tgt)\nttok_code=sum(len(d) for d in docs if (\"<code>\" in tok.decode(d) or \"<pre>\" in tok.decode(d) or \"</p>\" in tok.decode(d)))\nprint(f\"TARGET: {len(tgt)} docs, code-ish docs={tcode} ({tcode/len(tgt)*100:.0f}%), code-ish tokens~{ttok_code} ({ttok_code/len(dev)*100:.0f}% of eval)\")\n# how much of the POOL is code Q&A?\nncode=0; ntot=0; samp=[]\nwith open(\"data/pool.jsonl\") as f:\n for line in f:\n r=json.loads(line); ntot+=1; t=r[\"text\"]\n if \"<code>\" in t or \"<pre>\" in t:\n ncode+=1\n if len(samp)<4: samp.append((r[\"id\"],t[:220]))\nprint(f\"POOL: {ntot} docs, docs with <code>/<pre> = {ncode} ({ncode/ntot*100:.1f}%)\")\nfor i,s in samp: print(f\" id {i}: {s!r}\")\nPY",
"description": "Probe code-Q&A prevalence in target and pool"
}Bash result
TARGET: 2704 docs, code-ish docs=358 (13%), code-ish tokens~249474 (25% of eval)
POOL: 182016 docs, docs with <code>/<pre> = 128 (0.1%)
id 2033: 'WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Rela'
id 4313: 'Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology |\nThe Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on t'
id 10732: 'This page contains historical information about our SVN server. We now use Git instead.\nThe SVN URL for the original MPICH2 trunk was:\nhttps://svn.mcs.anl.gov/repos/mpi/mpich2/trunk. In place of\ntrunk, there were\ntags<co'
id 28721: "++ I'm commenting mostly just to bump this excellent piece of advice.\nSince port is rarely important and I like to use this idiom in addition to running a traditional webserver on port 80, I'd shorten it to use the defau"
[stdout]
TARGET: 2704 docs, code-ish docs=358 (13%), code-ish tokens~249474 (25% of eval)
POOL: 182016 docs, docs with <code>/<pre> = 128 (0.1%)
id 2033: 'WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Rela'
id 4313: 'Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology |\nThe Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on t'
id 10732: 'This page contains historical information about our SVN server. We now use Git instead.\nThe SVN URL for the original MPICH2 trunk was:\nhttps://svn.mcs.anl.gov/repos/mpi/mpich2/trunk. In place of\ntrunk, there were\ntags<co'
id 28721: "++ I'm commenting mostly just to bump this excellent piece of advice.\nSince port is rarely important and I like to use this idiom in addition to running a traditional webserver on port 80, I'd shorten it to use the defau"
[metadata] {"noOutputExpected": false}Critical finding: 25% of the eval tokens are technical/code Q&A — but the pool has only 128 docs (0.1%) with `<code>`/`<pre>`. That register is severely under-supplied; I can't fill 25% of training with it, but going from zero exposure to some should meaningfully cut perplexity on that quarter of the eval (an unseen register gets near-uniform, catastrophic ppl). Let me check the dev result and probe for broader technical content while waiting.
args
{
"command": "sleep 30; echo \"=== dev_v2 so far ===\"; grep -v \"sequence length\" /tmp/dev_v2.log; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Check dev_eval progress"
}Bash result
=== dev_v2 so far ===
packed 12000000 tokens from 26535 ids -> /tmp/dev_train.npy
remaining: 4308
[stdout]
=== dev_v2 so far ===
packed 12000000 tokens from 26535 ids -> /tmp/dev_train.npy
remaining: 4308
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; python3 - <<'PY'\nimport json, re\n# broader technical detector: code-ish symbols / programming keywords\ncodepat = re.compile(r\"</?(code|pre|p|div|span)>|def |import |function |public |void |\\bclass \\b|#include|println|System\\.|\\$\\w+|=>|\\{\\s|\\breturn \\b|SELECT |console\\.|std::\")\nn_tag=0; n_broad=0; ntot=0; samp=[]\nwith open(\"data/pool.jsonl\") as f:\n for line in f:\n r=json.loads(line); t=r[\"text\"]; ntot+=1\n tag = (\"<code>\" in t or \"<pre>\" in t or \"</p>\" in t)\n broad = tag or bool(codepat.search(t))\n # require it to still be substantial and somewhat english\n if broad: \n n_broad+=1\n if tag: n_tag+=1\n if len(samp)<6 and not tag and codepat.search(t): samp.append((r[\"id\"], t[:200]))\nprint(f\"pool docs: tag-code={n_tag} broad-technical={n_broad} ({n_broad/ntot*100:.1f}%)\")\nprint(\"broad-but-not-tag samples:\")\nfor i,s in samp: print(f\" id {i}: {s!r}\")\nPY",
"description": "Broader technical-content probe"
}Bash result
pool docs: tag-code=192 broad-technical=58383 (32.1%)
broad-but-not-tag samples:
id 8: '|Facility Type:||Full Service Restaurant|\n|Inspection date:||March 27, 2012|\n|Number of critical violations:||3|\n|Number of non-critical violations:||3|\nDefinition of critical and non critical violati'
id 16: 'Michigan unemployment claims workers are losing their jobs.\nNEW YORK (CNNMoney) -- Many jobless claims workers in the state of Michigan will soon be filing for unemployment themselves.\nAbout 400 state'
id 17: "A week ago, talk about Nebraska's running game centered on its stable of talented running backs. Now, you have to throw the quarterback into the mix as well. After redshirt freshman Taylor Martinez ga"
id 18: 'Amazon.com (AMZN) will have a difficult time meeting analyst expectations this quarter given its several significant headwinds. The two most prominent headwinds relate to revenue, namely the impositio'
id 21: 'I’ve a new release of NYTProf ready to upload but I’m stuck.\nThe CPAN Testers service is reporting a failure on a number of systems but I can’t reproduce it locally or work out the cause.\nCan you repr'
id 23: 'Information contained on this page is provided by companies via press release distributed through PR Newswire, an independent third-party content provider. PR Newswire, WorldNow and this Station make '
[stdout]
pool docs: tag-code=192 broad-technical=58383 (32.1%)
broad-but-not-tag samples:
id 8: '|Facility Type:||Full Service Restaurant|\n|Inspection date:||March 27, 2012|\n|Number of critical violations:||3|\n|Number of non-critical violations:||3|\nDefinition of critical and non critical violati'
id 16: 'Michigan unemployment claims workers are losing their jobs.\nNEW YORK (CNNMoney) -- Many jobless claims workers in the state of Michigan will soon be filing for unemployment themselves.\nAbout 400 state'
id 17: "A week ago, talk about Nebraska's running game centered on its stable of talented running backs. Now, you have to throw the quarterback into the mix as well. After redshirt freshman Taylor Martinez ga"
id 18: 'Amazon.com (AMZN) will have a difficult time meeting analyst expectations this quarter given its several significant headwinds. The two most prominent headwinds relate to revenue, namely the impositio'
id 21: 'I’ve a new release of NYTProf ready to upload but I’m stuck.\nThe CPAN Testers service is reporting a failure on a number of systems but I can’t reproduce it locally or work out the cause.\nCan you repr'
id 23: 'Information contained on this page is provided by companies via press release distributed through PR Newswire, an independent third-party content provider. PR Newswire, WorldNow and this Station make '
[metadata] {"noOutputExpected": false}dev_eval is training now (packed 12M from 26,535 ids). The broad technical regex is unreliable (32% — false positives on tables/$amounts); the trustworthy supply is just 192 tag-code docs. I'll include those as a guaranteed "code lane." Let me get the v2 number.
args
{
"command": "sleep 75; echo \"=== dev_v2 result ===\"; grep -Ev \"sequence length\" /tmp/dev_v2.log | tail -5; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Get v2 dev perplexity"
}Bash result
=== dev_v2 result ===
packed 12000000 tokens from 26535 ids -> /tmp/dev_train.npy
HELDOUT_PPL 538.553 loss 6.2889 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
remaining: 4178
[stdout]
=== dev_v2 result ===
packed 12000000 tokens from 26535 ids -> /tmp/dev_train.npy
HELDOUT_PPL 538.553 loss 6.2889 train_s 56 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
remaining: 4178
[metadata] {"noOutputExpected": false}v2 dev perplexity = 538.55. I can't interpret that without the random baseline — that's the critical reference. Let me run it now (a random selection is the "do-nothing" anchor). While it trains, I'll build v3 (clean-English backbone + guaranteed code lane).
args
{
"command": "cd /workspace; python3 - <<'PY'\nimport json, numpy as np\nrng=np.random.default_rng(1337)\nids=list(range(182016))\nrng.shuffle(ids)\njson.dump(ids[:40000], open(\"/tmp/random_sel.json\",\"w\"))\nprint(\"random selection: 40000 ids\")\nPY\nnohup bash dev_eval.sh /tmp/random_sel.json > /tmp/dev_rand.log 2>&1 &\necho \"random dev_eval launched pid $!\"",
"description": "Create and launch random baseline eval"
}Bash result
random selection: 40000 ids
random dev_eval launched pid 1958
[stdout]
random selection: 40000 ids
random dev_eval launched pid 1958
[metadata] {"noOutputExpected": false}Random baseline is training. Now let me upgrade curate.py to v3: add a guaranteed code lane (the ~192 tag-code docs — the only supply for the 25%-of-eval technical register) ahead of the clean-English backbone. I'll fold in tag detection so I only re-run the expensive pass once.
args
{
"replace_all": false,
"file_path": "/workspace/submission/curate.py",
"old_string": "def load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)\n\n# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------\nfinal = np.where(gate==1, sc, -1e9)\norder = np.argsort(-final)\nsel=[]; tot=0\nfor j in order:\n if final[j] <= -1e8: break\n sel.append(int(ids[j])); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)",
"new_string": "def load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]; tag=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n # technical/code Q&A lane: HTML/code tags are the only reliable signal for\n # the target's ~25%-of-tokens code register, which is scarce in the pool.\n is_tag = (\"<code>\" in text or \"<pre>\" in text or \"</p>\" in text)\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n tag.append(1 if is_tag else 0)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8),\n np.array(tag,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate, tag = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate, tag=tag)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}; tag-code {int(tag.sum())}\", flush=True)\n\n# ---------------- 4. emit: code lane first (guaranteed), then clean backbone --------\n# Code lane: every tag-code doc with a minimal amount of English prose, best-first.\ncode_ok = (tag==1) & (sc > -50) # has words + some target signal\ncode_order = np.argsort(-np.where(code_ok, sc, -1e9))\ncode_ids = [int(ids[j]) for j in code_order if code_ok[j]]\n# Backbone: gated clean-English docs, best-first (excluding ones already in code lane).\nfinal = np.where(gate==1, sc, -1e9)\nback_order = np.argsort(-final)\nseen = set(code_ids)\nsel=[]; tot=0\nfor i in code_ids: # code lane at the front\n sel.append(i); tot += int(ntok[np.searchsorted(ids, i)]) if False else 0\n# recompute code-lane token totals cleanly via id->row map\nrow = {int(ids[k]):k for k in range(len(ids))}\nsel=list(code_ids); tot=sum(int(ntok[row[i]]) for i in sel)\nfor j in back_order: # then clean backbone\n if final[j] <= -1e8: break\n i=int(ids[j])\n if i in seen: continue\n sel.append(i); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids ({len(code_ids)} code-lane), ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)"
}Edit result
The file /workspace/submission/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/submission/curate.py", "oldString": "def load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)\n\n# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------\nfinal = np.where(gate==1, sc, -1e9)\norder = np.argsort(-final)\nsel=[]; tot=0\nfor j in order:\n if final[j] <= -1e8: break\n sel.append(int(ids[j])); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)", "newString": "def load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]; tag=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n # technical/code Q&A lane: HTML/code tags are the only reliable signal for\n # the target's ~25%-of-tokens code register, which is scarce in the pool.\n is_tag = (\"<code>\" in text or \"<pre>\" in text or \"</p>\" in text)\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n tag.append(1 if is_tag else 0)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8),\n np.array(tag,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate, tag = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate, tag=tag)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}; tag-code {int(tag.sum())}\", flush=True)\n\n# ---------------- 4. emit: code lane first (guaranteed), then clean backbone --------\n# Code lane: every tag-code doc with a minimal amount of English prose, best-first.\ncode_ok = (tag==1) & (sc > -50) # has words + some target signal\ncode_order = np.argsort(-np.where(code_ok, sc, -1e9))\ncode_ids = [int(ids[j]) for j in code_order if code_ok[j]]\n# Backbone: gated clean-English docs, best-first (excluding ones already in code lane).\nfinal = np.where(gate==1, sc, -1e9)\nback_order = np.argsort(-final)\nseen = set(code_ids)\nsel=[]; tot=0\nfor i in code_ids: # code lane at the front\n sel.append(i); tot += int(ntok[np.searchsorted(ids, i)]) if False else 0\n# recompute code-lane token totals cleanly via id->row map\nrow = {int(ids[k]):k for k in range(len(ids))}\nsel=list(code_ids); tot=sum(int(ntok[row[i]]) for i in sel)\nfor j in back_order: # then clean backbone\n if final[j] <= -1e8: break\n i=int(ids[j])\n if i in seen: continue\n sel.append(i); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids ({len(code_ids)} code-lane), ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)", "originalFile": "\"\"\"Curate a 12M-token pretraining selection from a raw web pool by domain-matching\nto the disclosed broad high-quality target (Wikipedia / web prose / news / tech Q&A).\n\nCriterion (stated, reproducible), CCNet-style:\n 1. GATE — drop non-English / boilerplate / word-salad / degenerate docs with cheap\n structural filters (English stop-word ratio, letter ratio, unique-word ratio,\n word length, length). The stop-word-ratio gate is the workhorse: fluent English\n prose (all four target registers) is rich in {the,of,and,to,a,in,is,...} whereas\n CSS/JS dumps, foreign text and SEO word-salad are not.\n 2. SCORE — rank survivors by a TARGET bigram language model's mean per-word\n log-probability (interpolated bigram+unigram+uniform). Documents whose wording\n is *typical of the target* score high; this rewards clean encyclopedic / news /\n web-prose / Q&A English and demotes odd-but-English filler.\n 3. A small target-vs-pool log-ratio bonus adds discrimination toward\n target-distinctive content over generic web filler.\n Emit ids best-first until ~2x the 12M-token budget so packing never underfills.\n\nThe dev target only DEFINES the target word distribution (which generalises to the\nhidden official sample); no per-id labels are used. Per-doc features are cached to\n/tmp so the ranking/threshold can be re-derived without re-reading the pool.\n\"\"\"\nimport json, re, math, sys, os, numpy as np\nfrom collections import defaultdict\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE = \"/tmp/pool_feats.npz\"\n\nMAX_WORDS = 600 # words scanned per doc for scoring (bounds cost)\nBG_SAMPLE = 40000 # pool docs for background unigram model (log-ratio bonus)\nBUDGET = 12_000_000\nOVER = 2.0 # over-provide ids to this multiple of the budget\nCHARS_PER_TOK = 4.435\nEXPLORE = \"--explore\" in sys.argv\nRATIO_W = 0.35 # weight of target-vs-pool log-ratio bonus\n\nSTOP = set(\"the of and to a in is that it for on with as was are be by this at from \"\n \"or an not but have has had he she they we you i his her their our your its \"\n \"which who will would can could there been were said more one all if them \"\n \"when so what about into than then some other time up out only over also \"\n \"no do does did how new may these two his\".split())\nword_re = re.compile(r\"[a-z][a-z']+\") # alphabetic words only (for gates + LM)\n\ndef words(text):\n return word_re.findall(text.lower())\n\n# ---------------- 1. target bigram LM + unigram, from decoded dev ----------------\nfrom transformers import AutoTokenizer\ntok = AutoTokenizer.from_pretrained(\"gpt2\")\ndev = np.load(DEV).tolist(); EOS = 50256\ndocs, cur = [], []\nfor t in dev:\n if t == EOS:\n if cur: docs.append(cur); cur = []\n else: cur.append(t)\nif cur: docs.append(cur)\ndef clean(s): return s.replace(\" @-@ \", \"-\").replace(\" @,@ \", \",\").replace(\" @.@ \", \".\")\ntgt_uni = defaultdict(int); tgt_bi = defaultdict(int); tgt_N = 0\nfor d in docs:\n ws = words(clean(tok.decode(d)))\n for j,wd in enumerate(ws):\n tgt_uni[wd] += 1; tgt_N += 1\n if j: tgt_bi[(ws[j-1], wd)] += 1\nV = len(tgt_uni)\nprint(f\"target: {len(docs)} docs, {tgt_N} words, vocab {V}\", flush=True)\n\n# ---------------- 2. pool background unigram (for log-ratio bonus) ----------------\nbg_uni = defaultdict(int); bg_N = 0; nbg = 0\nwith open(POOL) as fh:\n for line in fh:\n if nbg >= BG_SAMPLE: break\n for wd in words(json.loads(line)[\"text\"])[:MAX_WORDS]:\n bg_uni[wd] += 1; bg_N += 1\n nbg += 1\nprint(f\"background: {nbg} docs, {bg_N} words\", flush=True)\n\n# precompute unigram log-probs and log-ratio per word\nL2, L1, L0 = 0.6, 0.399, 0.001 # bigram / unigram / uniform interpolation\ndef uni_logp(wd):\n return math.log(L1 * tgt_uni.get(wd,0)/tgt_N + L0/1.0 * 1.0/ (V+1))\n# log ratio target/pool for a word (smoothed), clipped\ndef logratio(wd):\n pt = (tgt_uni.get(wd,0)+0.5)/(tgt_N+0.5*V)\n pp = (bg_uni.get(wd,0)+0.5)/(bg_N+0.5*V)\n return max(-3.0, min(3.0, math.log(pt/pp)))\n\ndef score_doc(ws):\n \"\"\"mean per-word target-LM logprob (+ mean log-ratio bonus).\"\"\"\n n = len(ws)\n if n == 0: return -99.0, 0.0\n lp = 0.0; lr = 0.0; prev = None\n for wd in ws:\n pu = tgt_uni.get(wd,0)/tgt_N\n if prev is not None:\n cb = tgt_bi.get((prev,wd),0)\n pbi = cb/tgt_uni[prev] if tgt_uni.get(prev,0) else 0.0\n else:\n pbi = 0.0\n p = L2*pbi + L1*pu + L0*(1.0/(V+1))\n lp += math.log(p)\n lr += logratio(wd)\n prev = wd\n return lp/n, lr/n\n\n# ---------------- 3. single pass over pool: gates + score, cache ----------------\ndef load_and_score():\n ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]\n with open(POOL) as fh:\n for line in fh:\n r = json.loads(line); text = r[\"text\"]\n ws = words(text)[:MAX_WORDS]\n nw = len(ws)\n ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)\n if nw < 60:\n sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue\n head = text[:6000]\n alpha_frac = sum(c.isalpha() for c in head)/max(1,len(head))\n stop_ratio = sum(w in STOP for w in ws)/nw\n uniq_ratio = len(set(ws))/nw\n mwl = sum(len(w) for w in ws)/nw\n ok = (alpha_frac>0.60 and stop_ratio>0.20 and stop_ratio<0.65\n and uniq_ratio>0.34 and 3.0<=mwl<=9.0)\n lp, lr = score_doc(ws)\n sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)\n gate.append(1 if ok else 0)\n return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),\n np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))\n\nids, sc, lpar, lrar, ntok, gate = load_and_score()\nnp.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)\nprint(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)\n\n# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------\nfinal = np.where(gate==1, sc, -1e9)\norder = np.argsort(-final)\nsel=[]; tot=0\nfor j in order:\n if final[j] <= -1e8: break\n sel.append(int(ids[j])); tot += int(ntok[j])\n if tot >= OVER*BUDGET: break\njson.dump(sel, open(OUT,\"w\"))\nprint(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)\n\nif EXPLORE:\n id2text={}\n with open(POOL) as fh:\n for line in fh:\n r=json.loads(line); id2text[r[\"id\"]]=r[\"text\"]\n def show(rank_idxs, tag):\n for j in rank_idxs:\n i=int(ids[j])\n print(f\"[{tag} sc {sc[j]:.3f} lp {lpar[j]:.3f} lr {lrar[j]:.3f} id {i}] {id2text[i][:200].strip()[:200]!r}\")\n top=order[:20000]\n codef=np.mean([(\"<code>\" in id2text[int(ids[j])] or \"<pre>\" in id2text[int(ids[j])]) for j in top])\n print(f\"\\ntop-20k code/pre-frac={codef:.3f}\")\n print(\"\\n== TOP 8 ==\"); show(order[:8],\"top\")\n print(\"\\n== rank ~8000 ==\"); show(order[8000:8003],\"mid\")\n print(\"\\n== rank ~16000 ==\"); show(order[16000:16003],\"lo\")\n g=sc[gate==1]\n print(f\"\\ngated scores: p50 {np.percentile(g,50):.3f} p90 {np.percentile(g,90):.3f} p99 {np.percentile(g,99):.3f}\")\n", "structuredPatch": [{"oldStart": 105, "oldLines": 13, "newStart": 105, "newLines": 17, "lines": [" ", " # ---------------- 3. single pass over pool: gates + score, cache ----------------", " def load_and_score():", "- ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]", "+ ids=[]; sc=[]; lpar=[]; lrar=[]; ntok=[]; gate=[]; tag=[]", " with open(POOL) as fh:", " for line in fh:", " r = json.loads(line); text = r[\"text\"]", "+ # technical/code Q&A lane: HTML/code tags are the only reliable signal for", "+ # the target's ~25%-of-tokens code register, which is scarce in the pool.", "+ is_tag = (\"<code>\" in text or \"<pre>\" in text or \"</p>\" in text)", " ws = words(text)[:MAX_WORDS]", " nw = len(ws)", " ids.append(r[\"id\"]); ntok.append(int(len(text)/CHARS_PER_TOK)+1)", "+ tag.append(1 if is_tag else 0)", " if nw < 60:", " sc.append(-99.0); lpar.append(-99.0); lrar.append(0.0); gate.append(0); continue", " head = text[:6000]"]}, {"oldStart": 125, "oldLines": 22, "newStart": 129, "newLines": 36, "lines": [" sc.append(lp + RATIO_W*lr); lpar.append(lp); lrar.append(lr)", " gate.append(1 if ok else 0)", " return (np.array(ids), np.array(sc,dtype=np.float32), np.array(lpar,dtype=np.float32),", "- np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8))", "+ np.array(lrar,dtype=np.float32), np.array(ntok), np.array(gate,dtype=np.int8),", "+ np.array(tag,dtype=np.int8))", " ", "-ids, sc, lpar, lrar, ntok, gate = load_and_score()", "-np.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate)", "-print(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}\", flush=True)", "+ids, sc, lpar, lrar, ntok, gate, tag = load_and_score()", "+np.savez(CACHE, ids=ids, sc=sc, lpar=lpar, lrar=lrar, ntok=ntok, gate=gate, tag=tag)", "+print(f\"scored {len(ids)} docs; gate-pass {int(gate.sum())}; tag-code {int(tag.sum())}\", flush=True)", " ", "-# ---------------- 4. rank gated docs best-first, emit until ~OVER*budget ----------", "+# ---------------- 4. emit: code lane first (guaranteed), then clean backbone --------", "+# Code lane: every tag-code doc with a minimal amount of English prose, best-first.", "+code_ok = (tag==1) & (sc > -50) # has words + some target signal", "+code_order = np.argsort(-np.where(code_ok, sc, -1e9))", "+code_ids = [int(ids[j]) for j in code_order if code_ok[j]]", "+# Backbone: gated clean-English docs, best-first (excluding ones already in code lane).", " final = np.where(gate==1, sc, -1e9)", "-order = np.argsort(-final)", "+back_order = np.argsort(-final)", "+seen = set(code_ids)", " sel=[]; tot=0", "-for j in order:", "+for i in code_ids: # code lane at the front", "+ sel.append(i); tot += int(ntok[np.searchsorted(ids, i)]) if False else 0", "+# recompute code-lane token totals cleanly via id->row map", "+row = {int(ids[k]):k for k in range(len(ids))}", "+sel=list(code_ids); tot=sum(int(ntok[row[i]]) for i in sel)", "+for j in back_order: # then clean backbone", " if final[j] <= -1e8: break", "- sel.append(int(ids[j])); tot += int(ntok[j])", "+ i=int(ids[j])", "+ if i in seen: continue", "+ sel.append(i); tot += int(ntok[j])", " if tot >= OVER*BUDGET: break", " json.dump(sel, open(OUT,\"w\"))", "-print(f\"selection: {len(sel)} ids, ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)", "+print(f\"selection: {len(sel)} ids ({len(code_ids)} code-lane), ~{tot/1e6:.1f}M est tokens -> {OUT}\", flush=True)", " ", " if EXPLORE:", " id2text={}"]}], "userModified": false, "replaceAll": false}Now let me overlap: launch v3 curate (CPU, ~3min) while the random baseline trains (GPU).
args
{
"command": "cd /workspace; echo \"=== random dev so far ===\"; grep -Ev \"sequence length\" /tmp/dev_rand.log | tail -3\n# launch v3 curate (writes selection.json + cache)\nnohup python3 submission/curate.py > /tmp/curate_v3.log 2>&1 &\necho \"v3 curate launched pid $!\"; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Check random eval, launch v3 curate"
}Bash result
=== random dev so far ===
packed 12000000 tokens from 40000 ids -> /tmp/dev_train.npy
v3 curate launched pid 2069
remaining: 4088
[stdout]
=== random dev so far ===
packed 12000000 tokens from 40000 ids -> /tmp/dev_train.npy
v3 curate launched pid 2069
remaining: 4088
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"taskId": "2",
"status": "completed"
}TaskUpdate result
Updated task #2 status
[metadata] {"success": true, "taskId": "2", "updatedFields": ["status"], "statusChange": {"from": "in_progress", "to": "completed"}}args
{
"taskId": "3",
"status": "in_progress"
}TaskUpdate result
Updated task #3 status
[metadata] {"success": true, "taskId": "3", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 60; echo \"=== RANDOM baseline result ===\"; grep -E \"HELDOUT_PPL\" /tmp/dev_rand.log; echo \"=== v3 curate so far ===\"; grep -Ev \"sequence length\" /tmp/curate_v3.log; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Get random baseline PPL and v3 progress"
}Bash result
=== RANDOM baseline result ===
HELDOUT_PPL 457.226 loss 6.1252 train_s 56 /tmp/dev_train.npy
=== v3 curate so far ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
remaining: 4028
[stdout]
=== RANDOM baseline result ===
HELDOUT_PPL 457.226 loss 6.1252 train_s 56 /tmp/dev_train.npy
=== v3 curate so far ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
remaining: 4028
[metadata] {"noOutputExpected": false}Major finding: random baseline = 457, my v2 = 538 — my top-k selection is WORSE than random. The target is broad (4 registers); aggressive top-k on one quality axis over-concentrated on a narrow "clean common-word" slice and lost the diversity a broad target needs. Random wins by covering all registers.
This flips the strategy: remove only clear junk, preserve diversity, match register proportions — not top-k. Let me wait for v3's cache (has all per-doc features), then build selection variants from it instantly (no re-scoring) and test the key hypothesis: does junk-removal + diversity beat random?
args
{
"command": "sleep 45; echo \"=== v3 curate ===\"; grep -Ev \"sequence length\" /tmp/curate_v3.log; echo \"cache:\"; ls -la /tmp/pool_feats.npz 2>/dev/null; echo \"remaining: $(cat /workspace/.timer/remaining_secs)\"",
"description": "Wait for v3 curate + cache"
}Bash result
=== v3 curate ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
scored 182016 docs; gate-pass 142940; tag-code 192
selection: 26263 ids (192 code-lane), ~24.0M est tokens -> /workspace/submission/selection.json
cache:
-rw-r--r-- 1 root root 5462174 Aug 4 03:47 /tmp/pool_feats.npz
remaining: 3928
[stdout]
=== v3 curate ===
target: 2705 docs, 637233 words, vocab 37566
background: 40000 docs, 13007194 words
scored 182016 docs; gate-pass 142940; tag-code 192
selection: 26263 ids (192 code-lane), ~24.0M est tokens -> /workspace/submission/selection.json
cache:
-rw-r--r-- 1 root root 5462174 Aug 4 03:47 /tmp/pool_feats.npz
remaining: 3928
[metadata] {"noOutputExpected": false}