claude-opus-4-8
nano-gpt-data-curation — graded by the task's verifier against a measured baseline.
- backend
- modal
- variant
- claude claude-opus-4-8
- metric
- held-out multi-domain perplexity
- actions
- 102
- duration
- 1h 9m
- recorded
- 2026-08-04
run iddclm-nano-curation__TGPvq9r
Everything claude-opus-4-8 did on this task, in order.
download JSON ↗I'll start by understanding the environment, the data, and the constraints. Let me gather key information in parallel.
args
{
"command": "cat /workspace/.timer/remaining_secs 2>/dev/null; echo \"---timer above---\"; ls -la /workspace/; echo \"---workspace---\"; ls -la /workspace/data/; echo \"---data---\"; ls -la /workspace/submission/ 2>/dev/null; echo \"---submission---\"",
"description": "Check timer and list workspace contents"
}Bash result
5400
---timer above---
total 14
drwxr-xr-x 1 root root 100 Aug 4 03:22 .
drwxr-xr-x 1 root root 160 Aug 4 03:22 ..
drwxr-xr-x 1 root root 60 Aug 4 03:22 .live
drwxr-xr-x 1 root root 100 Aug 4 03:22 .metrics
drwxr-xr-x 1 root root 140 Aug 4 03:22 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---workspace---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---data---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
---submission---
[stdout]
5400
---timer above---
total 14
drwxr-xr-x 1 root root 100 Aug 4 03:22 .
drwxr-xr-x 1 root root 160 Aug 4 03:22 ..
drwxr-xr-x 1 root root 60 Aug 4 03:22 .live
drwxr-xr-x 1 root root 100 Aug 4 03:22 .metrics
drwxr-xr-x 1 root root 140 Aug 4 03:22 .timer
drwxr-xr-x 1 root root 57 Aug 3 22:37 data
-rwxrwxr-x 1 root root 714 Aug 3 22:34 dev_eval.sh
-rw-rw-r-- 1 root root 5706 Aug 3 22:34 model.py
-rw-rw-r-- 1 root root 869 Aug 3 22:34 pack_selection.py
drwxr-xr-x 1 root root 10 Aug 3 22:37 submission
-rwxrwxr-x 1 root root 1509 Aug 3 22:34 timer.sh
-rw-rw-r-- 1 root root 2962 Aug 3 22:34 train_nano.py
---workspace---
total 792657
drwxr-xr-x 1 root root 57 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
-r--r--r-- 1 root root 2000128 Aug 3 22:34 multi_dev.npy
-r--r--r-- 1 root root 809679384 Aug 3 22:35 pool.jsonl
---data---
total 1
drwxr-xr-x 1 root root 10 Aug 3 22:37 .
drwxr-xr-x 1 root root 100 Aug 4 03:22 ..
---submission---
[metadata] {"noOutputExpected": false}args
{
"command": "head -c 2000 /workspace/data/pool.jsonl; echo; echo \"---head above---\"; wc -l /workspace/data/pool.jsonl",
"description": "Preview pool data and count lines"
}Bash result
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
---head above---
182016 /workspace/data/pool.jsonl
[stdout]
{"id": 0, "text": "|Viewing Single Post From: Spoilers for the Week of February 11th|\n|Lil||Feb 1 2013, 09:58 AM|\nDon't care about Chloe/Taniel/Jen-Jen. Don't care about Sami, really, but hoping that we get some good \"SAMANTHA GENE!!\" Marlena Death-Stares out of it. And \"newfound\" feelings. Please. If only.\nSTEFANO!! STEFANO, STEFANO, STEFANO!!!! :cheer:\n|Spoilers for the Week of February 11th \u00b7 DAYS: News, Spoilers & Discussion|"}
{"id": 1, "text": "*sigh* Fundamentalist community, let me pass on some advice to you I learned from the atheistic community:\nIf you have set yourself on fire, do not run.\nOkay? Okay?? Please?\nLook, D, you had two months to say to Harvard in private emails, \"Im sorry, I shouldnt have been using that animation in my paid presentations. I wont use it again. I really do like 'Inner Life', though, and would love to use it in classroom presentations, from the BioVisions site, if that is acceptable.\"\nI sat here, for two months, waiting for that to happen, anything to happen, and it didnt. Two months, on your own terms, you could have written a similar post to yesterdays. I would have given you the benefit of the doubt-- maybe you didnt know the credits werent visible to the audience, and I wouldnt have said a word beyond this, as its Harvards problem, not mine. This would have been a funny joke to those of us involved in dealing with you people, but it would have been a PR non-issue for you.\nBut after you set yourself on fire, you didnt douse it out with a bucket of ice cold reality and accountability. You ran. And youre still running.\nWhy not just state \"I screwed up. Sorry everyone.\" and move on? Why the excuses? Why the denial? Why the passive language? Why the vague words and cryptic capitalizations? Why the writes and rewrites of your 'press release'? We know it wasnt written of your own volition, or it would have been done *before* Harvard had to take action. And, your behavior before this, regarding this issue, is
---head above---
182016 /workspace/data/pool.jsonl
[metadata] {"noOutputExpected": false}args
{
"command": "cat /workspace/train_nano.py 2>/dev/null | head -200",
"description": "Read training script"
}Bash result
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
[stdout]
"""From-scratch nano-GPT (30M) LM training on a 1D token .npy; report held-out perplexity.
Frozen recipe — the ONLY thing that varies between runs is the training token set."""
import argparse, math, time, numpy as np, torch
from model import GPT, GPTConfig
ap = argparse.ArgumentParser()
ap.add_argument("--train_npy", required=True)
ap.add_argument("--heldout_npy", required=True)
ap.add_argument("--out_json", required=True)
ap.add_argument("--max_iters", type=int, default=3000)
ap.add_argument("--warmup", type=int, default=150)
ap.add_argument("--batch", type=int, default=32)
ap.add_argument("--block", type=int, default=256)
ap.add_argument("--lr", type=float, default=6e-4)
ap.add_argument("--seed", type=int, default=1337)
a = ap.parse_args()
torch.manual_seed(a.seed); np.random.seed(a.seed)
dev = "cuda"
tr = np.load(a.train_npy); ho = np.load(a.heldout_npy)
tr = torch.from_numpy(tr.astype(np.int64)); ho = torch.from_numpy(ho.astype(np.int64))
rng = np.random.default_rng(a.seed)
model = GPT(GPTConfig(block_size=a.block, vocab_size=50257, n_layer=6, n_head=6,
n_embd=384, dropout=0.0, bias=False)).to(dev)
opt = model.configure_optimizers(0.1, a.lr, (0.9, 0.95), "cuda")
def lr_at(it):
if it < a.warmup: return a.lr * (it + 1) / (a.warmup + 1)
r = (it - a.warmup) / max(1, a.max_iters - a.warmup)
return 0.1 * a.lr + 0.5 * (1 + math.cos(math.pi * r)) * (a.lr - 0.1 * a.lr)
def get_batch(src):
ix = rng.integers(0, len(src) - a.block - 1, size=a.batch)
x = torch.stack([src[i:i+a.block] for i in ix]).to(dev)
y = torch.stack([src[i+1:i+1+a.block] for i in ix]).to(dev)
return x, y
t0 = time.time()
model.train()
for it in range(a.max_iters):
for g in opt.param_groups: g["lr"] = lr_at(it)
x, y = get_batch(tr)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, loss = model(x, y)
opt.zero_grad(set_to_none=True); loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0); opt.step()
train_s = time.time() - t0
# held-out perplexity: mean next-token CE over non-overlapping windows (window-weighted)
import json
model.eval()
nwin = (len(ho) - 1) // a.block
starts = [j * a.block for j in range(nwin)]
sum_loss, nseen = 0.0, 0
with torch.no_grad():
for i in range(0, nwin, a.batch):
bs = starts[i:i+a.batch]
x = torch.stack([ho[s:s+a.block] for s in bs]).to(dev)
y = torch.stack([ho[s+1:s+1+a.block] for s in bs]).to(dev)
with torch.autocast("cuda", dtype=torch.bfloat16):
_, l = model(x, y)
sum_loss += l.item() * len(bs); nseen += len(bs)
mean_loss = sum_loss / max(1, nseen)
ppl = math.exp(mean_loss)
out = {"train_npy": a.train_npy, "heldout_ppl": ppl, "mean_loss": mean_loss,
"max_iters": a.max_iters, "seed": a.seed, "train_s": round(train_s, 1)}
json.dump(out, open(a.out_json, "w"))
print(f"HELDOUT_PPL {ppl:.3f} loss {mean_loss:.4f} train_s {train_s:.0f} {a.train_npy}")
[metadata] {"noOutputExpected": false}90 minutes total budget (5400s). Training runs cost ~60s each out of that, so I need to be economical with runs. Let me read the remaining scaffolding and inspect the dev target.
args
{
"command": "echo \"=== pack_selection.py ===\"; cat /workspace/pack_selection.py; echo; echo \"=== dev_eval.sh ===\"; cat /workspace/dev_eval.sh; echo; echo \"=== timer.sh ===\"; cat /workspace/timer.sh",
"description": "Read packing, eval, timer scripts"
}Bash result
=== pack_selection.py ===
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
=== timer.sh ===
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[stdout]
=== pack_selection.py ===
"""Solver-side: pack a selection.json (ordered ids) into a 12M-token train array from pool.jsonl."""
import json, sys, numpy as np
from transformers import AutoTokenizer
POOL = "/workspace/data/pool.jsonl"
SEL = sys.argv[1] if len(sys.argv) > 1 else "/workspace/submission/selection.json"
OUT = sys.argv[2] if len(sys.argv) > 2 else "/tmp/dev_train.npy"
BUDGET = 12_000_000
tok = AutoTokenizer.from_pretrained("gpt2"); EOS = tok.eos_token_id
text = {}
for line in open(POOL):
r = json.loads(line); text[r["id"]] = r["text"]
sel = json.load(open(SEL))
parts, tot = [], 0
for i in sel:
ids = tok(text[i], add_special_tokens=False).input_ids
parts.extend(ids); parts.append(EOS); tot += len(ids) + 1
if tot >= BUDGET: break
arr = np.array(parts[:BUDGET], dtype=np.uint16)
np.save(OUT, arr)
print(f"packed {len(arr)} tokens from {len(sel)} ids -> {OUT}")
=== dev_eval.sh ===
#!/bin/bash
# Diagnostic: train the frozen 30M GPT on your selection and report DEV perplexity.
# Usage: bash dev_eval.sh [selection.json] (default /workspace/submission/selection.json)
# NOTE: the official target is a DIFFERENT (hidden) sample from the same HQ domain.
set -euo pipefail
export TOKENIZERS_PARALLELISM=false
SEL="${1:-/workspace/submission/selection.json}"
cd /workspace
python3 pack_selection.py "$SEL" /tmp/dev_train.npy
python3 train_nano.py --train_npy /tmp/dev_train.npy --heldout_npy /workspace/data/multi_dev.npy \
--out_json /tmp/dev.json --max_iters 3000 --seed 1337 2>&1 | grep HELDOUT_PPL
echo "(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)"
=== timer.sh ===
#!/usr/bin/env bash
# Budget timer daemon — writes /workspace/.timer/remaining_secs every 10s so the
# solver can poll time left instead of doing wall-clock arithmetic.
#
# Budget (seconds) comes from $TASK_BUDGET_SECS (set by the runner to the agent
# timeout). Idempotent via a PID file. On ScaleTrain the trial runner starts its
# own timer; this file is the fallback for the Modal/local harbor backends, which
# bootstrap /workspace/timer.sh on environment start. The fallback default below
# (7200 = 2 h) matches this task's [agent] timeout_sec and is only used if
# $TASK_BUDGET_SECS is unset.
set -u
TIMER_DIR="/workspace/.timer"
PID_FILE="$TIMER_DIR/timer.pid"
mkdir -p "$TIMER_DIR"
if [ -s "$PID_FILE" ] && kill -0 "$(cat "$PID_FILE" 2>/dev/null)" 2>/dev/null; then
exit 0 # already running
fi
echo $$ > "$PID_FILE"
trap 'rm -f "$PID_FILE"' EXIT INT TERM
START=$(date +%s)
BUDGET="${TASK_BUDGET_SECS:-7200}"
echo "$START" > "$TIMER_DIR/start_epoch"
echo "$BUDGET" > "$TIMER_DIR/budget_secs"
while true; do
NOW=$(date +%s)
ELAPSED=$((NOW - START))
REMAINING=$((BUDGET - ELAPSED))
[ "$REMAINING" -lt 0 ] && REMAINING=0
echo "$REMAINING" > "$TIMER_DIR/remaining_secs"
echo "$ELAPSED" > "$TIMER_DIR/elapsed_secs"
[ "$REMAINING" -le 1800 ] && [ ! -f "$TIMER_DIR/alert_30min" ] && touch "$TIMER_DIR/alert_30min"
[ "$REMAINING" -le 600 ] && [ ! -f "$TIMER_DIR/alert_10min" ] && touch "$TIMER_DIR/alert_10min"
[ "$REMAINING" -le 0 ] && break
sleep 10
done
[metadata] {"noOutputExpected": false}args
{
"command": "cat /workspace/model.py",
"description": "Read model definition"
}Bash result
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
[stdout]
"""A small GPT (decoder-only Transformer) used as the distillation *student*.
Self-contained re-implementation in this repo's own idiom (no external GPT code).
GPT-2 byte-level BPE vocabulary (50257) so the student shares the teacher's
tokenizer. Standard pre-norm Transformer with weight-tied embeddings.
"""
from dataclasses import dataclass
import math
import torch
import torch.nn as nn
import torch.nn.functional as F
@dataclass
class GPTConfig:
block_size: int = 256
vocab_size: int = 50257
n_layer: int = 6
n_head: int = 6
n_embd: int = 384
dropout: float = 0.0
bias: bool = False
class LayerNorm(nn.Module):
def __init__(self, ndim, bias):
super().__init__()
self.weight = nn.Parameter(torch.ones(ndim))
self.bias = nn.Parameter(torch.zeros(ndim)) if bias else None
def forward(self, x):
return F.layer_norm(x, self.weight.shape, self.weight, self.bias, 1e-5)
class CausalSelfAttention(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
assert cfg.n_embd % cfg.n_head == 0
self.c_attn = nn.Linear(cfg.n_embd, 3 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.n_head = cfg.n_head
self.n_embd = cfg.n_embd
self.dropout = cfg.dropout
def forward(self, x):
B, T, C = x.size()
q, k, v = self.c_attn(x).split(self.n_embd, dim=2)
q = q.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
k = k.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
v = v.view(B, T, self.n_head, C // self.n_head).transpose(1, 2)
y = F.scaled_dot_product_attention(
q, k, v, dropout_p=self.dropout if self.training else 0.0, is_causal=True
)
y = y.transpose(1, 2).contiguous().view(B, T, C)
return self.c_proj(y)
class MLP(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.c_fc = nn.Linear(cfg.n_embd, 4 * cfg.n_embd, bias=cfg.bias)
self.c_proj = nn.Linear(4 * cfg.n_embd, cfg.n_embd, bias=cfg.bias)
self.dropout = nn.Dropout(cfg.dropout)
def forward(self, x):
return self.dropout(self.c_proj(F.gelu(self.c_fc(x))))
class Block(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.ln_1 = LayerNorm(cfg.n_embd, cfg.bias)
self.attn = CausalSelfAttention(cfg)
self.ln_2 = LayerNorm(cfg.n_embd, cfg.bias)
self.mlp = MLP(cfg)
def forward(self, x):
x = x + self.attn(self.ln_1(x))
x = x + self.mlp(self.ln_2(x))
return x
class GPT(nn.Module):
def __init__(self, cfg: GPTConfig):
super().__init__()
self.cfg = cfg
self.transformer = nn.ModuleDict(dict(
wte=nn.Embedding(cfg.vocab_size, cfg.n_embd),
wpe=nn.Embedding(cfg.block_size, cfg.n_embd),
drop=nn.Dropout(cfg.dropout),
h=nn.ModuleList([Block(cfg) for _ in range(cfg.n_layer)]),
ln_f=LayerNorm(cfg.n_embd, cfg.bias),
))
self.lm_head = nn.Linear(cfg.n_embd, cfg.vocab_size, bias=False)
self.transformer.wte.weight = self.lm_head.weight # weight tying
self.apply(self._init_weights)
# GPT-2 style scaled init on residual projections
for pn, p in self.named_parameters():
if pn.endswith('c_proj.weight'):
nn.init.normal_(p, mean=0.0, std=0.02 / math.sqrt(2 * cfg.n_layer))
def _init_weights(self, module):
if isinstance(module, nn.Linear):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
if module.bias is not None:
nn.init.zeros_(module.bias)
elif isinstance(module, nn.Embedding):
nn.init.normal_(module.weight, mean=0.0, std=0.02)
def forward(self, idx, targets=None):
"""Return (logits, ce_loss). logits: [B, T, vocab]. ce_loss is the plain
next-token cross-entropy if `targets` given, else None. The distillation
loss is computed separately in kd_trainer.kd_loss."""
B, T = idx.size()
pos = torch.arange(0, T, dtype=torch.long, device=idx.device)
x = self.transformer.drop(self.transformer.wte(idx) + self.transformer.wpe(pos))
for block in self.transformer.h:
x = block(x)
x = self.transformer.ln_f(x)
logits = self.lm_head(x)
ce_loss = None
if targets is not None:
ce_loss = F.cross_entropy(
logits.view(-1, logits.size(-1)), targets.reshape(-1), ignore_index=-1
)
return logits, ce_loss
def configure_optimizers(self, weight_decay, lr, betas, device_type):
decay, no_decay = [], []
for n, p in self.named_parameters():
if not p.requires_grad:
continue
(decay if p.dim() >= 2 else no_decay).append(p)
groups = [
{'params': decay, 'weight_decay': weight_decay},
{'params': no_decay, 'weight_decay': 0.0},
]
fused = device_type == 'cuda'
return torch.optim.AdamW(groups, lr=lr, betas=betas, fused=fused)
@torch.no_grad()
def estimate_ppl(self, get_eval_batch, n_batches):
"""Mean per-token NLL over `n_batches` eval batches -> perplexity."""
was_training = self.training
self.eval()
losses = torch.zeros(n_batches)
for i in range(n_batches):
x, y = get_eval_batch()
_, loss = self(x, y)
losses[i] = loss.item()
if was_training:
self.train()
return math.exp(losses.mean().item())
[metadata] {"noOutputExpected": false}args
{
"command": "python3 -c \"\nimport numpy as np\na = np.load('data/multi_dev.npy')\nprint('shape', a.shape, 'dtype', a.dtype, 'min', a.min(), 'max', a.max())\nprint('first 50', a[:50].tolist())\n\"",
"description": "Inspect dev target npy"
}Bash result
shape (1000000,) dtype uint16 min 0 max 50256
first 50 [41698, 3549, 373, 17338, 12617, 351, 41149, 10565, 284, 2897, 3176, 1104, 837, 475, 584, 10976, 8302, 1327, 284, 1282, 416, 764, 15933, 837, 287, 3945, 41435, 837, 41149, 10565, 5545, 284, 262, 8111, 33636, 7023, 465, 3352, 329, 281, 33149, 22817, 837, 262, 3307, 286, 543, 837, 739, 262]
[stdout]
shape (1000000,) dtype uint16 min 0 max 50256
first 50 [41698, 3549, 373, 17338, 12617, 351, 41149, 10565, 284, 2897, 3176, 1104, 837, 475, 584, 10976, 8302, 1327, 284, 1282, 416, 764, 15933, 837, 287, 3945, 41435, 837, 41149, 10565, 5545, 284, 262, 8111, 33636, 7023, 465, 3352, 329, 281, 33149, 22817, 837, 262, 3307, 286, 543, 837, 739, 262]
[metadata] {"noOutputExpected": false}args
{
"command": "python3 -c \"import tiktoken; print('tiktoken ok')\" 2>&1 | head; python3 -c \"import torch; print('torch', torch.__version__, 'cuda', torch.cuda.is_available(), torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'no gpu')\" 2>&1 | head; nvidia-smi --query-gpu=memory.total,memory.used --format=csv 2>/dev/null | head",
"description": "Check tiktoken, torch, GPU"
}Bash result
Traceback (most recent call last):
File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'tiktoken'
torch 2.8.0+cu128 cuda True NVIDIA H100 80GB HBM3
memory.total [MiB], memory.used [MiB]
81559 MiB, 4 MiB
[stdout]
Traceback (most recent call last):
File "<string>", line 1, in <module>
ModuleNotFoundError: No module named 'tiktoken'
torch 2.8.0+cu128 cuda True NVIDIA H100 80GB HBM3
memory.total [MiB], memory.used [MiB]
81559 MiB, 4 MiB
[metadata] {"noOutputExpected": false}Key facts established: 90 min budget, H100 GPU, 182K docs (~200M tokens), need to fill 12M tokens (~6%). The packer uses HF gpt2 tokenizer and inserts EOS between docs. The dev target is 1M GPT-2 tokens I can decode to text. Let me decode the target and sample the pool to understand what "high-quality" looks like here.
args
{
"command": "cd /workspace; free -g | head -3; echo \"---\"; python3 -c \"\nimport transformers, sklearn, scipy, numpy\nprint('transformers', transformers.__version__)\nprint('sklearn', sklearn.__version__)\nprint('numpy', numpy.__version__)\n\" 2>&1 | head; echo \"--- nproc ---\"; nproc",
"description": "Check RAM, packages, cores"
}Bash result
total used free shared buff/cache available
Mem: 1024 0 1023 0 0 1023
Swap: 0 0 0
---
Traceback (most recent call last):
File "<string>", line 2, in <module>
ModuleNotFoundError: No module named 'sklearn'
--- nproc ---
16
[stdout]
total used free shared buff/cache available
Mem: 1024 0 1023 0 0 1023
Swap: 0 0 0
---
Traceback (most recent call last):
File "<string>", line 2, in <module>
ModuleNotFoundError: No module named 'sklearn'
--- nproc ---
16
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np\ntok = AutoTokenizer.from_pretrained('gpt2')\na = np.load('data/multi_dev.npy')\n# decode in chunks separated by EOS (50256)\neos = tok.eos_token_id\nprint('EOS id', eos, 'count in dev', int((a==eos).sum()))\n# split on EOS\nimport numpy as np\nidx = np.where(a==eos)[0]\nprint('num docs in dev (approx segments):', len(idx))\n# decode first few segments\nprev=0\nsegs=[]\nfor j in idx[:8]:\n segs.append(tok.decode(a[prev:j].tolist()))\n prev=j+1\nfor i,s in enumerate(segs):\n print('==== DEV SEG', i, 'len_chars', len(s), '====')\n print(s[:600].replace(chr(10),' '))\n print()\n\"",
"description": "Decode dev target segments to understand target"
}Bash result
EOS id 50256 count in dev 2704
num docs in dev (approx segments): 2704
==== DEV SEG 0 len_chars 1043 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
==== DEV SEG 1 len_chars 603 ====
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
==== DEV SEG 2 len_chars 1037 ====
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
==== DEV SEG 3 len_chars 435 ====
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused .
==== DEV SEG 4 len_chars 889 ====
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
==== DEV SEG 5 len_chars 603 ====
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
==== DEV SEG 6 len_chars 112 ====
In 1910 , Shackleton made a series of three recordings describing the expedition using an Edison Phonograph .
==== DEV SEG 7 len_chars 364 ====
Several mostly intact cases of whisky and brandy left behind in 1909 were recovered in 2010 , for analysis by a distilling company . A revival of the vintage ( and since lost ) formula for the particular brands found has been offered for sale with a portion of the proceeds to benefit the New Zealand Antarctic Heritage Trust which discovered the lost spirits .
[stdout]
EOS id 50256 count in dev 2704
num docs in dev (approx segments): 2704
==== DEV SEG 0 len_chars 1043 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The aim was the conquest of both the geographical South Pole and the South Magnetic Pole . Shackleton then worked hard to persuade others of his wealthy friends and acquaintances to contribute , includin
==== DEV SEG 1 len_chars 603 ====
On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attempts on the South Pole and South Magnetic Pole . However , before leaving England , he had been pressured to give an undertaking to Scott that he would not base himself in the McMurdo area , which Scott was claiming as his own field of work . Shackleton reluctantly agreed to look for winter quarters at either the Barrier Inlet ( which Discovery had briefly visited in 1902 ) or King Edward VII Land
==== DEV SEG 2 len_chars 1037 ====
To conserve coal , the ship was towed 1 @,@ 650 miles ( 2 @,@ 655 km ) by the steamer Koonya to the Antarctic ice , after Shackleton had persuaded the New Zealand government and the Union Steamship Company to share the cost . In accordance with Shackleton 's promise to Scott , the ship headed for the eastern sector of the Great Ice Barrier , arriving there on 21 January 1908 . They found that the Barrier Inlet had expanded to form a large bay , in which were hundreds of whales , which led to the immediate christening of the area as the Bay of Whales . It was noted that ice conditions were uns
==== DEV SEG 3 len_chars 435 ====
Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was eventually established at Cape Royds , about 24 miles ( 39 km ) north of Hut Point . The party was in high spirits , despite the difficult conditions ; Shackleton 's ability to communicate with each man kept the party happy and focused .
==== DEV SEG 4 len_chars 889 ====
The " Great Southern Journey " , as Frank Wild called it , began on 29 October 1908 . On 9 January 1909 , Shackleton and three companions ( Wild , Eric Marshall and Jameson Adams ) reached a new Farthest South latitude of 88 ° 23 ' S , a point only 112 miles ( 180 km ) from the Pole . En route the South Pole party discovered the Beardmore Glacier ( named after Shackleton 's patron ) and became the first persons to see and travel on the South Polar Plateau . Their return journey to McMurdo Sound was a race against starvation , on half @-@ rations for much of the way . At one point , Shackleton
==== DEV SEG 5 len_chars 603 ====
The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgeworth David , Douglas Mawson , and Alistair Mackay . Shackleton returned to the United Kingdom as a hero , and soon afterwards published his expedition account , Heart of the Antarctic . Emily Shackleton later recorded : " The only comment he made to me about not reaching the Pole was ' a live donkey is better than a dead lion , isn 't it ? ' and I said ' Yes darling , as far as I am concerned ' "
==== DEV SEG 6 len_chars 112 ====
In 1910 , Shackleton made a series of three recordings describing the expedition using an Edison Phonograph .
==== DEV SEG 7 len_chars 364 ====
Several mostly intact cases of whisky and brandy left behind in 1909 were recovered in 2010 , for analysis by a distilling company . A revival of the vintage ( and since lost ) formula for the particular brands found has been offered for sale with a portion of the proceeds to benefit the New Zealand Antarctic Heritage Trust which discovered the lost spirits .
[metadata] {"noOutputExpected": false}Two important findings: the dev target segment 0-7 is clearly Wikipedia (WikiText-103 formatting — note the @,@ / @-@ markers), and I need to verify the RAM situation (the free -g output is ambiguous). Let me confirm memory and sample the other target registers (news, Q&A, web prose) spread across the dev array.
args
{
"command": "free -h; echo \"--- cgroup limit ---\"; cat /sys/fs/cgroup/memory.max 2>/dev/null || cat /sys/fs/cgroup/memory/memory.limit_in_bytes 2>/dev/null; echo; grep MemTotal /proc/meminfo",
"description": "Verify true RAM and cgroup limit"
}Bash result
total used free shared buff/cache available
Mem: 1.0Ti 199Mi 1.0Ti 0B 266Mi 1.0Ti
Swap: 0B 0B 0B
--- cgroup limit ---
1099511627776
MemTotal: 1073741824 kB
[stdout]
total used free shared buff/cache available
Mem: 1.0Ti 199Mi 1.0Ti 0B 266Mi 1.0Ti
Swap: 0B 0B 0B
--- cgroup limit ---
1099511627776
MemTotal: 1073741824 kB
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np\ntok = AutoTokenizer.from_pretrained('gpt2')\na = np.load('data/multi_dev.npy')\neos = tok.eos_token_id\nidx = np.where(a==eos)[0]\nstarts = np.concatenate([[0], idx+1])\nends = np.concatenate([idx, [len(a)]])\n# sample segments spread across the array to see different registers\nimport numpy as np\npicks = [0, 300, 700, 1000, 1300, 1600, 1900, 2200, 2500, 2690]\nfor p in picks:\n s,e = starts[p], ends[p]\n txt = tok.decode(a[s:e].tolist())\n print('==== DEV SEG', p, 'chars', len(txt), '====')\n print(txt[:400].replace(chr(10),' '))\n print()\n\"",
"description": "Sample dev segments across the array for register variety"
}Bash result
==== DEV SEG 0 chars 1043 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The
==== DEV SEG 300 chars 311 ====
" Mystery Date " received steady viewership that was consistent with the ratings for the previous week . It received 2 @.@ 8 million viewers , down only from 2 @.@ 9 from " Tea Leaves " . The episode also received a 1 @.@ 0 rating in the important 18 @-@ 49 demographic , the same rating as the week before .
==== DEV SEG 700 chars 1171 ====
In 1968 , Lennon was told The Dairy Cottage was too cramped for them all , so he told Birch to buy a house , and he found a 4 @-@ bedroom house in Gateacre Park Drive , Liverpool . Lennon told Birch to furnish and decorate it , and to send all the bills to him . The Dykinses heard nothing from Lennon for years , until he phoned Baird in 1975 , and asked for mementos of his childhood life , such a
==== DEV SEG 1000 chars 532 ====
In April 2006 , a team of astronomers , believing that Oval BA might converge with the GRS that year , observed the storms through the Hubble Space Telescope . The storms pass each other about every two years , but the passings of 2002 and 2004 did not produce anything exciting . Dr. Amy Simon @-@ Miller , of the Goddard Space Flight Center , predicted the storms would have their closest passing
==== DEV SEG 1300 chars 685 ====
In 1925 , US 112 was originally proposed to run from Oshkosh to Fremont , Wisconsin , on what later became U.S. Route 110 . When it was initially designated in November 1926 , US 112 made a sharp turn to the southwest to connect to US 20 in Elkhart , Indiana . In 1931 , a new trunkline highway was designated between M @-@ 60 at Niles and US 112 at Union . This highway was numbered M @-@ 151 . In
==== DEV SEG 1600 chars 770 ====
A solar cell , or photovoltaic cell ( PV ) , is a device that converts light into electric current using the photovoltaic effect . The first solar cell was constructed by Charles Fritts in the 1880s . The German industrialist Ernst Werner von Siemens was among those who recognized the importance of this discovery . In 1931 , the German engineer Bruno Lange developed a photo cell using silver sele
==== DEV SEG 1900 chars 2802 ====
PETALING JAYA: Times are a-changing. Blue collar foreign workers in Malaysia are climbing the ladder faster than expected by opening businesses traditionally run by locals, making it harder for youths to earn a living, said an economist. The foreign workers start off working as cashiers in clothing stores, jewellery shops, restaurants, mechanic workshops, construction businesses and selling mobil
==== DEV SEG 2200 chars 4615 ====
After five days of scouring the life of Las Vegas gunman Stephen Paddock and chasing 1,000 leads, investigators confessed Friday they still don't know what drove him to mass murder, and they announced plans to put up billboards appealing for the public's help.In their effort to find any hint of his motive, investigators were looking into whether he was with a prostitute days before the shooting, s
==== DEV SEG 2500 chars 939 ====
<p>in C# winforms when we display a message box it has no title in the title bar and no title in its button that is in the task bar. </p> <p>What if i want to set title and icon for a message box.</p> <p>one option is that create a form that appears and behaves like a message box and i show and hide it when i want. yes that can be done but i want to modify the "MessageBox"</p> <p>Use a MessageB
==== DEV SEG 2690 chars 913 ====
<p>I need a for loop which would print from year 2000 to 2099.</p> <pre><code>[Bindable] private var yearValue:Array private function we():void { var i:Number; for(i=2000;i<=2099;i++){ yearValue = new Array(i); } } <mx:ComboBox id='year' labelField="Year" dataProvider="{yearValue}"> </mx:ComboBox> </code></pre> <p>When i populate i
[stdout]
==== DEV SEG 0 chars 1043 ====
Beardmore was sufficiently impressed with Shackleton to offer financial support , but other donations proved hard to come by . Nevertheless , in February 1907 , Shackleton presented to the Royal Geographic Society his plans for an Antarctic expedition , the details of which , under the name British Antarctic Expedition , were published in the Royal Society 's newsletter , Geographic Journal . The
==== DEV SEG 300 chars 311 ====
" Mystery Date " received steady viewership that was consistent with the ratings for the previous week . It received 2 @.@ 8 million viewers , down only from 2 @.@ 9 from " Tea Leaves " . The episode also received a 1 @.@ 0 rating in the important 18 @-@ 49 demographic , the same rating as the week before .
==== DEV SEG 700 chars 1171 ====
In 1968 , Lennon was told The Dairy Cottage was too cramped for them all , so he told Birch to buy a house , and he found a 4 @-@ bedroom house in Gateacre Park Drive , Liverpool . Lennon told Birch to furnish and decorate it , and to send all the bills to him . The Dykinses heard nothing from Lennon for years , until he phoned Baird in 1975 , and asked for mementos of his childhood life , such a
==== DEV SEG 1000 chars 532 ====
In April 2006 , a team of astronomers , believing that Oval BA might converge with the GRS that year , observed the storms through the Hubble Space Telescope . The storms pass each other about every two years , but the passings of 2002 and 2004 did not produce anything exciting . Dr. Amy Simon @-@ Miller , of the Goddard Space Flight Center , predicted the storms would have their closest passing
==== DEV SEG 1300 chars 685 ====
In 1925 , US 112 was originally proposed to run from Oshkosh to Fremont , Wisconsin , on what later became U.S. Route 110 . When it was initially designated in November 1926 , US 112 made a sharp turn to the southwest to connect to US 20 in Elkhart , Indiana . In 1931 , a new trunkline highway was designated between M @-@ 60 at Niles and US 112 at Union . This highway was numbered M @-@ 151 . In
==== DEV SEG 1600 chars 770 ====
A solar cell , or photovoltaic cell ( PV ) , is a device that converts light into electric current using the photovoltaic effect . The first solar cell was constructed by Charles Fritts in the 1880s . The German industrialist Ernst Werner von Siemens was among those who recognized the importance of this discovery . In 1931 , the German engineer Bruno Lange developed a photo cell using silver sele
==== DEV SEG 1900 chars 2802 ====
PETALING JAYA: Times are a-changing. Blue collar foreign workers in Malaysia are climbing the ladder faster than expected by opening businesses traditionally run by locals, making it harder for youths to earn a living, said an economist. The foreign workers start off working as cashiers in clothing stores, jewellery shops, restaurants, mechanic workshops, construction businesses and selling mobil
==== DEV SEG 2200 chars 4615 ====
After five days of scouring the life of Las Vegas gunman Stephen Paddock and chasing 1,000 leads, investigators confessed Friday they still don't know what drove him to mass murder, and they announced plans to put up billboards appealing for the public's help.In their effort to find any hint of his motive, investigators were looking into whether he was with a prostitute days before the shooting, s
==== DEV SEG 2500 chars 939 ====
<p>in C# winforms when we display a message box it has no title in the title bar and no title in its button that is in the task bar. </p> <p>What if i want to set title and icon for a message box.</p> <p>one option is that create a form that appears and behaves like a message box and i show and hide it when i want. yes that can be done but i want to modify the "MessageBox"</p> <p>Use a MessageB
==== DEV SEG 2690 chars 913 ====
<p>I need a for loop which would print from year 2000 to 2099.</p> <pre><code>[Bindable] private var yearValue:Array private function we():void { var i:Number; for(i=2000;i<=2099;i++){ yearValue = new Array(i); } } <mx:ComboBox id='year' labelField="Year" dataProvider="{yearValue}"> </mx:ComboBox> </code></pre> <p>When i populate i
[metadata] {"noOutputExpected": false}RAM is actually 1.0 TiB — no memory constraint, I misread. Now I have a clear picture of the 4 target registers:
- Wikipedia/encyclopedic (WikiText-103 formatted,
@,@markers) - News (datelines, journalistic prose — normal punctuation)
- Technical Q&A (StackOverflow-style with
<p>/<pre><code>/<) - General HQ web prose
Let me set up task tracking and run a broad streaming analysis of the pool to see its composition and how detectable these registers are.
args
{
"subject": "Analyze pool composition & registers",
"description": "Stream pool.jsonl, compute per-doc length/char stats and register indicators (code/HTML, wiki markers, news cues). Understand what fraction is target-like vs junk.",
"activeForm": "Analyzing pool composition"
}TaskCreate result
Task #1 created successfully: Analyze pool composition & registers
[metadata] {"task": {"id": "1", "subject": "Analyze pool composition & registers"}}args
{
"subject": "Build relevance+quality scorer (curate.py)",
"description": "Decode dev target as positives (normalize WikiText markers), random pool as negatives; train GPU logistic-regression classifier over hashed n-grams. Combine with heuristic quality filter. Produce register-balanced, ordered selection.json.",
"activeForm": "Building relevance+quality scorer"
}TaskCreate result
Task #2 created successfully: Build relevance+quality scorer (curate.py)
[metadata] {"task": {"id": "2", "subject": "Build relevance+quality scorer (curate.py)"}}args
{
"subject": "Validate selection via dev_eval training runs",
"description": "Run frozen train_nano on random baseline and on curated selection; compare held-out dev perplexity. Iterate on selection if time permits.",
"activeForm": "Validating selection with training runs"
}TaskCreate result
Task #3 created successfully: Validate selection via dev_eval training runs
[metadata] {"task": {"id": "3", "subject": "Validate selection via dev_eval training runs"}}args
{
"subject": "Write claim.md deliverable",
"description": "Hypothesis / Mechanism (observable other than final perplexity) / Falsification / Transfer.",
"activeForm": "Writing claim.md"
}TaskCreate result
Task #4 created successfully: Write claim.md deliverable
[metadata] {"task": {"id": "4", "subject": "Write claim.md deliverable"}}encrypted chain-of-thought — recorded, not exposed
args
{
"taskId": "1",
"status": "in_progress"
}TaskUpdate result
Updated task #1 status
[metadata] {"success": true, "taskId": "1", "updatedFields": ["status"], "statusChange": {"from": "pending", "to": "in_progress"}}args
{
"command": "cd /workspace; python3 -c \"\nimport json, re\nn=0\ntot_chars=0\nlens=[]\n# register indicator counters\nc_code=0; c_html=0; c_wiki=0; c_news=0; c_short=0; c_nonascii=0\nsamples_code=[]; samples_clean=[]; samples_junk=[]\nalpha_fracs=[]\nimport random\nrandom.seed(0)\nfor line in open('data/pool.jsonl'):\n r=json.loads(line); t=r['text']; L=len(t)\n n+=1; tot_chars+=L; \n if n<=60000: # sample for stats to keep it fast-ish but representative\n lens.append(L)\n alpha=sum(ch.isalpha() for ch in t[:2000]); alpha_fracs.append(alpha/max(1,min(2000,L)))\n has_code = ('<code>' in t) or ('</code>' in t) or ('def ' in t and '(' in t) or ('function(' in t) or ('public ' in t and '{' in t) or ('<' in t and '>' in t)\n has_html = ('<p>' in t) or ('<div' in t) or ('<a href' in t)\n has_wiki = (' @,@ ' in t) or (' @-@ ' in t) or (' @.@ ' in t)\n has_news = bool(re.match(r'^[A-Z][A-Z ]{3,30}[:\\-—]', t.strip()[:40])) or ('(Reuters)' in t) or ('(AP)' in t)\n if has_code: c_code+=1\n if has_html: c_html+=1\n if has_wiki: c_wiki+=1\n if has_news: c_news+=1\n if L<200: c_short+=1\n # collect a few samples\n if has_code and len(samples_code)<3 and n>100: samples_code.append(t[:300])\nprint('total docs', n)\nprint('total chars', tot_chars, ' ~tokens (chars/4)', tot_chars//4)\nimport numpy as np\nlens=np.array(lens); af=np.array(alpha_fracs)\nprint('char len: min %d p10 %d p50 %d p90 %d max %d mean %.0f'%(lens.min(), np.percentile(lens,10), np.percentile(lens,50), np.percentile(lens,90), lens.max(), lens.mean()))\nprint('alpha frac (first2k): p10 %.2f p50 %.2f p90 %.2f'%(np.percentile(af,10),np.percentile(af,50),np.percentile(af,90)))\nprint('--- register indicator counts (of %d) ---'%n)\nprint('has_code', c_code, '%.1f%%'%(100*c_code/n))\nprint('has_html', c_html, '%.1f%%'%(100*c_html/n))\nprint('has_wiki_markers', c_wiki, '%.1f%%'%(100*c_wiki/n))\nprint('has_news_dateline', c_news, '%.1f%%'%(100*c_news/n))\nprint('short_lt200chars', c_short, '%.1f%%'%(100*c_short/n))\nprint()\nfor i,s in enumerate(samples_code): print('CODE SAMPLE',i, repr(s[:250])); print()\n\"",
"description": "Stream pool: length/char stats and register indicator counts"
}Bash result
total docs 182016
total chars 770537151 ~tokens (chars/4) 192634287
char len: min 136 p10 529 p50 1786 p90 6035 max 522573 mean 3094
alpha frac (first2k): p10 0.76 p50 0.79 p90 0.82
--- register indicator counts (of 182016) ---
has_code 1984 1.1%
has_html 537 0.3%
has_wiki_markers 0 0.0%
has_news_dateline 1893 1.0%
short_lt200chars 2151 1.2%
CODE SAMPLE 0 'WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Related posts:Win a Trip to the St'
CODE SAMPLE 1 'Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology |\nThe Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on the web is defined, making it p'
CODE SAMPLE 2 'Toshiba Satellite A505-S6040 TruBrite 16.0-Inch Laptop (Black/Silver)February 9th, 2010 | Posted by in Toshiba\nToshiba Satellite A505-S6040 TruBrite 16.0-Inch Laptop (Black/Silver) for Sale\nFor all you want to see, hear and experience, the Satellite '
[stdout]
total docs 182016
total chars 770537151 ~tokens (chars/4) 192634287
char len: min 136 p10 529 p50 1786 p90 6035 max 522573 mean 3094
alpha frac (first2k): p10 0.76 p50 0.79 p90 0.82
--- register indicator counts (of 182016) ---
has_code 1984 1.1%
has_html 537 0.3%
has_wiki_markers 0 0.0%
has_news_dateline 1893 1.0%
short_lt200chars 2151 1.2%
CODE SAMPLE 0 'WATCH #FORUM Full Length & Free For a Limited Time videos by Lorene Voskinarian October 24th, 2012No Comments Tweet Forum is making their latest film #FORUM free to watch for a very limited time. Check it out below. Related posts:Win a Trip to the St'
CODE SAMPLE 1 'Methods | Statistics | Clinical | Educational | Industrial | Professional items | World psychology |\nThe Semantic Web is an evolving extension of the World Wide Web in which the semantics of information and services on the web is defined, making it p'
CODE SAMPLE 2 'Toshiba Satellite A505-S6040 TruBrite 16.0-Inch Laptop (Black/Silver)February 9th, 2010 | Posted by in Toshiba\nToshiba Satellite A505-S6040 TruBrite 16.0-Inch Laptop (Black/Silver) for Sale\nFor all you want to see, hear and experience, the Satellite '
[metadata] {"noOutputExpected": false}The pool is raw web text with almost no formatting markers (0% wiki markers, 1% code, 0.3% HTML) — so I can't tag pool docs by surface markers; register matching must be content/vocabulary-based. The pool samples are dominated by commercial/forum spam ("Toshiba laptop for sale", "free videos", soap-opera forums), so filtering junk is a big lever. Let me measure the target's true register mixture (token shares) from the dev set, which I'll need to match for balance.
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 -c \"\nfrom transformers import AutoTokenizer\nimport numpy as np, re\ntok = AutoTokenizer.from_pretrained('gpt2')\na = np.load('data/multi_dev.npy')\neos = tok.eos_token_id\nidx = np.where(a==eos)[0]\nstarts = np.concatenate([[0], idx+1]); ends = np.concatenate([idx, [len(a)]])\ndef label(txt):\n if ('<p>' in txt) or ('<code>' in txt) or ('<' in txt) or ('<pre>' in txt) or ('</' in txt): return 'qa'\n if (' @,@ ' in txt) or (' @-@ ' in txt) or (' @.@ ' in txt): return 'wiki'\n # news: dateline ALLCAPS: or agency tags, or leading location\n head=txt.strip()[:60]\n if re.match(r'^[A-Z][A-Za-z .]{2,25}[:\\-—]', head) or ('(Reuters)' in txt) or ('(AP)' in txt) or ('PETALING' in head): return 'news'\n return 'web'\nfrom collections import Counter\ncnt=Counter(); toks=Counter()\nsegtexts={'qa':[],'wiki':[],'news':[],'web':[]}\nfor s,e in zip(starts,ends):\n ntok=e-s\n if ntok<3: continue\n txt=tok.decode(a[s:e].tolist())\n lb=label(txt)\n cnt[lb]+=1; toks[lb]+=ntok\n segtexts[lb].append(txt)\ntot=sum(toks.values())\nprint('register #segs #tokens token_share')\nfor k in ['wiki','news','qa','web']:\n print('%-6s %5d %7d %.1f%%'%(k, cnt[k], toks[k], 100*toks[k]/tot))\nprint('total tokens (approx w/o eos):', tot)\n# show a couple 'web' and 'news' samples to verify labeling\nprint('--- WEB samples ---')\nfor t in segtexts['web'][:3]: print(repr(t[:200]))\nprint('--- NEWS samples ---')\nfor t in segtexts['news'][:3]: print(repr(t[:200]))\n\"",
"description": "Label dev segments by register and measure token shares"
}Bash result
register #segs #tokens token_share
wiki 966 163490 16.4%
news 47 36106 3.6%
qa 360 251243 25.2%
web 1332 546457 54.8%
total tokens (approx w/o eos): 997296
--- WEB samples ---
" On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attem"
" Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was even"
" The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgew"
--- NEWS samples ---
' CSI : Crime Scene Investigation executive producer and showrunner Carol Mendelsohn considers Grissom the center of the show . Grissom has received positive responses from critics , ranking number 82 '
"Authorities have had to truck water supplies to 18 of Cambodia's 25 provinces, with some 2.5 million people affected by the drought\n\nBy Prak Chan Thul\n\nPHNOM PENH, May 24 (Reuters) - As the worst drou"
'If you were a hip-hop fan who grew up in the 90s, chances are more likely than not that the dominance of Roc-a-Fella is embedded in your memory. June of this past year marked the 20th anniversary of J'
[stdout]
register #segs #tokens token_share
wiki 966 163490 16.4%
news 47 36106 3.6%
qa 360 251243 25.2%
web 1332 546457 54.8%
total tokens (approx w/o eos): 997296
--- WEB samples ---
" On 1 January 1908 , Nimrod sailed for the Antarctic from Lyttelton Harbour , New Zealand . Shackleton 's original plans had envisaged using the old Discovery base in McMurdo Sound to launch his attem"
" Nimrod arrived at McMurdo Sound on 29 January , but was stopped by ice 16 miles ( 26 km ) north of Discovery 's old base at Hut Point . After considerable weather delays , Shackleton 's base was even"
" The expedition 's other main accomplishments included the first ascent of Mount Erebus , and the discovery of the approximate location of the South Magnetic Pole , reached on 16 January 1909 by Edgew"
--- NEWS samples ---
' CSI : Crime Scene Investigation executive producer and showrunner Carol Mendelsohn considers Grissom the center of the show . Grissom has received positive responses from critics , ranking number 82 '
"Authorities have had to truck water supplies to 18 of Cambodia's 25 provinces, with some 2.5 million people affected by the drought\n\nBy Prak Chan Thul\n\nPHNOM PENH, May 24 (Reuters) - As the worst drou"
'If you were a hip-hop fan who grew up in the 90s, chances are more likely than not that the dominance of Roc-a-Fella is embedded in your memory. June of this past year marked the 20th anniversary of J'
[metadata] {"noOutputExpected": false}Now I understand the target mixture. My surface labeling is leaky (wiki-without-markers leaks into "web"), but the reliable signal is technical Q&A ≈ 25% of tokens (HTML/code markers are unambiguous), matching the "equal parts" design. The other 75% (encyclopedic + news + HQ web) is all clean informative prose.
Strategy: two content classifiers over hashed n-grams — a prose head (encyclopedic/news/web) and a technical head (programming/Q&A) — trained on decoded dev positives vs. random-pool negatives, with register-aware junk filtering, then a 75/25 token-balanced selection. Let me benchmark featurization speed before committing.
args
{
"command": "cd /workspace; python3 -c \"\nimport json, re, time, zlib\nD=1<<20\nwre=re.compile(r\\\"[a-z0-9']+\\\")\ndef feats(text, cap=1000):\n w=wre.findall(text.lower())[:cap]\n idx=set()\n for tk in w:\n idx.add(zlib.crc32(tk.encode())&(D-1))\n for i in range(len(w)-1):\n idx.add(zlib.crc32((w[i]+' '+w[i+1]).encode())&(D-1))\n return idx\nt0=time.time(); n=0; nnz=0\nfor line in open('data/pool.jsonl'):\n r=json.loads(line); s=feats(r['text']); nnz+=len(s); n+=1\n if n>=5000: break\ndt=time.time()-t0\nprint('5000 docs in %.2fs -> full 182k est %.1fs, mean nnz %.0f'%(dt, dt*182016/5000, nnz/n))\n\"",
"description": "Benchmark single-process featurization speed"
}Bash result
5000 docs in 1.50s -> full 182k est 54.7s, mean nnz 555
[stdout]
5000 docs in 1.50s -> full 182k est 54.7s, mean nnz 555
[metadata] {"noOutputExpected": false}Featurization is ~55s single-process — fast enough. Word-regex features are naturally robust to WikiText markers (punctuation produces no word tokens), so the classifier learns pure content vocabulary that transfers to raw pool text.
Let me parallelize: launch a random-baseline dev_eval in the background (uses GPU) while I build curate.py (uses CPU). First, generate a random selection and start its training run.
args
{
"command": "cd /workspace; python3 -c \"\nimport json, random\nrandom.seed(1337)\nids=list(range(182016)); random.shuffle(ids)\nsel=ids[:30000]\njson.dump(sel, open('/workspace/submission/selection_random.json','w'))\nprint('wrote random selection', len(sel))\n\"",
"description": "Create random baseline selection"
}Bash result
wrote random selection 30000
[stdout]
wrote random selection 30000
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; (bash dev_eval.sh /workspace/submission/selection_random.json > /tmp/baseline_random.log 2>&1; echo \"DONE_EXIT $?\" >> /tmp/baseline_random.log) ",
"description": "Run random baseline dev_eval in background",
"run_in_background": true
}Bash result
Command running in background with ID: beym2btkh. Output is being written to: /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/beym2btkh.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "beym2btkh", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}Baseline is training in the background. Now let me write curate.py. I'll use an NBSVM-style log-odds scorer (target-register vs. pool-background over hashed n-grams) — closed-form, fast, reproducible, no fragile hyperparameter tuning — with two heads (prose + technical) and register-aware quality filtering.
args
{
"file_path": "/workspace/curate.py",
"content": "#!/usr/bin/env python3\n\"\"\"\nCurate a raw web pool for pretraining a small LM whose eval target is a BROAD,\nmulti-domain high-quality English mixture (equal parts encyclopedic/Wikipedia,\ngeneral high-quality web prose, news, technical Q&A).\n\nSTATED CRITERION\n----------------\nKeep the pool documents that are (a) clean, fluent, information-dense English and\n(b) distributionally close to the disclosed target. We measure target-closeness\nwith an NBSVM-style log-odds score: for hashed word unigram+bigram features, weigh\neach feature by log P(feature | target register) - log P(feature | pool background).\nA document's score is the sum of its feature weights == its log-likelihood ratio of\nbeing drawn from the target register vs. the generic web pool (a DSIR-style signal).\n\nThe target is a fixed mixture whose one unambiguous slice is technical Q&A (~25% of\ntarget tokens, identifiable by HTML/code markup). Programming/Q&A prose is\nstylistically distinct from encyclopedic/news/web prose, so we score with TWO heads:\n - PROSE head: positives = decoded dev segments WITHOUT code/HTML markup\n - TECH head: positives = decoded dev segments WITH code/HTML markup\nEach pool doc is assigned to the head it looks most like, passes a register-aware\njunk filter, and is ranked by that head's score. We then fill the 12M-token budget\ntoken-balanced 75% prose / 25% tech to match the target mixture, interleaving so any\nearly budget cutoff stays balanced. Remaining ranked ids are appended as overflow.\n\nThe positives come only from the disclosed dev target; the selection is over disjoint\npool ids. Output is an ordered id list, produced entirely by this script (no hand\npicking).\n\"\"\"\nimport json, re, zlib, sys, time\nimport numpy as np\nfrom scipy import sparse\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nTECH_SHARE = 0.25 # target token share of technical Q&A register\nD = 1 << 20 # hashed feature dimension\nWORD_CAP = 1200 # cap words/doc for featurization + stats (bounds cost)\nALPHA = 1.0 # NB smoothing\nN_OVERFLOW_IDS = 45_000 # total ids to emit (well over budget for safety)\n\nwre = re.compile(r\"[a-z0-9']+\")\nSTOP = set(\"the a an and or but if then else of to in on at by for with from as is are was \"\n \"were be been being it its this that these those he she they we you i his her their \"\n \"our your not no do does did has have had will would can could should may might must \"\n \"there here what which who whom whose when where why how than so such about into over \"\n \"after before between out up down off again more most other some any each\".split())\n\ndef doc_features(text):\n \"\"\"Return (feature_index_array_uint32, stats_dict) for one document.\"\"\"\n low = text.lower()\n w = wre.findall(low)\n if len(w) > WORD_CAP:\n w = w[:WORD_CAP]\n nw = len(w)\n idx = set()\n for tk in w:\n idx.add(zlib.crc32(tk.encode()) & (D - 1))\n for i in range(nw - 1):\n idx.add(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (D - 1))\n # char-level quality stats over a bounded prefix\n seg = text[:4000]\n L = max(1, len(seg))\n alpha = sum(c.isalpha() for c in seg)\n digit = sum(c.isdigit() for c in seg)\n ascii_ = sum(ord(c) < 128 for c in seg)\n space = sum(c.isspace() for c in seg)\n sym = L - alpha - digit - space\n stops = sum(1 for x in w if x in STOP)\n uniq = len(set(w))\n mwl = (sum(len(x) for x in w) / nw) if nw else 0.0\n stats = dict(\n nw=nw,\n tok_est=max(1, len(text) // 4),\n alpha=alpha / L,\n digit=digit / L,\n ascii=ascii_ / L,\n sym=sym / L,\n stopfrac=(stops / nw) if nw else 0.0,\n uniq=(uniq / nw) if nw else 0.0,\n mwl=mwl,\n has_code=((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)\n or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)),\n )\n return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stats\n\n\ndef build_csr(list_of_idx_arrays, D):\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n for i, a in enumerate(list_of_idx_arrays):\n indptr[i + 1] = indptr[i] + len(a)\n indices = np.concatenate(list_of_idx_arrays) if list_of_idx_arrays else np.zeros(0, np.uint32)\n data = np.ones(len(indices), dtype=np.float32)\n return sparse.csr_matrix((data, indices, indptr), shape=(len(list_of_idx_arrays), D))\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):\n \"\"\"NBSVM/DSIR-style per-feature log-odds: log p(f|target) - log p(f|background).\"\"\"\n p = (pos_df + ALPHA) / (npos + 2 * ALPHA)\n q = (bg_df + ALPHA) / (nbg + 2 * ALPHA)\n return np.log(p) - np.log(q)\n\n\ndef main():\n t0 = time.time()\n print(\"[1] decoding dev target -> positive register sets\", flush=True)\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n dev = np.load(DEV)\n cut = np.where(dev == eos)[0]\n starts = np.concatenate([[0], cut + 1]); ends = np.concatenate([cut, [len(dev)]])\n pos_prose_idx, pos_tech_idx = [], []\n for s, e in zip(starts, ends):\n if e - s < 3:\n continue\n txt = tok.decode(dev[s:e].tolist())\n a, st = doc_features(txt)\n (pos_tech_idx if st[\"has_code\"] else pos_prose_idx).append(a)\n print(f\" dev positives: prose={len(pos_prose_idx)} tech={len(pos_tech_idx)}\", flush=True)\n\n print(\"[2] streaming pool: featurize + quality stats\", flush=True)\n ids, idx_arrays, stats = [], [], []\n n = 0\n for line in open(POOL):\n r = json.loads(line)\n a, st = doc_features(r[\"text\"])\n ids.append(r[\"id\"]); idx_arrays.append(a); stats.append(st)\n n += 1\n print(f\" pooled {n} docs in {time.time()-t0:.0f}s\", flush=True)\n\n ids = np.array(ids, dtype=np.int64)\n X = build_csr(idx_arrays, D)\n bg_df = np.asarray(X.sum(axis=0)).ravel() # document frequency over pool\n nbg = X.shape[0]\n\n # per-head log-odds weights\n Xp = build_csr(pos_prose_idx, D); Xt = build_csr(pos_tech_idx, D)\n r_prose = nb_logodds(np.asarray(Xp.sum(0)).ravel(), bg_df, Xp.shape[0], nbg)\n r_tech = nb_logodds(np.asarray(Xt.sum(0)).ravel(), bg_df, Xt.shape[0], nbg)\n s_prose = X.dot(r_prose)\n s_tech = X.dot(r_tech)\n\n # gather stats into arrays\n def col(k): return np.array([s[k] for s in stats])\n nw = col(\"nw\"); tok_est = col(\"tok_est\"); alpha = col(\"alpha\"); digit = col(\"digit\")\n ascii_ = col(\"ascii\"); sym = col(\"sym\"); stopfrac = col(\"stopfrac\"); uniq = col(\"uniq\")\n mwl = col(\"mwl\"); has_code = col(\"has_code\").astype(bool)\n\n # register-aware junk filters\n prose_ok = ((nw >= 50) & (alpha >= 0.60) & (stopfrac >= 0.12) & (uniq >= 0.30) &\n (mwl <= 11.0) & (ascii_ >= 0.90) & (sym <= 0.15))\n tech_ok = ((nw >= 30) & (alpha >= 0.35) & (uniq >= 0.25) & (ascii_ >= 0.85) &\n (mwl <= 14.0))\n\n # assign each doc to the head it scores higher on (standardized)\n zp = (s_prose - s_prose.mean()) / (s_prose.std() + 1e-9)\n zt = (s_tech - s_tech.mean()) / (s_tech.std() + 1e-9)\n is_tech = (zt > zp)\n\n # candidate lists\n tech_mask = is_tech & tech_ok\n prose_mask = (~is_tech) & prose_ok\n tech_order = np.argsort(-s_tech[tech_mask])\n prose_order = np.argsort(-s_prose[prose_mask])\n tech_pos = np.where(tech_mask)[0][tech_order]\n prose_pos = np.where(prose_mask)[0][prose_order]\n print(f\"[3] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"\n f\"(dropped {n-len(prose_pos)-len(tech_pos)} by filter/assignment)\", flush=True)\n\n # exact-duplicate guard via cheap fingerprint of stats (nw,alpha rounded) + first score;\n # true dedup done on text hash below during emission.\n # token-balanced interleave to 75/25, then overflow.\n tech_budget = BUDGET * TECH_SHARE\n prose_budget = BUDGET * (1 - TECH_SHARE)\n sel, seen = [], set()\n pt = pi = 0\n tacc = pacc = 0\n # phase 1: fill budget balanced\n while (tacc < tech_budget or pacc < prose_budget):\n take_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))\n if take_tech and pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen:\n seen.add(ids[j]); sel.append(j); tacc += tok_est[j]\n elif (not take_tech) and pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen:\n seen.add(ids[j]); sel.append(j); pacc += tok_est[j]\n else:\n # one side exhausted; drain the other\n if pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen: seen.add(ids[j]); sel.append(j); tacc += tok_est[j]\n elif pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen: seen.add(ids[j]); sel.append(j); pacc += tok_est[j]\n else:\n break\n print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M tokens\", flush=True)\n\n # phase 2: overflow — append remaining ranked docs (prose then tech interleaved) for margin\n rest = []\n while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) + len(rest) < N_OVERFLOW_IDS:\n for _ in range(3):\n if pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen: seen.add(ids[j]); rest.append(j)\n if pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen: seen.add(ids[j]); rest.append(j)\n sel.extend(rest)\n\n out_ids = [int(ids[j]) for j in sel]\n json.dump(out_ids, open(OUT, \"w\"))\n est_tokens = sum(int(tok_est[j]) for j in sel)\n print(f\"[4] wrote {len(out_ids)} ids -> {OUT} (~{est_tokens/1e6:.0f}M est tokens) \"\n f\"in {time.time()-t0:.0f}s\", flush=True)\n\n # diagnostics: peek at top docs of each head\n if \"--peek\" in sys.argv:\n texts = {}\n for line in open(POOL):\n r = json.loads(line); texts[r[\"id\"]] = r[\"text\"]\n print(\"\\n===== TOP PROSE =====\")\n for j in prose_pos[:4]:\n print(f\"[score {s_prose[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")\n print(\"\\n===== TOP TECH =====\")\n for j in tech_pos[:4]:\n print(f\"[score {s_tech[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")\n print(\"\\n===== BOTTOM (lowest prose score, any) =====\")\n worst = np.argsort(s_prose)[:3]\n for j in worst:\n print(f\"[score {s_prose[j]:.1f}] {texts[int(ids[j])][:200]!r}\")\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
File created successfully at: /workspace/curate.py (file state is current in your context — no need to Read it back)
[metadata] {"type": "create", "filePath": "/workspace/curate.py", "content": "#!/usr/bin/env python3\n\"\"\"\nCurate a raw web pool for pretraining a small LM whose eval target is a BROAD,\nmulti-domain high-quality English mixture (equal parts encyclopedic/Wikipedia,\ngeneral high-quality web prose, news, technical Q&A).\n\nSTATED CRITERION\n----------------\nKeep the pool documents that are (a) clean, fluent, information-dense English and\n(b) distributionally close to the disclosed target. We measure target-closeness\nwith an NBSVM-style log-odds score: for hashed word unigram+bigram features, weigh\neach feature by log P(feature | target register) - log P(feature | pool background).\nA document's score is the sum of its feature weights == its log-likelihood ratio of\nbeing drawn from the target register vs. the generic web pool (a DSIR-style signal).\n\nThe target is a fixed mixture whose one unambiguous slice is technical Q&A (~25% of\ntarget tokens, identifiable by HTML/code markup). Programming/Q&A prose is\nstylistically distinct from encyclopedic/news/web prose, so we score with TWO heads:\n - PROSE head: positives = decoded dev segments WITHOUT code/HTML markup\n - TECH head: positives = decoded dev segments WITH code/HTML markup\nEach pool doc is assigned to the head it looks most like, passes a register-aware\njunk filter, and is ranked by that head's score. We then fill the 12M-token budget\ntoken-balanced 75% prose / 25% tech to match the target mixture, interleaving so any\nearly budget cutoff stays balanced. Remaining ranked ids are appended as overflow.\n\nThe positives come only from the disclosed dev target; the selection is over disjoint\npool ids. Output is an ordered id list, produced entirely by this script (no hand\npicking).\n\"\"\"\nimport json, re, zlib, sys, time\nimport numpy as np\nfrom scipy import sparse\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nBUDGET = 12_000_000\nTECH_SHARE = 0.25 # target token share of technical Q&A register\nD = 1 << 20 # hashed feature dimension\nWORD_CAP = 1200 # cap words/doc for featurization + stats (bounds cost)\nALPHA = 1.0 # NB smoothing\nN_OVERFLOW_IDS = 45_000 # total ids to emit (well over budget for safety)\n\nwre = re.compile(r\"[a-z0-9']+\")\nSTOP = set(\"the a an and or but if then else of to in on at by for with from as is are was \"\n \"were be been being it its this that these those he she they we you i his her their \"\n \"our your not no do does did has have had will would can could should may might must \"\n \"there here what which who whom whose when where why how than so such about into over \"\n \"after before between out up down off again more most other some any each\".split())\n\ndef doc_features(text):\n \"\"\"Return (feature_index_array_uint32, stats_dict) for one document.\"\"\"\n low = text.lower()\n w = wre.findall(low)\n if len(w) > WORD_CAP:\n w = w[:WORD_CAP]\n nw = len(w)\n idx = set()\n for tk in w:\n idx.add(zlib.crc32(tk.encode()) & (D - 1))\n for i in range(nw - 1):\n idx.add(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (D - 1))\n # char-level quality stats over a bounded prefix\n seg = text[:4000]\n L = max(1, len(seg))\n alpha = sum(c.isalpha() for c in seg)\n digit = sum(c.isdigit() for c in seg)\n ascii_ = sum(ord(c) < 128 for c in seg)\n space = sum(c.isspace() for c in seg)\n sym = L - alpha - digit - space\n stops = sum(1 for x in w if x in STOP)\n uniq = len(set(w))\n mwl = (sum(len(x) for x in w) / nw) if nw else 0.0\n stats = dict(\n nw=nw,\n tok_est=max(1, len(text) // 4),\n alpha=alpha / L,\n digit=digit / L,\n ascii=ascii_ / L,\n sym=sym / L,\n stopfrac=(stops / nw) if nw else 0.0,\n uniq=(uniq / nw) if nw else 0.0,\n mwl=mwl,\n has_code=((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)\n or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)),\n )\n return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stats\n\n\ndef build_csr(list_of_idx_arrays, D):\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n for i, a in enumerate(list_of_idx_arrays):\n indptr[i + 1] = indptr[i] + len(a)\n indices = np.concatenate(list_of_idx_arrays) if list_of_idx_arrays else np.zeros(0, np.uint32)\n data = np.ones(len(indices), dtype=np.float32)\n return sparse.csr_matrix((data, indices, indptr), shape=(len(list_of_idx_arrays), D))\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):\n \"\"\"NBSVM/DSIR-style per-feature log-odds: log p(f|target) - log p(f|background).\"\"\"\n p = (pos_df + ALPHA) / (npos + 2 * ALPHA)\n q = (bg_df + ALPHA) / (nbg + 2 * ALPHA)\n return np.log(p) - np.log(q)\n\n\ndef main():\n t0 = time.time()\n print(\"[1] decoding dev target -> positive register sets\", flush=True)\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n dev = np.load(DEV)\n cut = np.where(dev == eos)[0]\n starts = np.concatenate([[0], cut + 1]); ends = np.concatenate([cut, [len(dev)]])\n pos_prose_idx, pos_tech_idx = [], []\n for s, e in zip(starts, ends):\n if e - s < 3:\n continue\n txt = tok.decode(dev[s:e].tolist())\n a, st = doc_features(txt)\n (pos_tech_idx if st[\"has_code\"] else pos_prose_idx).append(a)\n print(f\" dev positives: prose={len(pos_prose_idx)} tech={len(pos_tech_idx)}\", flush=True)\n\n print(\"[2] streaming pool: featurize + quality stats\", flush=True)\n ids, idx_arrays, stats = [], [], []\n n = 0\n for line in open(POOL):\n r = json.loads(line)\n a, st = doc_features(r[\"text\"])\n ids.append(r[\"id\"]); idx_arrays.append(a); stats.append(st)\n n += 1\n print(f\" pooled {n} docs in {time.time()-t0:.0f}s\", flush=True)\n\n ids = np.array(ids, dtype=np.int64)\n X = build_csr(idx_arrays, D)\n bg_df = np.asarray(X.sum(axis=0)).ravel() # document frequency over pool\n nbg = X.shape[0]\n\n # per-head log-odds weights\n Xp = build_csr(pos_prose_idx, D); Xt = build_csr(pos_tech_idx, D)\n r_prose = nb_logodds(np.asarray(Xp.sum(0)).ravel(), bg_df, Xp.shape[0], nbg)\n r_tech = nb_logodds(np.asarray(Xt.sum(0)).ravel(), bg_df, Xt.shape[0], nbg)\n s_prose = X.dot(r_prose)\n s_tech = X.dot(r_tech)\n\n # gather stats into arrays\n def col(k): return np.array([s[k] for s in stats])\n nw = col(\"nw\"); tok_est = col(\"tok_est\"); alpha = col(\"alpha\"); digit = col(\"digit\")\n ascii_ = col(\"ascii\"); sym = col(\"sym\"); stopfrac = col(\"stopfrac\"); uniq = col(\"uniq\")\n mwl = col(\"mwl\"); has_code = col(\"has_code\").astype(bool)\n\n # register-aware junk filters\n prose_ok = ((nw >= 50) & (alpha >= 0.60) & (stopfrac >= 0.12) & (uniq >= 0.30) &\n (mwl <= 11.0) & (ascii_ >= 0.90) & (sym <= 0.15))\n tech_ok = ((nw >= 30) & (alpha >= 0.35) & (uniq >= 0.25) & (ascii_ >= 0.85) &\n (mwl <= 14.0))\n\n # assign each doc to the head it scores higher on (standardized)\n zp = (s_prose - s_prose.mean()) / (s_prose.std() + 1e-9)\n zt = (s_tech - s_tech.mean()) / (s_tech.std() + 1e-9)\n is_tech = (zt > zp)\n\n # candidate lists\n tech_mask = is_tech & tech_ok\n prose_mask = (~is_tech) & prose_ok\n tech_order = np.argsort(-s_tech[tech_mask])\n prose_order = np.argsort(-s_prose[prose_mask])\n tech_pos = np.where(tech_mask)[0][tech_order]\n prose_pos = np.where(prose_mask)[0][prose_order]\n print(f\"[3] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"\n f\"(dropped {n-len(prose_pos)-len(tech_pos)} by filter/assignment)\", flush=True)\n\n # exact-duplicate guard via cheap fingerprint of stats (nw,alpha rounded) + first score;\n # true dedup done on text hash below during emission.\n # token-balanced interleave to 75/25, then overflow.\n tech_budget = BUDGET * TECH_SHARE\n prose_budget = BUDGET * (1 - TECH_SHARE)\n sel, seen = [], set()\n pt = pi = 0\n tacc = pacc = 0\n # phase 1: fill budget balanced\n while (tacc < tech_budget or pacc < prose_budget):\n take_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))\n if take_tech and pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen:\n seen.add(ids[j]); sel.append(j); tacc += tok_est[j]\n elif (not take_tech) and pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen:\n seen.add(ids[j]); sel.append(j); pacc += tok_est[j]\n else:\n # one side exhausted; drain the other\n if pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen: seen.add(ids[j]); sel.append(j); tacc += tok_est[j]\n elif pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen: seen.add(ids[j]); sel.append(j); pacc += tok_est[j]\n else:\n break\n print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M tokens\", flush=True)\n\n # phase 2: overflow — append remaining ranked docs (prose then tech interleaved) for margin\n rest = []\n while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) + len(rest) < N_OVERFLOW_IDS:\n for _ in range(3):\n if pi < len(prose_pos):\n j = prose_pos[pi]; pi += 1\n if ids[j] not in seen: seen.add(ids[j]); rest.append(j)\n if pt < len(tech_pos):\n j = tech_pos[pt]; pt += 1\n if ids[j] not in seen: seen.add(ids[j]); rest.append(j)\n sel.extend(rest)\n\n out_ids = [int(ids[j]) for j in sel]\n json.dump(out_ids, open(OUT, \"w\"))\n est_tokens = sum(int(tok_est[j]) for j in sel)\n print(f\"[4] wrote {len(out_ids)} ids -> {OUT} (~{est_tokens/1e6:.0f}M est tokens) \"\n f\"in {time.time()-t0:.0f}s\", flush=True)\n\n # diagnostics: peek at top docs of each head\n if \"--peek\" in sys.argv:\n texts = {}\n for line in open(POOL):\n r = json.loads(line); texts[r[\"id\"]] = r[\"text\"]\n print(\"\\n===== TOP PROSE =====\")\n for j in prose_pos[:4]:\n print(f\"[score {s_prose[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")\n print(\"\\n===== TOP TECH =====\")\n for j in tech_pos[:4]:\n print(f\"[score {s_tech[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")\n print(\"\\n===== BOTTOM (lowest prose score, any) =====\")\n worst = np.argsort(s_prose)[:3]\n for j in worst:\n print(f\"[score {s_prose[j]:.1f}] {texts[int(ids[j])][:200]!r}\")\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [], "originalFile": null, "userModified": false}Baseline finished. Let me check its result and run curate.py with --peek to validate the scorer.
args
{
"command": "cat /tmp/baseline_random.log",
"description": "Read random baseline result"
}Bash result
Token indices sequence length is longer than the specified maximum sequence length for this model (2126 > 1024). Running this sequence through the model will result in indexing errors
packed 12000000 tokens from 30000 ids -> /tmp/dev_train.npy
HELDOUT_PPL 476.714 loss 6.1669 train_s 54 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
DONE_EXIT 0
[stdout]
Token indices sequence length is longer than the specified maximum sequence length for this model (2126 > 1024). Running this sequence through the model will result in indexing errors
packed 12000000 tokens from 30000 ids -> /tmp/dev_train.npy
HELDOUT_PPL 476.714 loss 6.1669 train_s 54 /tmp/dev_train.npy
(dev PPL on multi_dev; official uses a hidden disjoint sample of the same HQ domain)
DONE_EXIT 0
[metadata] {"noOutputExpected": false}args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 curate.py --peek 2>&1 | tail -50",
"description": "Run curate.py with peek diagnostics"
}Bash result
Traceback (most recent call last):
File "/workspace/curate.py", line 32, in <module>
from scipy import sparse
ModuleNotFoundError: No module named 'scipy'
[stdout]
Traceback (most recent call last):
File "/workspace/curate.py", line 32, in <module>
from scipy import sparse
ModuleNotFoundError: No module named 'scipy'
[metadata] {"noOutputExpected": false}Random baseline PPL = 476.7 — that's my number to beat. But scipy isn't available. No problem — I'll implement the sparse ops with pure numpy (bincount for document-frequency, add.reduceat for per-doc scoring), which is actually faster. Let me rewrite those parts.
args
{
"replace_all": false,
"file_path": "/workspace/curate.py",
"old_string": "import json, re, zlib, sys, time\nimport numpy as np\nfrom scipy import sparse\nfrom transformers import AutoTokenizer",
"new_string": "import json, re, zlib, sys, time\nimport numpy as np\nfrom transformers import AutoTokenizer"
}Edit result
The file /workspace/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/curate.py", "oldString": "import json, re, zlib, sys, time\nimport numpy as np\nfrom scipy import sparse\nfrom transformers import AutoTokenizer", "newString": "import json, re, zlib, sys, time\nimport numpy as np\nfrom transformers import AutoTokenizer", "originalFile": null, "structuredPatch": [{"oldStart": 29, "oldLines": 7, "newStart": 29, "newLines": 6, "lines": [" \"\"\"", " import json, re, zlib, sys, time", " import numpy as np", "-from scipy import sparse", " from transformers import AutoTokenizer", " ", " POOL = \"/workspace/data/pool.jsonl\""]}], "userModified": false, "replaceAll": false}encrypted chain-of-thought — recorded, not exposed
args
{
"replace_all": false,
"file_path": "/workspace/curate.py",
"old_string": "def build_csr(list_of_idx_arrays, D):\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n for i, a in enumerate(list_of_idx_arrays):\n indptr[i + 1] = indptr[i] + len(a)\n indices = np.concatenate(list_of_idx_arrays) if list_of_idx_arrays else np.zeros(0, np.uint32)\n data = np.ones(len(indices), dtype=np.float32)\n return sparse.csr_matrix((data, indices, indptr), shape=(len(list_of_idx_arrays), D))\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):",
"new_string": "def concat_idx(list_of_idx_arrays):\n \"\"\"Flatten per-doc unique feature indices into (allidx, indptr) CSR-row layout.\"\"\"\n lens = np.fromiter((len(a) for a in list_of_idx_arrays), dtype=np.int64,\n count=len(list_of_idx_arrays))\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n np.cumsum(lens, out=indptr[1:])\n allidx = (np.concatenate(list_of_idx_arrays).astype(np.int64)\n if list_of_idx_arrays else np.zeros(0, np.int64))\n return allidx, indptr\n\n\ndef df_from(allidx):\n \"\"\"Document frequency per feature (idx arrays are per-doc unique -> bincount == df).\"\"\"\n return np.bincount(allidx, minlength=D).astype(np.float64)\n\n\ndef score_docs(r, allidx, indptr):\n \"\"\"Per-doc sum of feature weights r over the CSR rows == log-likelihood-ratio score.\"\"\"\n if len(allidx) == 0:\n return np.zeros(len(indptr) - 1)\n return np.add.reduceat(r[allidx], indptr[:-1])\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):"
}Edit result
The file /workspace/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/curate.py", "oldString": "def build_csr(list_of_idx_arrays, D):\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n for i, a in enumerate(list_of_idx_arrays):\n indptr[i + 1] = indptr[i] + len(a)\n indices = np.concatenate(list_of_idx_arrays) if list_of_idx_arrays else np.zeros(0, np.uint32)\n data = np.ones(len(indices), dtype=np.float32)\n return sparse.csr_matrix((data, indices, indptr), shape=(len(list_of_idx_arrays), D))\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):", "newString": "def concat_idx(list_of_idx_arrays):\n \"\"\"Flatten per-doc unique feature indices into (allidx, indptr) CSR-row layout.\"\"\"\n lens = np.fromiter((len(a) for a in list_of_idx_arrays), dtype=np.int64,\n count=len(list_of_idx_arrays))\n indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)\n np.cumsum(lens, out=indptr[1:])\n allidx = (np.concatenate(list_of_idx_arrays).astype(np.int64)\n if list_of_idx_arrays else np.zeros(0, np.int64))\n return allidx, indptr\n\n\ndef df_from(allidx):\n \"\"\"Document frequency per feature (idx arrays are per-doc unique -> bincount == df).\"\"\"\n return np.bincount(allidx, minlength=D).astype(np.float64)\n\n\ndef score_docs(r, allidx, indptr):\n \"\"\"Per-doc sum of feature weights r over the CSR rows == log-likelihood-ratio score.\"\"\"\n if len(allidx) == 0:\n return np.zeros(len(indptr) - 1)\n return np.add.reduceat(r[allidx], indptr[:-1])\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):", "originalFile": null, "structuredPatch": [{"oldStart": 87, "oldLines": 15, "newStart": 87, "newLines": 29, "lines": [" return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stats", " ", " ", "-def build_csr(list_of_idx_arrays, D):", "+def concat_idx(list_of_idx_arrays):", "+ \"\"\"Flatten per-doc unique feature indices into (allidx, indptr) CSR-row layout.\"\"\"", "+ lens = np.fromiter((len(a) for a in list_of_idx_arrays), dtype=np.int64,", "+ count=len(list_of_idx_arrays))", " indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)", "- for i, a in enumerate(list_of_idx_arrays):", "- indptr[i + 1] = indptr[i] + len(a)", "- indices = np.concatenate(list_of_idx_arrays) if list_of_idx_arrays else np.zeros(0, np.uint32)", "- data = np.ones(len(indices), dtype=np.float32)", "- return sparse.csr_matrix((data, indices, indptr), shape=(len(list_of_idx_arrays), D))", "+ np.cumsum(lens, out=indptr[1:])", "+ allidx = (np.concatenate(list_of_idx_arrays).astype(np.int64)", "+ if list_of_idx_arrays else np.zeros(0, np.int64))", "+ return allidx, indptr", " ", " ", "+def df_from(allidx):", "+ \"\"\"Document frequency per feature (idx arrays are per-doc unique -> bincount == df).\"\"\"", "+ return np.bincount(allidx, minlength=D).astype(np.float64)", "+", "+", "+def score_docs(r, allidx, indptr):", "+ \"\"\"Per-doc sum of feature weights r over the CSR rows == log-likelihood-ratio score.\"\"\"", "+ if len(allidx) == 0:", "+ return np.zeros(len(indptr) - 1)", "+ return np.add.reduceat(r[allidx], indptr[:-1])", "+", "+", " def nb_logodds(pos_df, bg_df, npos, nbg):", " \"\"\"NBSVM/DSIR-style per-feature log-odds: log p(f|target) - log p(f|background).\"\"\"", " p = (pos_df + ALPHA) / (npos + 2 * ALPHA)"]}], "userModified": false, "replaceAll": false}encrypted chain-of-thought — recorded, not exposed
args
{
"replace_all": false,
"file_path": "/workspace/curate.py",
"old_string": " ids = np.array(ids, dtype=np.int64)\n X = build_csr(idx_arrays, D)\n bg_df = np.asarray(X.sum(axis=0)).ravel() # document frequency over pool\n nbg = X.shape[0]\n\n # per-head log-odds weights\n Xp = build_csr(pos_prose_idx, D); Xt = build_csr(pos_tech_idx, D)\n r_prose = nb_logodds(np.asarray(Xp.sum(0)).ravel(), bg_df, Xp.shape[0], nbg)\n r_tech = nb_logodds(np.asarray(Xt.sum(0)).ravel(), bg_df, Xt.shape[0], nbg)\n s_prose = X.dot(r_prose)\n s_tech = X.dot(r_tech)",
"new_string": " ids = np.array(ids, dtype=np.int64)\n allidx, indptr = concat_idx(idx_arrays)\n bg_df = df_from(allidx) # document frequency over pool\n nbg = len(idx_arrays)\n\n # per-head log-odds weights\n pa, _ = concat_idx(pos_prose_idx); ta, _ = concat_idx(pos_tech_idx)\n r_prose = nb_logodds(df_from(pa), bg_df, len(pos_prose_idx), nbg)\n r_tech = nb_logodds(df_from(ta), bg_df, len(pos_tech_idx), nbg)\n s_prose = score_docs(r_prose, allidx, indptr)\n s_tech = score_docs(r_tech, allidx, indptr)"
}Edit result
The file /workspace/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"filePath": "/workspace/curate.py", "oldString": " ids = np.array(ids, dtype=np.int64)\n X = build_csr(idx_arrays, D)\n bg_df = np.asarray(X.sum(axis=0)).ravel() # document frequency over pool\n nbg = X.shape[0]\n\n # per-head log-odds weights\n Xp = build_csr(pos_prose_idx, D); Xt = build_csr(pos_tech_idx, D)\n r_prose = nb_logodds(np.asarray(Xp.sum(0)).ravel(), bg_df, Xp.shape[0], nbg)\n r_tech = nb_logodds(np.asarray(Xt.sum(0)).ravel(), bg_df, Xt.shape[0], nbg)\n s_prose = X.dot(r_prose)\n s_tech = X.dot(r_tech)", "newString": " ids = np.array(ids, dtype=np.int64)\n allidx, indptr = concat_idx(idx_arrays)\n bg_df = df_from(allidx) # document frequency over pool\n nbg = len(idx_arrays)\n\n # per-head log-odds weights\n pa, _ = concat_idx(pos_prose_idx); ta, _ = concat_idx(pos_tech_idx)\n r_prose = nb_logodds(df_from(pa), bg_df, len(pos_prose_idx), nbg)\n r_tech = nb_logodds(df_from(ta), bg_df, len(pos_tech_idx), nbg)\n s_prose = score_docs(r_prose, allidx, indptr)\n s_tech = score_docs(r_tech, allidx, indptr)", "originalFile": null, "structuredPatch": [{"oldStart": 145, "oldLines": 16, "newStart": 145, "newLines": 16, "lines": [" print(f\" pooled {n} docs in {time.time()-t0:.0f}s\", flush=True)", " ", " ids = np.array(ids, dtype=np.int64)", "- X = build_csr(idx_arrays, D)", "- bg_df = np.asarray(X.sum(axis=0)).ravel() # document frequency over pool", "- nbg = X.shape[0]", "+ allidx, indptr = concat_idx(idx_arrays)", "+ bg_df = df_from(allidx) # document frequency over pool", "+ nbg = len(idx_arrays)", " ", " # per-head log-odds weights", "- Xp = build_csr(pos_prose_idx, D); Xt = build_csr(pos_tech_idx, D)", "- r_prose = nb_logodds(np.asarray(Xp.sum(0)).ravel(), bg_df, Xp.shape[0], nbg)", "- r_tech = nb_logodds(np.asarray(Xt.sum(0)).ravel(), bg_df, Xt.shape[0], nbg)", "- s_prose = X.dot(r_prose)", "- s_tech = X.dot(r_tech)", "+ pa, _ = concat_idx(pos_prose_idx); ta, _ = concat_idx(pos_tech_idx)", "+ r_prose = nb_logodds(df_from(pa), bg_df, len(pos_prose_idx), nbg)", "+ r_tech = nb_logodds(df_from(ta), bg_df, len(pos_tech_idx), nbg)", "+ s_prose = score_docs(r_prose, allidx, indptr)", "+ s_tech = score_docs(r_tech, allidx, indptr)", " ", " # gather stats into arrays", " def col(k): return np.array([s[k] for s in stats])"]}], "userModified": false, "replaceAll": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 curate.py --peek 2>&1 | tail -45",
"description": "Run curate.py with numpy backend + peek"
}Bash result
Command did not complete within its 120s timeout and was moved to the background (ID: by0deunmy). Output is being written to: /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "by0deunmy", "timedOutAfterMs": 120000, "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "sleep 20; cat /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output 2>/dev/null | tail -45",
"description": "Check curate.py progress"
}Bash result
(Bash completed with no output)
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cat /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output 2>/dev/null | tail -50; echo \"---remaining secs---\"; cat /workspace/.timer/remaining_secs",
"description": "Read curate output and timer"
}Bash result
---remaining secs---
4569
[stdout]
---remaining secs---
4569
[metadata] {"noOutputExpected": false}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "ls -la /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output; echo \"---\"; wc -c /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output; echo \"--- procs ---\"; ps aux | grep -E \"curate|python3\" | grep -v grep | head",
"description": "Check output file size and running processes"
}Bash result
-rw-r--r-- 1 root root 0 Aug 4 03:34 /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output
---
0 /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output
--- procs ---
root 154 0.0 0.0 12432 5680 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid exit "$rc"
root 156 0.0 0.0 12432 6060 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.2 0.0 6226752 544740 ? Sl 03:22 0:35 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
root 1240 0.0 0.0 12564 4612 ? Ss 03:34 0:00 /bin/bash -c source /logs/agent/sessions/shell-snapshots/snapshot-bash-1785813781393-2ipmy9.sh 2>/dev/null || true && shopt -u extglob 2>/dev/null || true && { \builtin unalias -- 'unsetenv'; \builtin unset -f -- 'unsetenv'; } >/dev/null 2>&1 || true && eval 'cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 curate.py --peek 2>&1 | tail -45' < /dev/null && pwd -P >| /tmp/claude-f486-cwd
root 1242 95.4 0.4 5811040 5193856 ? Rl 03:34 2:27 python3 curate.py --peek
[stdout]
-rw-r--r-- 1 root root 0 Aug 4 03:34 /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output
---
0 /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output
--- procs ---
root 154 0.0 0.0 12432 5680 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid exit "$rc"
root 156 0.0 0.0 12432 6060 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.2 0.0 6226752 544740 ? Sl 03:22 0:35 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
root 1240 0.0 0.0 12564 4612 ? Ss 03:34 0:00 /bin/bash -c source /logs/agent/sessions/shell-snapshots/snapshot-bash-1785813781393-2ipmy9.sh 2>/dev/null || true && shopt -u extglob 2>/dev/null || true && { \builtin unalias -- 'unsetenv'; \builtin unset -f -- 'unsetenv'; } >/dev/null 2>&1 || true && eval 'cd /workspace; export TOKENIZERS_PARALLELISM=false; python3 curate.py --peek 2>&1 | tail -45' < /dev/null && pwd -P >| /tmp/claude-f486-cwd
root 1242 95.4 0.4 5811040 5193856 ? Rl 03:34 2:27 python3 curate.py --peek
[metadata] {"noOutputExpected": false}The job is running (PID 1242, 95% CPU) — my | tail -45 buffered all output until EOF, so nothing showed. Let me wait for it to finish.
args
{
"command": "for i in $(seq 1 20); do if ! kill -0 1242 2>/dev/null; then echo \"process finished\"; break; fi; sleep 5; done; echo \"=== output ===\"; cat /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/by0deunmy.output 2>/dev/null | tail -45",
"description": "Wait for curate.py to finish and read output"
}Bash result
process finished
=== output ===
[1] decoding dev target -> positive register sets
dev positives: prose=2345 tech=360
[2] streaming pool: featurize + quality stats
pooled 182016 docs in 159s
[3] candidates: prose=95295 tech=72836 (dropped 13885 by filter/assignment)
budget phase: 30079 docs, tech~3.0M prose~9.0M tokens
[4] wrote 45003 ids -> /workspace/submission/selection.json (~32M est tokens) in 163s
===== TOP PROSE =====
[score 515.6 nw 975] '|PREDECESSORS AND SHORT HISTORY:\nin the 16th century by Rao Gopalji fl.1550, 4th son of Raja Prithviraj\nSinghji I of Jaipur. Rulers were…\n- Rao GOPAL SINGH,\nThakur Saheb of Chomu fl.1550, married\nand had issue.\n- Thakur NATHU SINGH,\nThakur '
[score 456.8 nw 1200] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army prisoners of war|\n|D'
[score 399.9 nw 1200] 'an, also written as Khurasan (Middle Persian: Khwarāsān, Persian: خراسان بزرگ or خراسان کهن listen (help·info)), is a historical region lying in the northeast of Persia. "In pre-Islamic and early Islamic times, the term "Khurassan" frequent'
[score 386.4 nw 623] 'a<|endoftext|>Czech Tobe alit it staleness checker unguardedly. lignifies sigillate that unsnarl hellishly? gassy Thacher doodle, her precesses mundanely. wainscoted and sex-limited Ricki bestialize her Cowes executed and typecast proficien'
===== TOP TECH =====
[score 4870.0 nw 1200] "? #debill\nWhat Digital Economy Bill? #debill\nLinks\n@doc1online has declared that they do not recognise the Digital Economy Bill\nI choose not to recognise the UK's Digital Economy Bill #whatdebill #debill http://whatdebill.org\nView tweet\nDec"
[score 4832.4 nw 1200] " Blogger.<|endoftext|>What Digital Economy Bill? #debill\nWhat Digital Economy Bill? #debill\nLinks\n@doc1online has declared that they do not recognise the Digital Economy Bill\nI choose not to recognise the UK's Digital Economy Bill #whatdebi"
[score 4343.1 nw 1200] ' LastFm - User - GetTopTags API :: Temboo\nSorry, Temboo requires JavaScript! Please enable it in your browser and try again.\nPRODUCTS\nKosmos IoT\nBuild industrial Internet of Things systems\nAPIs\nConnect to any API in any language on any devi'
[score 4343.1 nw 1200] ' LastFm - User - GetTopTags API :: Temboo\nSorry, Temboo requires JavaScript! Please enable it in your browser and try again.\nPRODUCTS\nKosmos IoT\nBuild industrial Internet of Things systems\nAPIs\nConnect to any API in any language on any devi'
===== BOTTOM (lowest prose score, any) =====
[score -1344.3] 'es.<|endoftext|>Unique Exquisite Embroidered Pillow Cover | Discovered\nAll categories\nAll categories\nJewelry\nAccessories\nBags\nHome Decor\nSign in\nJoin\n0\nAll categories\nAll categories\nJewelry\nAccessorie'
[score -1344.3] 'es.<|endoftext|>Unique Exquisite Embroidered Pillow Cover | Discovered\nAll categories\nAll categories\nJewelry\nAccessories\nBags\nHome Decor\nSign in\nJoin\n0\nAll categories\nAll categories\nJewelry\nAccessorie'
[score -1335.3] ', RI<|endoftext|>Beware: Betaling Google Chrome is Malware! - Ask Dave Taylor\nFacebook\nInstagram\nLinkedin\nPinterest\nTwitter\nYouTube\nHome\nVideos\nMost Popular\nTop Categories\nBooks\nAbout Dave\nAsk Me\n> Do'
[stdout]
process finished
=== output ===
[1] decoding dev target -> positive register sets
dev positives: prose=2345 tech=360
[2] streaming pool: featurize + quality stats
pooled 182016 docs in 159s
[3] candidates: prose=95295 tech=72836 (dropped 13885 by filter/assignment)
budget phase: 30079 docs, tech~3.0M prose~9.0M tokens
[4] wrote 45003 ids -> /workspace/submission/selection.json (~32M est tokens) in 163s
===== TOP PROSE =====
[score 515.6 nw 975] '|PREDECESSORS AND SHORT HISTORY:\nin the 16th century by Rao Gopalji fl.1550, 4th son of Raja Prithviraj\nSinghji I of Jaipur. Rulers were…\n- Rao GOPAL SINGH,\nThakur Saheb of Chomu fl.1550, married\nand had issue.\n- Thakur NATHU SINGH,\nThakur '
[score 456.8 nw 1200] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army prisoners of war|\n|D'
[score 399.9 nw 1200] 'an, also written as Khurasan (Middle Persian: Khwarāsān, Persian: خراسان بزرگ or خراسان کهن listen (help·info)), is a historical region lying in the northeast of Persia. "In pre-Islamic and early Islamic times, the term "Khurassan" frequent'
[score 386.4 nw 623] 'a<|endoftext|>Czech Tobe alit it staleness checker unguardedly. lignifies sigillate that unsnarl hellishly? gassy Thacher doodle, her precesses mundanely. wainscoted and sex-limited Ricki bestialize her Cowes executed and typecast proficien'
===== TOP TECH =====
[score 4870.0 nw 1200] "? #debill\nWhat Digital Economy Bill? #debill\nLinks\n@doc1online has declared that they do not recognise the Digital Economy Bill\nI choose not to recognise the UK's Digital Economy Bill #whatdebill #debill http://whatdebill.org\nView tweet\nDec"
[score 4832.4 nw 1200] " Blogger.<|endoftext|>What Digital Economy Bill? #debill\nWhat Digital Economy Bill? #debill\nLinks\n@doc1online has declared that they do not recognise the Digital Economy Bill\nI choose not to recognise the UK's Digital Economy Bill #whatdebi"
[score 4343.1 nw 1200] ' LastFm - User - GetTopTags API :: Temboo\nSorry, Temboo requires JavaScript! Please enable it in your browser and try again.\nPRODUCTS\nKosmos IoT\nBuild industrial Internet of Things systems\nAPIs\nConnect to any API in any language on any devi'
[score 4343.1 nw 1200] ' LastFm - User - GetTopTags API :: Temboo\nSorry, Temboo requires JavaScript! Please enable it in your browser and try again.\nPRODUCTS\nKosmos IoT\nBuild industrial Internet of Things systems\nAPIs\nConnect to any API in any language on any devi'
===== BOTTOM (lowest prose score, any) =====
[score -1344.3] 'es.<|endoftext|>Unique Exquisite Embroidered Pillow Cover | Discovered\nAll categories\nAll categories\nJewelry\nAccessories\nBags\nHome Decor\nSign in\nJoin\n0\nAll categories\nAll categories\nJewelry\nAccessorie'
[score -1344.3] 'es.<|endoftext|>Unique Exquisite Embroidered Pillow Cover | Discovered\nAll categories\nAll categories\nJewelry\nAccessories\nBags\nHome Decor\nSign in\nJoin\n0\nAll categories\nAll categories\nJewelry\nAccessorie'
[score -1335.3] ', RI<|endoftext|>Beware: Betaling Google Chrome is Malware! - Ask Dave Taylor\nFacebook\nInstagram\nLinkedin\nPinterest\nTwitter\nYouTube\nHome\nVideos\nMost Popular\nTop Categories\nBooks\nAbout Dave\nAsk Me\n> Do'
[metadata] {"noOutputExpected": false}The pipeline works and beats the baseline conceptually, but the peek reveals three fixable problems:
- Rare-word gibberish (SEO word-salad like "lignifies sigillate unsnarl hellishly") scores high on the prose head — classic sum-of-log-odds failure.
- Tech-keyword boilerplate (
#debilltweets, "requires JavaScript" nav menus) dominates the tech head. - Near-duplicate docs waste budget.
The fixes: length-normalized mean of clipped log-odds (kills additive rare-word inflation), a stopword-fraction floor + short-line filter (kills word-salad and menu boilerplate), and text-fingerprint dedup. I'll also add a feature cache so I can iterate on thresholds in seconds instead of re-featurizing (159s). Let me rewrite curate.py.
args
{
"file_path": "/workspace/curate.py",
"content": "#!/usr/bin/env python3\n\"\"\"\nCurate a raw web pool for pretraining a small LM whose eval target is a BROAD,\nmulti-domain high-quality English mixture (equal parts encyclopedic/Wikipedia,\ngeneral high-quality web prose, news, technical Q&A).\n\nSTATED CRITERION\n----------------\nKeep pool documents that are (a) clean, fluent, information-dense English and\n(b) distributionally close to the disclosed target register mix.\n\nTarget-closeness is a DSIR / NBSVM-style log-likelihood ratio. For hashed word\nunigram+bigram features we weight each feature f by\n r[f] = log P(f | target register) - log P(f | pool background)\nestimated from feature document-frequencies (Laplace-smoothed). A document's score\nis the MEAN over its features of clip(r[f], -5, 5). Using the mean (not the sum)\nremoves length bias, and clipping stops a handful of ultra-rare \"fancy\" words from\nletting SEO word-salad win -- the two dominant failure modes of a raw sum.\n\nThe target is a fixed mixture whose one unambiguous slice is technical Q&A (~25% of\ntarget tokens, identifiable by HTML/code markup in the decoded dev set). Programming/\nQ&A text is stylistically distinct from encyclopedic/news/web prose, so we score with\nTWO heads sharing the same features:\n PROSE head : positives = decoded dev segments WITHOUT code/HTML markup\n TECH head : positives = decoded dev segments WITH code/HTML markup\nEach pool doc is assigned to whichever head it looks most like (standardized score),\nmust pass a register-aware junk filter (drops word-salad, menu/boilerplate, non-English,\nmarkup dumps, repetitive spam), is deduplicated by a content fingerprint, and is ranked\nby that head's score. We then fill the 12M-token budget token-balanced 75% prose / 25%\ntech to match the target mixture, interleaving so any early budget cutoff stays balanced.\nRemaining ranked ids are appended as overflow so the packer never runs short.\n\nPositives come only from the disclosed dev target; the selection is over disjoint pool\nids. The output is an ordered id list produced entirely by this script.\n\"\"\"\nimport json, re, zlib, sys, time, os\nimport numpy as np\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE_DIR = \"/tmp/curate_cache\"\nBUDGET = 12_000_000\nTECH_SHARE = 0.25 # target token share of technical Q&A register\nD = 1 << 20 # hashed feature dimension\nWORD_CAP = 1200 # cap words/doc for featurization + stats (bounds cost)\nALPHA = 1.0 # NB Laplace smoothing\nCLIP = 5.0 # per-feature log-odds clip\nN_OVERFLOW_IDS = 45_000 # total ids to emit (well over budget for safety)\n\nwre = re.compile(r\"[a-z0-9']+\")\nSTOP = set(\"the a an and or but if then else of to in on at by for with from as is are was \"\n \"were be been being it its this that these those he she they we you i his her their \"\n \"our your not no do does did has have had will would can could should may might must \"\n \"there here what which who whom whose when where why how than so such about into over \"\n \"after before between out up down off again more most other some any each also\".split())\nNSTAT = 11 # number of per-doc stat columns\n\n\ndef doc_features(text):\n \"\"\"Return (unique_feature_idx uint32[], stat_vector float32[NSTAT], fingerprint uint32).\"\"\"\n low = text.lower()\n w = wre.findall(low)\n if len(w) > WORD_CAP:\n w = w[:WORD_CAP]\n nw = len(w)\n idx = set()\n for tk in w:\n idx.add(zlib.crc32(tk.encode()) & (D - 1))\n for i in range(nw - 1):\n idx.add(zlib.crc32((w[i] + \"\\x1f\" + w[i + 1]).encode()) & (D - 1))\n\n seg = text[:4000]\n L = max(1, len(seg))\n alpha = sum(c.isalpha() for c in seg)\n digit = sum(c.isdigit() for c in seg)\n ascii_ = sum(ord(c) < 128 for c in seg)\n space = sum(c.isspace() for c in seg)\n sym = L - alpha - digit - space\n stops = sum(1 for x in w if x in STOP)\n uniq = len(set(w))\n mwl = (sum(len(x) for x in w) / nw) if nw else 0.0\n # line structure: fraction of non-empty lines that are \"short\" (< 5 words)\n lines = [ln for ln in text[:8000].split(\"\\n\") if ln.strip()]\n if lines:\n frac_short = sum(1 for ln in lines if len(ln.split()) < 5) / len(lines)\n else:\n frac_short = 1.0\n has_code = 1.0 if ((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)\n or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)) else 0.0\n stat = np.array([nw, max(1, len(text) // 4), alpha / L, digit / L, ascii_ / L,\n sym / L, (stops / nw) if nw else 0.0, (uniq / nw) if nw else 0.0,\n mwl, frac_short, has_code], dtype=np.float32)\n fp = zlib.crc32((\" \".join(w[:25] + w[-25:]) + \"|\" + str(nw)).encode()) & 0xFFFFFFFF\n return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stat, np.uint32(fp)\n\n\ndef concat_idx(arrs):\n lens = np.fromiter((len(a) for a in arrs), dtype=np.int64, count=len(arrs))\n indptr = np.zeros(len(arrs) + 1, dtype=np.int64)\n np.cumsum(lens, out=indptr[1:])\n allidx = np.concatenate(arrs).astype(np.int64) if arrs else np.zeros(0, np.int64)\n return allidx, indptr\n\n\ndef df_from(allidx):\n return np.bincount(allidx, minlength=D).astype(np.float64)\n\n\ndef mean_score(r, allidx, indptr):\n \"\"\"Per-doc MEAN of clipped feature weights r (== length-normalized log-lik ratio).\"\"\"\n if len(allidx) == 0:\n return np.zeros(len(indptr) - 1)\n vals = np.clip(r[allidx], -CLIP, CLIP)\n ssum = np.add.reduceat(vals, indptr[:-1])\n cnt = np.diff(indptr).astype(np.float64)\n cnt[cnt == 0] = 1.0\n return ssum / cnt\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):\n p = (pos_df + ALPHA) / (npos + 2 * ALPHA)\n q = (bg_df + ALPHA) / (nbg + 2 * ALPHA)\n return np.log(p) - np.log(q)\n\n\ndef build_cache():\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n dev = np.load(DEV)\n cut = np.where(dev == eos)[0]\n starts = np.concatenate([[0], cut + 1]); ends = np.concatenate([cut, [len(dev)]])\n pos_prose, pos_tech = [], []\n for s, e in zip(starts, ends):\n if e - s < 3:\n continue\n a, st, _ = doc_features(tok.decode(dev[s:e].tolist()))\n (pos_tech if st[10] > 0 else pos_prose).append(a)\n print(f\" dev positives: prose={len(pos_prose)} tech={len(pos_tech)}\", flush=True)\n pa, _ = concat_idx(pos_prose); ta, _ = concat_idx(pos_tech)\n prose_df = df_from(pa); tech_df = df_from(ta)\n\n ids, arrs, stats, fps, heads = [], [], [], [], []\n for line in open(POOL):\n r = json.loads(line)\n a, st, fp = doc_features(r[\"text\"])\n ids.append(r[\"id\"]); arrs.append(a); stats.append(st); fps.append(fp)\n heads.append(r[\"text\"][:220])\n ids = np.array(ids, dtype=np.int64)\n allidx, indptr = concat_idx(arrs)\n stats = np.stack(stats); fps = np.array(fps, dtype=np.uint32)\n os.makedirs(CACHE_DIR, exist_ok=True)\n np.save(f\"{CACHE_DIR}/ids.npy\", ids)\n np.save(f\"{CACHE_DIR}/allidx.npy\", allidx)\n np.save(f\"{CACHE_DIR}/indptr.npy\", indptr)\n np.save(f\"{CACHE_DIR}/stats.npy\", stats)\n np.save(f\"{CACHE_DIR}/fps.npy\", fps)\n np.save(f\"{CACHE_DIR}/prose_df.npy\", prose_df)\n np.save(f\"{CACHE_DIR}/tech_df.npy\", tech_df)\n np.save(f\"{CACHE_DIR}/meta.npy\", np.array([len(pos_prose), len(pos_tech)], dtype=np.int64))\n with open(f\"{CACHE_DIR}/heads.json\", \"w\") as f:\n json.dump(heads, f)\n return ids, allidx, indptr, stats, fps, prose_df, tech_df, len(pos_prose), len(pos_tech), heads\n\n\ndef load_cache():\n ids = np.load(f\"{CACHE_DIR}/ids.npy\")\n allidx = np.load(f\"{CACHE_DIR}/allidx.npy\")\n indptr = np.load(f\"{CACHE_DIR}/indptr.npy\")\n stats = np.load(f\"{CACHE_DIR}/stats.npy\")\n fps = np.load(f\"{CACHE_DIR}/fps.npy\")\n prose_df = np.load(f\"{CACHE_DIR}/prose_df.npy\")\n tech_df = np.load(f\"{CACHE_DIR}/tech_df.npy\")\n meta = np.load(f\"{CACHE_DIR}/meta.npy\")\n heads = json.load(open(f\"{CACHE_DIR}/heads.json\"))\n return ids, allidx, indptr, stats, fps, prose_df, tech_df, int(meta[0]), int(meta[1]), heads\n\n\ndef main():\n t0 = time.time()\n if os.path.exists(f\"{CACHE_DIR}/ids.npy\") and \"--rebuild\" not in sys.argv:\n print(\"[1] loading feature cache\", flush=True)\n ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = load_cache()\n else:\n print(\"[1] featurizing dev positives + pool (one-time)\", flush=True)\n ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = build_cache()\n n = len(ids)\n bg_df = df_from(allidx); nbg = n\n print(f\" N={n} feats in {time.time()-t0:.0f}s\", flush=True)\n\n r_prose = nb_logodds(prose_df, bg_df, npp, nbg)\n r_tech = nb_logodds(tech_df, bg_df, npt, nbg)\n s_prose = mean_score(r_prose, allidx, indptr)\n s_tech = mean_score(r_tech, allidx, indptr)\n\n nw, tok_est, alpha, digit, ascii_, sym, stopfrac, uniq, mwl, fshort, has_code = \\\n [stats[:, i] for i in range(NSTAT)]\n\n # register-aware junk filters\n prose_ok = ((nw >= 50) & (alpha >= 0.62) & (stopfrac >= 0.20) & (uniq >= 0.30) &\n (uniq <= 0.92) & (mwl <= 10.0) & (ascii_ >= 0.92) & (sym <= 0.12) &\n (fshort <= 0.55) & (digit <= 0.20))\n tech_ok = ((nw >= 40) & (alpha >= 0.40) & (stopfrac >= 0.10) & (uniq >= 0.28) &\n (uniq <= 0.92) & (ascii_ >= 0.90) & (mwl <= 12.0) & (fshort <= 0.75))\n\n zp = (s_prose - s_prose.mean()) / (s_prose.std() + 1e-9)\n zt = (s_tech - s_tech.mean()) / (s_tech.std() + 1e-9)\n is_tech = zt > zp\n\n tech_mask = is_tech & tech_ok\n prose_mask = (~is_tech) & prose_ok\n tech_pos = np.where(tech_mask)[0][np.argsort(-s_tech[tech_mask])]\n prose_pos = np.where(prose_mask)[0][np.argsort(-s_prose[prose_mask])]\n print(f\"[2] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"\n f\"(dropped {n-len(prose_pos)-len(tech_pos)})\", flush=True)\n\n # token-balanced interleave to 75/25 with content-fingerprint dedup, then overflow\n tech_budget = BUDGET * TECH_SHARE\n prose_budget = BUDGET * (1 - TECH_SHARE)\n seen_fp = set()\n sel = []\n pt = pi = 0\n tacc = pacc = 0\n\n def take(order, ptr):\n while ptr < len(order):\n j = order[ptr]; ptr += 1\n fp = int(fps[j])\n if fp in seen_fp:\n continue\n seen_fp.add(fp)\n return j, ptr\n return None, ptr\n\n while (tacc < tech_budget or pacc < prose_budget) and (pt < len(tech_pos) or pi < len(prose_pos)):\n want_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))\n if want_tech and pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j); tacc += tok_est[j]\n elif pi < len(prose_pos):\n j, pi = take(prose_pos, pi)\n if j is not None: sel.append(j); pacc += tok_est[j]\n elif pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j); tacc += tok_est[j]\n print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M est tokens\",\n flush=True)\n\n # overflow: append remaining ranked docs (3 prose : 1 tech), deduped\n while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) < N_OVERFLOW_IDS:\n for _ in range(3):\n if pi < len(prose_pos):\n j, pi = take(prose_pos, pi)\n if j is not None: sel.append(j)\n if pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j)\n\n out_ids = [int(ids[j]) for j in sel]\n assert len(out_ids) == len(set(out_ids)), \"duplicate ids\"\n json.dump(out_ids, open(OUT, \"w\"))\n est = sum(int(tok_est[j]) for j in sel)\n print(f\"[3] wrote {len(out_ids)} ids -> {OUT} (~{est/1e6:.0f}M est tokens) in {time.time()-t0:.0f}s\",\n flush=True)\n\n if \"--peek\" in sys.argv:\n print(\"\\n===== TOP PROSE =====\")\n for j in prose_pos[:6]:\n print(f\"[s{s_prose[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"\n f\"fs{fshort[j]:.2f}] {heads[j]!r}\")\n print(\"\\n===== TOP TECH =====\")\n for j in tech_pos[:6]:\n print(f\"[s{s_tech[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"\n f\"fs{fshort[j]:.2f}] {heads[j]!r}\")\n print(\"\\n===== MID PROSE (rank ~8000) =====\")\n for j in prose_pos[8000:8004]:\n print(f\"[s{s_prose[j]:.2f}] {heads[j]!r}\")\n\n\nif __name__ == \"__main__\":\n main()\n"
}Write result
The file /workspace/curate.py has been updated successfully. (file state is current in your context — no need to Read it back)
[metadata] {"type": "update", "filePath": "/workspace/curate.py", "content": "#!/usr/bin/env python3\n\"\"\"\nCurate a raw web pool for pretraining a small LM whose eval target is a BROAD,\nmulti-domain high-quality English mixture (equal parts encyclopedic/Wikipedia,\ngeneral high-quality web prose, news, technical Q&A).\n\nSTATED CRITERION\n----------------\nKeep pool documents that are (a) clean, fluent, information-dense English and\n(b) distributionally close to the disclosed target register mix.\n\nTarget-closeness is a DSIR / NBSVM-style log-likelihood ratio. For hashed word\nunigram+bigram features we weight each feature f by\n r[f] = log P(f | target register) - log P(f | pool background)\nestimated from feature document-frequencies (Laplace-smoothed). A document's score\nis the MEAN over its features of clip(r[f], -5, 5). Using the mean (not the sum)\nremoves length bias, and clipping stops a handful of ultra-rare \"fancy\" words from\nletting SEO word-salad win -- the two dominant failure modes of a raw sum.\n\nThe target is a fixed mixture whose one unambiguous slice is technical Q&A (~25% of\ntarget tokens, identifiable by HTML/code markup in the decoded dev set). Programming/\nQ&A text is stylistically distinct from encyclopedic/news/web prose, so we score with\nTWO heads sharing the same features:\n PROSE head : positives = decoded dev segments WITHOUT code/HTML markup\n TECH head : positives = decoded dev segments WITH code/HTML markup\nEach pool doc is assigned to whichever head it looks most like (standardized score),\nmust pass a register-aware junk filter (drops word-salad, menu/boilerplate, non-English,\nmarkup dumps, repetitive spam), is deduplicated by a content fingerprint, and is ranked\nby that head's score. We then fill the 12M-token budget token-balanced 75% prose / 25%\ntech to match the target mixture, interleaving so any early budget cutoff stays balanced.\nRemaining ranked ids are appended as overflow so the packer never runs short.\n\nPositives come only from the disclosed dev target; the selection is over disjoint pool\nids. The output is an ordered id list produced entirely by this script.\n\"\"\"\nimport json, re, zlib, sys, time, os\nimport numpy as np\nfrom transformers import AutoTokenizer\n\nPOOL = \"/workspace/data/pool.jsonl\"\nDEV = \"/workspace/data/multi_dev.npy\"\nOUT = \"/workspace/submission/selection.json\"\nCACHE_DIR = \"/tmp/curate_cache\"\nBUDGET = 12_000_000\nTECH_SHARE = 0.25 # target token share of technical Q&A register\nD = 1 << 20 # hashed feature dimension\nWORD_CAP = 1200 # cap words/doc for featurization + stats (bounds cost)\nALPHA = 1.0 # NB Laplace smoothing\nCLIP = 5.0 # per-feature log-odds clip\nN_OVERFLOW_IDS = 45_000 # total ids to emit (well over budget for safety)\n\nwre = re.compile(r\"[a-z0-9']+\")\nSTOP = set(\"the a an and or but if then else of to in on at by for with from as is are was \"\n \"were be been being it its this that these those he she they we you i his her their \"\n \"our your not no do does did has have had will would can could should may might must \"\n \"there here what which who whom whose when where why how than so such about into over \"\n \"after before between out up down off again more most other some any each also\".split())\nNSTAT = 11 # number of per-doc stat columns\n\n\ndef doc_features(text):\n \"\"\"Return (unique_feature_idx uint32[], stat_vector float32[NSTAT], fingerprint uint32).\"\"\"\n low = text.lower()\n w = wre.findall(low)\n if len(w) > WORD_CAP:\n w = w[:WORD_CAP]\n nw = len(w)\n idx = set()\n for tk in w:\n idx.add(zlib.crc32(tk.encode()) & (D - 1))\n for i in range(nw - 1):\n idx.add(zlib.crc32((w[i] + \"\\x1f\" + w[i + 1]).encode()) & (D - 1))\n\n seg = text[:4000]\n L = max(1, len(seg))\n alpha = sum(c.isalpha() for c in seg)\n digit = sum(c.isdigit() for c in seg)\n ascii_ = sum(ord(c) < 128 for c in seg)\n space = sum(c.isspace() for c in seg)\n sym = L - alpha - digit - space\n stops = sum(1 for x in w if x in STOP)\n uniq = len(set(w))\n mwl = (sum(len(x) for x in w) / nw) if nw else 0.0\n # line structure: fraction of non-empty lines that are \"short\" (< 5 words)\n lines = [ln for ln in text[:8000].split(\"\\n\") if ln.strip()]\n if lines:\n frac_short = sum(1 for ln in lines if len(ln.split()) < 5) / len(lines)\n else:\n frac_short = 1.0\n has_code = 1.0 if ((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)\n or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)) else 0.0\n stat = np.array([nw, max(1, len(text) // 4), alpha / L, digit / L, ascii_ / L,\n sym / L, (stops / nw) if nw else 0.0, (uniq / nw) if nw else 0.0,\n mwl, frac_short, has_code], dtype=np.float32)\n fp = zlib.crc32((\" \".join(w[:25] + w[-25:]) + \"|\" + str(nw)).encode()) & 0xFFFFFFFF\n return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stat, np.uint32(fp)\n\n\ndef concat_idx(arrs):\n lens = np.fromiter((len(a) for a in arrs), dtype=np.int64, count=len(arrs))\n indptr = np.zeros(len(arrs) + 1, dtype=np.int64)\n np.cumsum(lens, out=indptr[1:])\n allidx = np.concatenate(arrs).astype(np.int64) if arrs else np.zeros(0, np.int64)\n return allidx, indptr\n\n\ndef df_from(allidx):\n return np.bincount(allidx, minlength=D).astype(np.float64)\n\n\ndef mean_score(r, allidx, indptr):\n \"\"\"Per-doc MEAN of clipped feature weights r (== length-normalized log-lik ratio).\"\"\"\n if len(allidx) == 0:\n return np.zeros(len(indptr) - 1)\n vals = np.clip(r[allidx], -CLIP, CLIP)\n ssum = np.add.reduceat(vals, indptr[:-1])\n cnt = np.diff(indptr).astype(np.float64)\n cnt[cnt == 0] = 1.0\n return ssum / cnt\n\n\ndef nb_logodds(pos_df, bg_df, npos, nbg):\n p = (pos_df + ALPHA) / (npos + 2 * ALPHA)\n q = (bg_df + ALPHA) / (nbg + 2 * ALPHA)\n return np.log(p) - np.log(q)\n\n\ndef build_cache():\n tok = AutoTokenizer.from_pretrained(\"gpt2\")\n eos = tok.eos_token_id\n dev = np.load(DEV)\n cut = np.where(dev == eos)[0]\n starts = np.concatenate([[0], cut + 1]); ends = np.concatenate([cut, [len(dev)]])\n pos_prose, pos_tech = [], []\n for s, e in zip(starts, ends):\n if e - s < 3:\n continue\n a, st, _ = doc_features(tok.decode(dev[s:e].tolist()))\n (pos_tech if st[10] > 0 else pos_prose).append(a)\n print(f\" dev positives: prose={len(pos_prose)} tech={len(pos_tech)}\", flush=True)\n pa, _ = concat_idx(pos_prose); ta, _ = concat_idx(pos_tech)\n prose_df = df_from(pa); tech_df = df_from(ta)\n\n ids, arrs, stats, fps, heads = [], [], [], [], []\n for line in open(POOL):\n r = json.loads(line)\n a, st, fp = doc_features(r[\"text\"])\n ids.append(r[\"id\"]); arrs.append(a); stats.append(st); fps.append(fp)\n heads.append(r[\"text\"][:220])\n ids = np.array(ids, dtype=np.int64)\n allidx, indptr = concat_idx(arrs)\n stats = np.stack(stats); fps = np.array(fps, dtype=np.uint32)\n os.makedirs(CACHE_DIR, exist_ok=True)\n np.save(f\"{CACHE_DIR}/ids.npy\", ids)\n np.save(f\"{CACHE_DIR}/allidx.npy\", allidx)\n np.save(f\"{CACHE_DIR}/indptr.npy\", indptr)\n np.save(f\"{CACHE_DIR}/stats.npy\", stats)\n np.save(f\"{CACHE_DIR}/fps.npy\", fps)\n np.save(f\"{CACHE_DIR}/prose_df.npy\", prose_df)\n np.save(f\"{CACHE_DIR}/tech_df.npy\", tech_df)\n np.save(f\"{CACHE_DIR}/meta.npy\", np.array([len(pos_prose), len(pos_tech)], dtype=np.int64))\n with open(f\"{CACHE_DIR}/heads.json\", \"w\") as f:\n json.dump(heads, f)\n return ids, allidx, indptr, stats, fps, prose_df, tech_df, len(pos_prose), len(pos_tech), heads\n\n\ndef load_cache():\n ids = np.load(f\"{CACHE_DIR}/ids.npy\")\n allidx = np.load(f\"{CACHE_DIR}/allidx.npy\")\n indptr = np.load(f\"{CACHE_DIR}/indptr.npy\")\n stats = np.load(f\"{CACHE_DIR}/stats.npy\")\n fps = np.load(f\"{CACHE_DIR}/fps.npy\")\n prose_df = np.load(f\"{CACHE_DIR}/prose_df.npy\")\n tech_df = np.load(f\"{CACHE_DIR}/tech_df.npy\")\n meta = np.load(f\"{CACHE_DIR}/meta.npy\")\n heads = json.load(open(f\"{CACHE_DIR}/heads.json\"))\n return ids, allidx, indptr, stats, fps, prose_df, tech_df, int(meta[0]), int(meta[1]), heads\n\n\ndef main():\n t0 = time.time()\n if os.path.exists(f\"{CACHE_DIR}/ids.npy\") and \"--rebuild\" not in sys.argv:\n print(\"[1] loading feature cache\", flush=True)\n ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = load_cache()\n else:\n print(\"[1] featurizing dev positives + pool (one-time)\", flush=True)\n ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = build_cache()\n n = len(ids)\n bg_df = df_from(allidx); nbg = n\n print(f\" N={n} feats in {time.time()-t0:.0f}s\", flush=True)\n\n r_prose = nb_logodds(prose_df, bg_df, npp, nbg)\n r_tech = nb_logodds(tech_df, bg_df, npt, nbg)\n s_prose = mean_score(r_prose, allidx, indptr)\n s_tech = mean_score(r_tech, allidx, indptr)\n\n nw, tok_est, alpha, digit, ascii_, sym, stopfrac, uniq, mwl, fshort, has_code = \\\n [stats[:, i] for i in range(NSTAT)]\n\n # register-aware junk filters\n prose_ok = ((nw >= 50) & (alpha >= 0.62) & (stopfrac >= 0.20) & (uniq >= 0.30) &\n (uniq <= 0.92) & (mwl <= 10.0) & (ascii_ >= 0.92) & (sym <= 0.12) &\n (fshort <= 0.55) & (digit <= 0.20))\n tech_ok = ((nw >= 40) & (alpha >= 0.40) & (stopfrac >= 0.10) & (uniq >= 0.28) &\n (uniq <= 0.92) & (ascii_ >= 0.90) & (mwl <= 12.0) & (fshort <= 0.75))\n\n zp = (s_prose - s_prose.mean()) / (s_prose.std() + 1e-9)\n zt = (s_tech - s_tech.mean()) / (s_tech.std() + 1e-9)\n is_tech = zt > zp\n\n tech_mask = is_tech & tech_ok\n prose_mask = (~is_tech) & prose_ok\n tech_pos = np.where(tech_mask)[0][np.argsort(-s_tech[tech_mask])]\n prose_pos = np.where(prose_mask)[0][np.argsort(-s_prose[prose_mask])]\n print(f\"[2] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"\n f\"(dropped {n-len(prose_pos)-len(tech_pos)})\", flush=True)\n\n # token-balanced interleave to 75/25 with content-fingerprint dedup, then overflow\n tech_budget = BUDGET * TECH_SHARE\n prose_budget = BUDGET * (1 - TECH_SHARE)\n seen_fp = set()\n sel = []\n pt = pi = 0\n tacc = pacc = 0\n\n def take(order, ptr):\n while ptr < len(order):\n j = order[ptr]; ptr += 1\n fp = int(fps[j])\n if fp in seen_fp:\n continue\n seen_fp.add(fp)\n return j, ptr\n return None, ptr\n\n while (tacc < tech_budget or pacc < prose_budget) and (pt < len(tech_pos) or pi < len(prose_pos)):\n want_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))\n if want_tech and pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j); tacc += tok_est[j]\n elif pi < len(prose_pos):\n j, pi = take(prose_pos, pi)\n if j is not None: sel.append(j); pacc += tok_est[j]\n elif pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j); tacc += tok_est[j]\n print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M est tokens\",\n flush=True)\n\n # overflow: append remaining ranked docs (3 prose : 1 tech), deduped\n while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) < N_OVERFLOW_IDS:\n for _ in range(3):\n if pi < len(prose_pos):\n j, pi = take(prose_pos, pi)\n if j is not None: sel.append(j)\n if pt < len(tech_pos):\n j, pt = take(tech_pos, pt)\n if j is not None: sel.append(j)\n\n out_ids = [int(ids[j]) for j in sel]\n assert len(out_ids) == len(set(out_ids)), \"duplicate ids\"\n json.dump(out_ids, open(OUT, \"w\"))\n est = sum(int(tok_est[j]) for j in sel)\n print(f\"[3] wrote {len(out_ids)} ids -> {OUT} (~{est/1e6:.0f}M est tokens) in {time.time()-t0:.0f}s\",\n flush=True)\n\n if \"--peek\" in sys.argv:\n print(\"\\n===== TOP PROSE =====\")\n for j in prose_pos[:6]:\n print(f\"[s{s_prose[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"\n f\"fs{fshort[j]:.2f}] {heads[j]!r}\")\n print(\"\\n===== TOP TECH =====\")\n for j in tech_pos[:6]:\n print(f\"[s{s_tech[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"\n f\"fs{fshort[j]:.2f}] {heads[j]!r}\")\n print(\"\\n===== MID PROSE (rank ~8000) =====\")\n for j in prose_pos[8000:8004]:\n print(f\"[s{s_prose[j]:.2f}] {heads[j]!r}\")\n\n\nif __name__ == \"__main__\":\n main()\n", "structuredPatch": [{"oldStart": 6, "oldLines": 39, "newStart": 6, "newLines": 47, "lines": [" ", " STATED CRITERION", " ----------------", "-Keep the pool documents that are (a) clean, fluent, information-dense English and", "-(b) distributionally close to the disclosed target. We measure target-closeness", "-with an NBSVM-style log-odds score: for hashed word unigram+bigram features, weigh", "-each feature by log P(feature | target register) - log P(feature | pool background).", "-A document's score is the sum of its feature weights == its log-likelihood ratio of", "-being drawn from the target register vs. the generic web pool (a DSIR-style signal).", "+Keep pool documents that are (a) clean, fluent, information-dense English and", "+(b) distributionally close to the disclosed target register mix.", " ", "+Target-closeness is a DSIR / NBSVM-style log-likelihood ratio. For hashed word", "+unigram+bigram features we weight each feature f by", "+ r[f] = log P(f | target register) - log P(f | pool background)", "+estimated from feature document-frequencies (Laplace-smoothed). A document's score", "+is the MEAN over its features of clip(r[f], -5, 5). Using the mean (not the sum)", "+removes length bias, and clipping stops a handful of ultra-rare \"fancy\" words from", "+letting SEO word-salad win -- the two dominant failure modes of a raw sum.", "+", " The target is a fixed mixture whose one unambiguous slice is technical Q&A (~25% of", "-target tokens, identifiable by HTML/code markup). Programming/Q&A prose is", "-stylistically distinct from encyclopedic/news/web prose, so we score with TWO heads:", "- - PROSE head: positives = decoded dev segments WITHOUT code/HTML markup", "- - TECH head: positives = decoded dev segments WITH code/HTML markup", "-Each pool doc is assigned to the head it looks most like, passes a register-aware", "-junk filter, and is ranked by that head's score. We then fill the 12M-token budget", "-token-balanced 75% prose / 25% tech to match the target mixture, interleaving so any", "-early budget cutoff stays balanced. Remaining ranked ids are appended as overflow.", "+target tokens, identifiable by HTML/code markup in the decoded dev set). Programming/", "+Q&A text is stylistically distinct from encyclopedic/news/web prose, so we score with", "+TWO heads sharing the same features:", "+ PROSE head : positives = decoded dev segments WITHOUT code/HTML markup", "+ TECH head : positives = decoded dev segments WITH code/HTML markup", "+Each pool doc is assigned to whichever head it looks most like (standardized score),", "+must pass a register-aware junk filter (drops word-salad, menu/boilerplate, non-English,", "+markup dumps, repetitive spam), is deduplicated by a content fingerprint, and is ranked", "+by that head's score. We then fill the 12M-token budget token-balanced 75% prose / 25%", "+tech to match the target mixture, interleaving so any early budget cutoff stays balanced.", "+Remaining ranked ids are appended as overflow so the packer never runs short.", " ", "-The positives come only from the disclosed dev target; the selection is over disjoint", "-pool ids. Output is an ordered id list, produced entirely by this script (no hand", "-picking).", "+Positives come only from the disclosed dev target; the selection is over disjoint pool", "+ids. The output is an ordered id list produced entirely by this script.", " \"\"\"", "-import json, re, zlib, sys, time", "+import json, re, zlib, sys, time, os", " import numpy as np", " from transformers import AutoTokenizer", " ", " POOL = \"/workspace/data/pool.jsonl\"", " DEV = \"/workspace/data/multi_dev.npy\"", " OUT = \"/workspace/submission/selection.json\"", "+CACHE_DIR = \"/tmp/curate_cache\"", " BUDGET = 12_000_000", " TECH_SHARE = 0.25 # target token share of technical Q&A register", " D = 1 << 20 # hashed feature dimension", " WORD_CAP = 1200 # cap words/doc for featurization + stats (bounds cost)", "-ALPHA = 1.0 # NB smoothing", "+ALPHA = 1.0 # NB Laplace smoothing", "+CLIP = 5.0 # per-feature log-odds clip", " N_OVERFLOW_IDS = 45_000 # total ids to emit (well over budget for safety)", " ", " wre = re.compile(r\"[a-z0-9']+\")"]}, {"oldStart": 46, "oldLines": 10, "newStart": 54, "newLines": 12, "lines": [" \"were be been being it its this that these those he she they we you i his her their \"", " \"our your not no do does did has have had will would can could should may might must \"", " \"there here what which who whom whose when where why how than so such about into over \"", "- \"after before between out up down off again more most other some any each\".split())", "+ \"after before between out up down off again more most other some any each also\".split())", "+NSTAT = 11 # number of per-doc stat columns", " ", "+", " def doc_features(text):", "- \"\"\"Return (feature_index_array_uint32, stats_dict) for one document.\"\"\"", "+ \"\"\"Return (unique_feature_idx uint32[], stat_vector float32[NSTAT], fingerprint uint32).\"\"\"", " low = text.lower()", " w = wre.findall(low)", " if len(w) > WORD_CAP:"]}, {"oldStart": 59, "oldLines": 8, "newStart": 69, "newLines": 8, "lines": [" for tk in w:", " idx.add(zlib.crc32(tk.encode()) & (D - 1))", " for i in range(nw - 1):", "- idx.add(zlib.crc32((w[i] + \" \" + w[i + 1]).encode()) & (D - 1))", "- # char-level quality stats over a bounded prefix", "+ idx.add(zlib.crc32((w[i] + \"\\x1f\" + w[i + 1]).encode()) & (D - 1))", "+", " seg = text[:4000]", " L = max(1, len(seg))", " alpha = sum(c.isalpha() for c in seg)"]}, {"oldStart": 71, "oldLines": 182, "newStart": 81, "newLines": 201, "lines": [" stops = sum(1 for x in w if x in STOP)", " uniq = len(set(w))", " mwl = (sum(len(x) for x in w) / nw) if nw else 0.0", "- stats = dict(", "- nw=nw,", "- tok_est=max(1, len(text) // 4),", "- alpha=alpha / L,", "- digit=digit / L,", "- ascii=ascii_ / L,", "- sym=sym / L,", "- stopfrac=(stops / nw) if nw else 0.0,", "- uniq=(uniq / nw) if nw else 0.0,", "- mwl=mwl,", "- has_code=((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)", "- or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)),", "- )", "- return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stats", "+ # line structure: fraction of non-empty lines that are \"short\" (< 5 words)", "+ lines = [ln for ln in text[:8000].split(\"\\n\") if ln.strip()]", "+ if lines:", "+ frac_short = sum(1 for ln in lines if len(ln.split()) < 5) / len(lines)", "+ else:", "+ frac_short = 1.0", "+ has_code = 1.0 if ((\"<code>\" in text) or (\"</code>\" in text) or (\"<pre>\" in text)", "+ or (\"<\" in text) or (\"</\" in text) or (\"<p>\" in text)) else 0.0", "+ stat = np.array([nw, max(1, len(text) // 4), alpha / L, digit / L, ascii_ / L,", "+ sym / L, (stops / nw) if nw else 0.0, (uniq / nw) if nw else 0.0,", "+ mwl, frac_short, has_code], dtype=np.float32)", "+ fp = zlib.crc32((\" \".join(w[:25] + w[-25:]) + \"|\" + str(nw)).encode()) & 0xFFFFFFFF", "+ return np.fromiter(idx, dtype=np.uint32, count=len(idx)), stat, np.uint32(fp)", " ", " ", "-def concat_idx(list_of_idx_arrays):", "- \"\"\"Flatten per-doc unique feature indices into (allidx, indptr) CSR-row layout.\"\"\"", "- lens = np.fromiter((len(a) for a in list_of_idx_arrays), dtype=np.int64,", "- count=len(list_of_idx_arrays))", "- indptr = np.zeros(len(list_of_idx_arrays) + 1, dtype=np.int64)", "+def concat_idx(arrs):", "+ lens = np.fromiter((len(a) for a in arrs), dtype=np.int64, count=len(arrs))", "+ indptr = np.zeros(len(arrs) + 1, dtype=np.int64)", " np.cumsum(lens, out=indptr[1:])", "- allidx = (np.concatenate(list_of_idx_arrays).astype(np.int64)", "- if list_of_idx_arrays else np.zeros(0, np.int64))", "+ allidx = np.concatenate(arrs).astype(np.int64) if arrs else np.zeros(0, np.int64)", " return allidx, indptr", " ", " ", " def df_from(allidx):", "- \"\"\"Document frequency per feature (idx arrays are per-doc unique -> bincount == df).\"\"\"", " return np.bincount(allidx, minlength=D).astype(np.float64)", " ", " ", "-def score_docs(r, allidx, indptr):", "- \"\"\"Per-doc sum of feature weights r over the CSR rows == log-likelihood-ratio score.\"\"\"", "+def mean_score(r, allidx, indptr):", "+ \"\"\"Per-doc MEAN of clipped feature weights r (== length-normalized log-lik ratio).\"\"\"", " if len(allidx) == 0:", " return np.zeros(len(indptr) - 1)", "- return np.add.reduceat(r[allidx], indptr[:-1])", "+ vals = np.clip(r[allidx], -CLIP, CLIP)", "+ ssum = np.add.reduceat(vals, indptr[:-1])", "+ cnt = np.diff(indptr).astype(np.float64)", "+ cnt[cnt == 0] = 1.0", "+ return ssum / cnt", " ", " ", " def nb_logodds(pos_df, bg_df, npos, nbg):", "- \"\"\"NBSVM/DSIR-style per-feature log-odds: log p(f|target) - log p(f|background).\"\"\"", " p = (pos_df + ALPHA) / (npos + 2 * ALPHA)", " q = (bg_df + ALPHA) / (nbg + 2 * ALPHA)", " return np.log(p) - np.log(q)", " ", " ", "-def main():", "- t0 = time.time()", "- print(\"[1] decoding dev target -> positive register sets\", flush=True)", "+def build_cache():", " tok = AutoTokenizer.from_pretrained(\"gpt2\")", " eos = tok.eos_token_id", " dev = np.load(DEV)", " cut = np.where(dev == eos)[0]", " starts = np.concatenate([[0], cut + 1]); ends = np.concatenate([cut, [len(dev)]])", "- pos_prose_idx, pos_tech_idx = [], []", "+ pos_prose, pos_tech = [], []", " for s, e in zip(starts, ends):", " if e - s < 3:", " continue", "- txt = tok.decode(dev[s:e].tolist())", "- a, st = doc_features(txt)", "- (pos_tech_idx if st[\"has_code\"] else pos_prose_idx).append(a)", "- print(f\" dev positives: prose={len(pos_prose_idx)} tech={len(pos_tech_idx)}\", flush=True)", "+ a, st, _ = doc_features(tok.decode(dev[s:e].tolist()))", "+ (pos_tech if st[10] > 0 else pos_prose).append(a)", "+ print(f\" dev positives: prose={len(pos_prose)} tech={len(pos_tech)}\", flush=True)", "+ pa, _ = concat_idx(pos_prose); ta, _ = concat_idx(pos_tech)", "+ prose_df = df_from(pa); tech_df = df_from(ta)", " ", "- print(\"[2] streaming pool: featurize + quality stats\", flush=True)", "- ids, idx_arrays, stats = [], [], []", "- n = 0", "+ ids, arrs, stats, fps, heads = [], [], [], [], []", " for line in open(POOL):", " r = json.loads(line)", "- a, st = doc_features(r[\"text\"])", "- ids.append(r[\"id\"]); idx_arrays.append(a); stats.append(st)", "- n += 1", "- print(f\" pooled {n} docs in {time.time()-t0:.0f}s\", flush=True)", "-", "+ a, st, fp = doc_features(r[\"text\"])", "+ ids.append(r[\"id\"]); arrs.append(a); stats.append(st); fps.append(fp)", "+ heads.append(r[\"text\"][:220])", " ids = np.array(ids, dtype=np.int64)", "- allidx, indptr = concat_idx(idx_arrays)", "- bg_df = df_from(allidx) # document frequency over pool", "- nbg = len(idx_arrays)", "+ allidx, indptr = concat_idx(arrs)", "+ stats = np.stack(stats); fps = np.array(fps, dtype=np.uint32)", "+ os.makedirs(CACHE_DIR, exist_ok=True)", "+ np.save(f\"{CACHE_DIR}/ids.npy\", ids)", "+ np.save(f\"{CACHE_DIR}/allidx.npy\", allidx)", "+ np.save(f\"{CACHE_DIR}/indptr.npy\", indptr)", "+ np.save(f\"{CACHE_DIR}/stats.npy\", stats)", "+ np.save(f\"{CACHE_DIR}/fps.npy\", fps)", "+ np.save(f\"{CACHE_DIR}/prose_df.npy\", prose_df)", "+ np.save(f\"{CACHE_DIR}/tech_df.npy\", tech_df)", "+ np.save(f\"{CACHE_DIR}/meta.npy\", np.array([len(pos_prose), len(pos_tech)], dtype=np.int64))", "+ with open(f\"{CACHE_DIR}/heads.json\", \"w\") as f:", "+ json.dump(heads, f)", "+ return ids, allidx, indptr, stats, fps, prose_df, tech_df, len(pos_prose), len(pos_tech), heads", " ", "- # per-head log-odds weights", "- pa, _ = concat_idx(pos_prose_idx); ta, _ = concat_idx(pos_tech_idx)", "- r_prose = nb_logodds(df_from(pa), bg_df, len(pos_prose_idx), nbg)", "- r_tech = nb_logodds(df_from(ta), bg_df, len(pos_tech_idx), nbg)", "- s_prose = score_docs(r_prose, allidx, indptr)", "- s_tech = score_docs(r_tech, allidx, indptr)", " ", "- # gather stats into arrays", "- def col(k): return np.array([s[k] for s in stats])", "- nw = col(\"nw\"); tok_est = col(\"tok_est\"); alpha = col(\"alpha\"); digit = col(\"digit\")", "- ascii_ = col(\"ascii\"); sym = col(\"sym\"); stopfrac = col(\"stopfrac\"); uniq = col(\"uniq\")", "- mwl = col(\"mwl\"); has_code = col(\"has_code\").astype(bool)", "+def load_cache():", "+ ids = np.load(f\"{CACHE_DIR}/ids.npy\")", "+ allidx = np.load(f\"{CACHE_DIR}/allidx.npy\")", "+ indptr = np.load(f\"{CACHE_DIR}/indptr.npy\")", "+ stats = np.load(f\"{CACHE_DIR}/stats.npy\")", "+ fps = np.load(f\"{CACHE_DIR}/fps.npy\")", "+ prose_df = np.load(f\"{CACHE_DIR}/prose_df.npy\")", "+ tech_df = np.load(f\"{CACHE_DIR}/tech_df.npy\")", "+ meta = np.load(f\"{CACHE_DIR}/meta.npy\")", "+ heads = json.load(open(f\"{CACHE_DIR}/heads.json\"))", "+ return ids, allidx, indptr, stats, fps, prose_df, tech_df, int(meta[0]), int(meta[1]), heads", " ", "+", "+def main():", "+ t0 = time.time()", "+ if os.path.exists(f\"{CACHE_DIR}/ids.npy\") and \"--rebuild\" not in sys.argv:", "+ print(\"[1] loading feature cache\", flush=True)", "+ ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = load_cache()", "+ else:", "+ print(\"[1] featurizing dev positives + pool (one-time)\", flush=True)", "+ ids, allidx, indptr, stats, fps, prose_df, tech_df, npp, npt, heads = build_cache()", "+ n = len(ids)", "+ bg_df = df_from(allidx); nbg = n", "+ print(f\" N={n} feats in {time.time()-t0:.0f}s\", flush=True)", "+", "+ r_prose = nb_logodds(prose_df, bg_df, npp, nbg)", "+ r_tech = nb_logodds(tech_df, bg_df, npt, nbg)", "+ s_prose = mean_score(r_prose, allidx, indptr)", "+ s_tech = mean_score(r_tech, allidx, indptr)", "+", "+ nw, tok_est, alpha, digit, ascii_, sym, stopfrac, uniq, mwl, fshort, has_code = \\", "+ [stats[:, i] for i in range(NSTAT)]", "+", " # register-aware junk filters", "- prose_ok = ((nw >= 50) & (alpha >= 0.60) & (stopfrac >= 0.12) & (uniq >= 0.30) &", "- (mwl <= 11.0) & (ascii_ >= 0.90) & (sym <= 0.15))", "- tech_ok = ((nw >= 30) & (alpha >= 0.35) & (uniq >= 0.25) & (ascii_ >= 0.85) &", "- (mwl <= 14.0))", "+ prose_ok = ((nw >= 50) & (alpha >= 0.62) & (stopfrac >= 0.20) & (uniq >= 0.30) &", "+ (uniq <= 0.92) & (mwl <= 10.0) & (ascii_ >= 0.92) & (sym <= 0.12) &", "+ (fshort <= 0.55) & (digit <= 0.20))", "+ tech_ok = ((nw >= 40) & (alpha >= 0.40) & (stopfrac >= 0.10) & (uniq >= 0.28) &", "+ (uniq <= 0.92) & (ascii_ >= 0.90) & (mwl <= 12.0) & (fshort <= 0.75))", " ", "- # assign each doc to the head it scores higher on (standardized)", " zp = (s_prose - s_prose.mean()) / (s_prose.std() + 1e-9)", " zt = (s_tech - s_tech.mean()) / (s_tech.std() + 1e-9)", "- is_tech = (zt > zp)", "+ is_tech = zt > zp", " ", "- # candidate lists", " tech_mask = is_tech & tech_ok", " prose_mask = (~is_tech) & prose_ok", "- tech_order = np.argsort(-s_tech[tech_mask])", "- prose_order = np.argsort(-s_prose[prose_mask])", "- tech_pos = np.where(tech_mask)[0][tech_order]", "- prose_pos = np.where(prose_mask)[0][prose_order]", "- print(f\"[3] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"", "- f\"(dropped {n-len(prose_pos)-len(tech_pos)} by filter/assignment)\", flush=True)", "+ tech_pos = np.where(tech_mask)[0][np.argsort(-s_tech[tech_mask])]", "+ prose_pos = np.where(prose_mask)[0][np.argsort(-s_prose[prose_mask])]", "+ print(f\"[2] candidates: prose={len(prose_pos)} tech={len(tech_pos)} \"", "+ f\"(dropped {n-len(prose_pos)-len(tech_pos)})\", flush=True)", " ", "- # exact-duplicate guard via cheap fingerprint of stats (nw,alpha rounded) + first score;", "- # true dedup done on text hash below during emission.", "- # token-balanced interleave to 75/25, then overflow.", "+ # token-balanced interleave to 75/25 with content-fingerprint dedup, then overflow", " tech_budget = BUDGET * TECH_SHARE", " prose_budget = BUDGET * (1 - TECH_SHARE)", "- sel, seen = [], set()", "+ seen_fp = set()", "+ sel = []", " pt = pi = 0", " tacc = pacc = 0", "- # phase 1: fill budget balanced", "- while (tacc < tech_budget or pacc < prose_budget):", "- take_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))", "- if take_tech and pt < len(tech_pos):", "- j = tech_pos[pt]; pt += 1", "- if ids[j] not in seen:", "- seen.add(ids[j]); sel.append(j); tacc += tok_est[j]", "- elif (not take_tech) and pi < len(prose_pos):", "- j = prose_pos[pi]; pi += 1", "- if ids[j] not in seen:", "- seen.add(ids[j]); sel.append(j); pacc += tok_est[j]", "- else:", "- # one side exhausted; drain the other", "- if pt < len(tech_pos):", "- j = tech_pos[pt]; pt += 1", "- if ids[j] not in seen: seen.add(ids[j]); sel.append(j); tacc += tok_est[j]", "- elif pi < len(prose_pos):", "- j = prose_pos[pi]; pi += 1", "- if ids[j] not in seen: seen.add(ids[j]); sel.append(j); pacc += tok_est[j]", "- else:", "- break", "- print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M tokens\", flush=True)", " ", "- # phase 2: overflow — append remaining ranked docs (prose then tech interleaved) for margin", "- rest = []", "- while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) + len(rest) < N_OVERFLOW_IDS:", "+ def take(order, ptr):", "+ while ptr < len(order):", "+ j = order[ptr]; ptr += 1", "+ fp = int(fps[j])", "+ if fp in seen_fp:", "+ continue", "+ seen_fp.add(fp)", "+ return j, ptr", "+ return None, ptr", "+", "+ while (tacc < tech_budget or pacc < prose_budget) and (pt < len(tech_pos) or pi < len(prose_pos)):", "+ want_tech = (tacc / max(1, tech_budget)) <= (pacc / max(1, prose_budget))", "+ if want_tech and pt < len(tech_pos):", "+ j, pt = take(tech_pos, pt)", "+ if j is not None: sel.append(j); tacc += tok_est[j]", "+ elif pi < len(prose_pos):", "+ j, pi = take(prose_pos, pi)", "+ if j is not None: sel.append(j); pacc += tok_est[j]", "+ elif pt < len(tech_pos):", "+ j, pt = take(tech_pos, pt)", "+ if j is not None: sel.append(j); tacc += tok_est[j]", "+ print(f\" budget phase: {len(sel)} docs, tech~{tacc/1e6:.1f}M prose~{pacc/1e6:.1f}M est tokens\",", "+ flush=True)", "+", "+ # overflow: append remaining ranked docs (3 prose : 1 tech), deduped", "+ while (pi < len(prose_pos) or pt < len(tech_pos)) and len(sel) < N_OVERFLOW_IDS:", " for _ in range(3):", " if pi < len(prose_pos):", "- j = prose_pos[pi]; pi += 1", "- if ids[j] not in seen: seen.add(ids[j]); rest.append(j)", "+ j, pi = take(prose_pos, pi)", "+ if j is not None: sel.append(j)", " if pt < len(tech_pos):", "- j = tech_pos[pt]; pt += 1", "- if ids[j] not in seen: seen.add(ids[j]); rest.append(j)", "- sel.extend(rest)", "+ j, pt = take(tech_pos, pt)", "+ if j is not None: sel.append(j)", " ", " out_ids = [int(ids[j]) for j in sel]", "+ assert len(out_ids) == len(set(out_ids)), \"duplicate ids\"", " json.dump(out_ids, open(OUT, \"w\"))", "- est_tokens = sum(int(tok_est[j]) for j in sel)", "- print(f\"[4] wrote {len(out_ids)} ids -> {OUT} (~{est_tokens/1e6:.0f}M est tokens) \"", "- f\"in {time.time()-t0:.0f}s\", flush=True)", "+ est = sum(int(tok_est[j]) for j in sel)", "+ print(f\"[3] wrote {len(out_ids)} ids -> {OUT} (~{est/1e6:.0f}M est tokens) in {time.time()-t0:.0f}s\",", "+ flush=True)", " ", "- # diagnostics: peek at top docs of each head", " if \"--peek\" in sys.argv:", "- texts = {}", "- for line in open(POOL):", "- r = json.loads(line); texts[r[\"id\"]] = r[\"text\"]", " print(\"\\n===== TOP PROSE =====\")", "- for j in prose_pos[:4]:", "- print(f\"[score {s_prose[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")", "+ for j in prose_pos[:6]:", "+ print(f\"[s{s_prose[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"", "+ f\"fs{fshort[j]:.2f}] {heads[j]!r}\")", " print(\"\\n===== TOP TECH =====\")", "- for j in tech_pos[:4]:", "- print(f\"[score {s_tech[j]:.1f} nw {int(nw[j])}] {texts[int(ids[j])][:240]!r}\")", "- print(\"\\n===== BOTTOM (lowest prose score, any) =====\")", "- worst = np.argsort(s_prose)[:3]", "- for j in worst:", "- print(f\"[score {s_prose[j]:.1f}] {texts[int(ids[j])][:200]!r}\")", "+ for j in tech_pos[:6]:", "+ print(f\"[s{s_tech[j]:.2f} nw{int(nw[j])} sw{stopfrac[j]:.2f} uq{uniq[j]:.2f} \"", "+ f\"fs{fshort[j]:.2f}] {heads[j]!r}\")", "+ print(\"\\n===== MID PROSE (rank ~8000) =====\")", "+ for j in prose_pos[8000:8004]:", "+ print(f\"[s{s_prose[j]:.2f}] {heads[j]!r}\")", " ", " ", " if __name__ == \"__main__\":"]}], "originalFile": null, "userModified": false}Now let me run the improved curate.py (it will build the cache this first time, ~160s). I'll write output to a file directly this time.
args
{
"command": "cd /workspace; export TOKENIZERS_PARALLELISM=false; (python3 curate.py --peek > /tmp/curate_run.log 2>&1; echo \"EXIT $?\" >> /tmp/curate_run.log) &\necho \"launched pid $!\"",
"description": "Run improved curate.py in background",
"run_in_background": true
}Bash result
Command running in background with ID: bo4n0n6xk. Output is being written to: /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/bo4n0n6xk.output. You will be notified when it completes. To check interim output, use Read on that file path.
Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "bo4n0n6xk", "backgroundCwdHint": "Session cwd remains /workspace; directory changes made by the backgrounded command do not apply to subsequent commands."}While curate.py runs, let me reconsider one risk: if the pool has little genuine technical content, forcing 25% tech could scrape low-quality tech-ish docs. I'll inspect the tech candidates in the peek and be ready to test the split via training. Let me check progress.
args
{
"command": "for i in $(seq 1 30); do if grep -q \"EXIT\" /tmp/curate_run.log 2>/dev/null; then break; fi; sleep 6; done; cat /tmp/curate_run.log",
"description": "Wait for curate.py and read output"
}Bash result
Command did not complete within its 120s timeout and was moved to the background (ID: b15t493zm). Output is being written to: /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/b15t493zm.output. You will be notified when it completes. To check interim output, use Read on that file path.
[metadata] {"noOutputExpected": false, "backgroundTaskId": "b15t493zm", "timedOutAfterMs": 120000}encrypted chain-of-thought — recorded, not exposed
args
{
"command": "cat /tmp/curate_run.log",
"description": "Read curate.py peek output"
}Bash result
[1] featurizing dev positives + pool (one-time)
dev positives: prose=2345 tech=360
[stdout]
[1] featurizing dev positives + pool (one-time)
dev positives: prose=2345 tech=360
[metadata] {"noOutputExpected": false}The log seems truncated. Let me check the actual state — cache files, both output locations, and whether the process is really done.
args
{
"command": "echo \"=== curate_run.log ===\"; cat /tmp/curate_run.log; echo \"=== harness output ===\"; cat /tmp/claude-0/-workspace/241b2f82-9a7f-471b-beef-757ffa753d6d/tasks/bo4n0n6xk.output 2>/dev/null | tail -40; echo \"=== cache dir ===\"; ls -la /tmp/curate_cache/ 2>/dev/null; echo \"=== procs ===\"; ps aux | grep curate | grep -v grep | head; echo \"=== selection ===\"; ls -la /workspace/submission/; python3 -c \"import json; s=json.load(open('/workspace/submission/selection.json')); print('selection ids:', len(s), 'first5', s[:5])\" 2>&1 | head",
"description": "Inspect logs, cache, procs, selection"
}Bash result
=== curate_run.log ===
[1] featurizing dev positives + pool (one-time)
dev positives: prose=2345 tech=360
N=182016 feats in 168s
[2] candidates: prose=75962 tech=58637 (dropped 47417)
budget phase: 8676 docs, tech~3.0M prose~9.0M est tokens
[3] wrote 45000 ids -> /workspace/submission/selection.json (~54M est tokens) in 170s
===== TOP PROSE =====
[s0.40 nw209 sw0.34 uq0.42 fs0.00] '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India. Justices Sanjay Kishan Kaul, Navin Sinha, Mohan M Shantanagoudar, De'
[s0.38 nw124 sw0.25 uq0.73 fs0.00] 'Lahore:In the prospects of stoned to death incident of a Khanewal Woman, Inspector General of Police Punjab, Mohammad Habib-ur-Rehman has recommended the Government of Punjab for Judicial Inquiry into Maryam Bibi’s case.'
[s0.37 nw1200 sw0.35 uq0.37 fs0.36] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army '
[s0.36 nw96 sw0.45 uq0.54 fs0.00] 'ggor Domm II (375BDW-51ADW) is widely considered to be the greatest ruler in the history of the Nethereigons. He reigned the Northern Dividend of the Nethereigons from 310BDW to 7ADW, and then ruled the United Nethereigo'
[s0.33 nw170 sw0.36 uq0.62 fs0.00] 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the vandalising of its party headquarters office in the state by the Trina'
[s0.33 nw212 sw0.36 uq0.54 fs0.12] "New Delhi | ANI: Yes Bank promoter Rana Kapoor informed the Enforcement Directorate (ED) that he was 'forced' by a then Congress Union Minister to buy an MF Hussain painting from Priyanka Gandhi Vadra for Rs 2 crores.\nAs"
===== TOP TECH =====
[s2.00 nw1200 sw0.11 uq0.34 fs0.75] 'ineadóir In-athnuaite/Leabaithe a nascadh\nIt appears JavaScript is disabled. To get the most out of the website we recommend enabling JavaScript in your browser.\nÚsáidtear fianáin ar an láithreán gréasáin seo, trí leanúi'
[s1.94 nw1200 sw0.17 uq0.29 fs0.68] 'ontoer<|endoftext|>Placa Segurança Área De Barulho Intenso | AfixGraf\nJavaScript seems to be disabled in your browser.\nVocê precisa habilitar o Javascript no seu navegador para utilizar as funcionalidades deste site.\nLoj'
[s1.93 nw451 sw0.19 uq0.70 fs0.00] 'Follow-up and lochial Stephanus clamours his cutinization undressings unsold firstly. mesmeric Ted overtrusts, resep cara mengolah ubi jalar his czaritza enthralling eyeleting occasionally. biform Dennie bodges, her rese'
[s1.89 nw1200 sw0.18 uq0.63 fs0.00] 'Consumptive Emory pubs, his Indo-Iranian abjured cuing tongue-in-cheek. synchronous Renaldo overmanning, her dating girl chart falsify very fiercely. mushy Dabney staned her abjuring vernalises concretely? jurisdictive A'
[s1.88 nw1200 sw0.23 uq0.73 fs0.02] 'Archaeological dating methods\nFutilities must undress thitherward convened dismissively towerish tomato.\nPossessory tongs being very mechanically fugued calendering.\nBurdensomely tomentum consciously scavenged the differ'
[s1.85 nw586 sw0.18 uq0.72 fs0.00] 'herical Prasad scared, her ironize very darkly. jocular Raimund filters her headhunts mow dear? nodulated and access 2010 datenbank tutorial slightest Gustavo cool his venging or lethargize pathetically. unsolid acciai d'
===== MID PROSE (rank ~8000) =====
[s-0.08] ' (Reuters) - Canadian telecom services provider Mitel Networks Corp said it would buy mobile networking software maker Mavenir Systems Inc for $560 million in cash and stock.\nMitel earlier tried to buy ShoreTel Inc but t'
[s-0.08] 'Yankees general manager Brian Cashman paid a visit to his publicly humiliated wife and kids yesterday — the first time he’s been spotted in public since a serial stalker was charged with extorting money from him after an'
[s-0.08] '<|endoftext|>Eight months ago, Politico tagged a little-known trade association called Freedom Partners as “the Koch brothers’ secret bank.” The Arlington-Va.-based organization, founded in 2011, had been bankrolling a s'
[s-0.08] 'ORY PAGE. Britannica does not currently have an article on this topic.\nLearn about this topic in these articles:\nThe myths of the Izumo Cycle then begin to appear in the narration. Having angered the heavenly gods and ha'
EXIT 0
=== harness output ===
launched pid 1712
=== cache dir ===
total 986147
drwxr-xr-x 1 root root 220 Aug 4 03:43 .
drwxrwxrwx 1 root root 220 Aug 4 03:43 ..
-rw-r--r-- 1 root root 939045864 Aug 4 03:43 allidx.npy
-rw-r--r-- 1 root root 728192 Aug 4 03:43 fps.npy
-rw-r--r-- 1 root root 42338517 Aug 4 03:43 heads.json
-rw-r--r-- 1 root root 1456256 Aug 4 03:43 ids.npy
-rw-r--r-- 1 root root 1456264 Aug 4 03:43 indptr.npy
-rw-r--r-- 1 root root 144 Aug 4 03:43 meta.npy
-rw-r--r-- 1 root root 8388736 Aug 4 03:43 prose_df.npy
-rw-r--r-- 1 root root 8008832 Aug 4 03:43 stats.npy
-rw-r--r-- 1 root root 8388736 Aug 4 03:43 tech_df.npy
=== procs ===
root 154 0.0 0.0 12432 5680 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid exit "$rc"
root 156 0.0 0.0 12432 6060 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.0 0.0 6226752 550884 ? Sl 03:22 0:49 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
=== selection ===
total 530
drwxr-xr-x 1 root root 80 Aug 4 03:37 .
drwxr-xr-x 1 root root 260 Aug 4 03:40 ..
-rw-r--r-- 1 root root 320669 Aug 4 03:43 selection.json
-rw-r--r-- 1 root root 221560 Aug 4 03:31 selection_random.json
selection ids: 45000 first5 [180643, 58452, 23673, 82269, 88750]
[stdout]
=== curate_run.log ===
[1] featurizing dev positives + pool (one-time)
dev positives: prose=2345 tech=360
N=182016 feats in 168s
[2] candidates: prose=75962 tech=58637 (dropped 47417)
budget phase: 8676 docs, tech~3.0M prose~9.0M est tokens
[3] wrote 45000 ids -> /workspace/submission/selection.json (~54M est tokens) in 170s
===== TOP PROSE =====
[s0.40 nw209 sw0.34 uq0.42 fs0.00] '<|endoftext|>New Delhi: Five new Supreme Court judges were today sworn in, taking the top court’s strength to 28, including the Chief Justice of India. Justices Sanjay Kishan Kaul, Navin Sinha, Mohan M Shantanagoudar, De'
[s0.38 nw124 sw0.25 uq0.73 fs0.00] 'Lahore:In the prospects of stoned to death incident of a Khanewal Woman, Inspector General of Police Punjab, Mohammad Habib-ur-Rehman has recommended the Government of Punjab for Judicial Inquiry into Maryam Bibi’s case.'
[s0.37 nw1200 sw0.35 uq0.37 fs0.36] 'Hill 303 massacre\n|Hill 303 massacre|\nBodies of massacre victims gathered near Waegwan, South Korea, many with their hands still bound\n|Location||Hill 303, Waegwan, South Korea|\n|Date||August 17, 1950\n|Target||U.S. Army '
[s0.36 nw96 sw0.45 uq0.54 fs0.00] 'ggor Domm II (375BDW-51ADW) is widely considered to be the greatest ruler in the history of the Nethereigons. He reigned the Northern Dividend of the Nethereigons from 310BDW to 7ADW, and then ruled the United Nethereigo'
[s0.33 nw170 sw0.36 uq0.62 fs0.00] 'The Bharatiya Janata Party (BJP) leaders of the West Bengal unit will meet Union Home Minister Rajnath Singh on Thursday to lodge protest against the vandalising of its party headquarters office in the state by the Trina'
[s0.33 nw212 sw0.36 uq0.54 fs0.12] "New Delhi | ANI: Yes Bank promoter Rana Kapoor informed the Enforcement Directorate (ED) that he was 'forced' by a then Congress Union Minister to buy an MF Hussain painting from Priyanka Gandhi Vadra for Rs 2 crores.\nAs"
===== TOP TECH =====
[s2.00 nw1200 sw0.11 uq0.34 fs0.75] 'ineadóir In-athnuaite/Leabaithe a nascadh\nIt appears JavaScript is disabled. To get the most out of the website we recommend enabling JavaScript in your browser.\nÚsáidtear fianáin ar an láithreán gréasáin seo, trí leanúi'
[s1.94 nw1200 sw0.17 uq0.29 fs0.68] 'ontoer<|endoftext|>Placa Segurança Área De Barulho Intenso | AfixGraf\nJavaScript seems to be disabled in your browser.\nVocê precisa habilitar o Javascript no seu navegador para utilizar as funcionalidades deste site.\nLoj'
[s1.93 nw451 sw0.19 uq0.70 fs0.00] 'Follow-up and lochial Stephanus clamours his cutinization undressings unsold firstly. mesmeric Ted overtrusts, resep cara mengolah ubi jalar his czaritza enthralling eyeleting occasionally. biform Dennie bodges, her rese'
[s1.89 nw1200 sw0.18 uq0.63 fs0.00] 'Consumptive Emory pubs, his Indo-Iranian abjured cuing tongue-in-cheek. synchronous Renaldo overmanning, her dating girl chart falsify very fiercely. mushy Dabney staned her abjuring vernalises concretely? jurisdictive A'
[s1.88 nw1200 sw0.23 uq0.73 fs0.02] 'Archaeological dating methods\nFutilities must undress thitherward convened dismissively towerish tomato.\nPossessory tongs being very mechanically fugued calendering.\nBurdensomely tomentum consciously scavenged the differ'
[s1.85 nw586 sw0.18 uq0.72 fs0.00] 'herical Prasad scared, her ironize very darkly. jocular Raimund filters her headhunts mow dear? nodulated and access 2010 datenbank tutorial slightest Gustavo cool his venging or lethargize pathetically. unsolid acciai d'
===== MID PROSE (rank ~8000) =====
[s-0.08] ' (Reuters) - Canadian telecom services provider Mitel Networks Corp said it would buy mobile networking software maker Mavenir Systems Inc for $560 million in cash and stock.\nMitel earlier tried to buy ShoreTel Inc but t'
[s-0.08] 'Yankees general manager Brian Cashman paid a visit to his publicly humiliated wife and kids yesterday — the first time he’s been spotted in public since a serial stalker was charged with extorting money from him after an'
[s-0.08] '<|endoftext|>Eight months ago, Politico tagged a little-known trade association called Freedom Partners as “the Koch brothers’ secret bank.” The Arlington-Va.-based organization, founded in 2011, had been bankrolling a s'
[s-0.08] 'ORY PAGE. Britannica does not currently have an article on this topic.\nLearn about this topic in these articles:\nThe myths of the Izumo Cycle then begin to appear in the narration. Having angered the heavenly gods and ha'
EXIT 0
=== harness output ===
launched pid 1712
=== cache dir ===
total 986147
drwxr-xr-x 1 root root 220 Aug 4 03:43 .
drwxrwxrwx 1 root root 220 Aug 4 03:43 ..
-rw-r--r-- 1 root root 939045864 Aug 4 03:43 allidx.npy
-rw-r--r-- 1 root root 728192 Aug 4 03:43 fps.npy
-rw-r--r-- 1 root root 42338517 Aug 4 03:43 heads.json
-rw-r--r-- 1 root root 1456256 Aug 4 03:43 ids.npy
-rw-r--r-- 1 root root 1456264 Aug 4 03:43 indptr.npy
-rw-r--r-- 1 root root 144 Aug 4 03:43 meta.npy
-rw-r--r-- 1 root root 8388736 Aug 4 03:43 prose_df.npy
-rw-r--r-- 1 root root 8008832 Aug 4 03:43 stats.npy
-rw-r--r-- 1 root root 8388736 Aug 4 03:43 tech_df.npy
=== procs ===
root 154 0.0 0.0 12432 5680 ? Ss 03:22 0:00 bash -c rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid setsid bash -c 'set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '"'"'# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'"'"'"'"'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'"'"'"'"'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'"'"'"'"'"'"'t win and doesn'"'"'"'"'"'"'"'"'t distort averages. '"'"' 2>&1 </dev/null | tee /logs/agent/claude-code.txt' & child="$!" echo "$child" > /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid wait "$child" rc="$?" rm -f /tmp/ri-exec-c0b673edae74493cb6ed87fa33506a60.pid exit "$rc"
root 156 0.0 0.0 12432 6060 ? Ss 03:22 0:00 bash -c set -o pipefail; export PATH="$HOME/.local/bin:/usr/local/bin:$PATH"; claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- '# Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don'"'"'t assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool'"'"'s documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can'"'"'t win and doesn'"'"'t distort averages. ' 2>&1 </dev/null | tee /logs/agent/claude-code.txt
root 157 4.0 0.0 6226752 550884 ? Sl 03:22 0:49 claude --verbose --output-format=stream-json --permission-mode=bypassPermissions --disallowedTools WebSearch WebFetch --print -- # Task: Curate a raw web corpus to train the best small language model ## Context You are assembling the pretraining data for a small language model. You have a large pool of **raw web documents** at `/workspace/data/pool.jsonl` ??? one JSON object per line: `{"id": <int>, "text": "..."}`. The documents are unlabeled. The pool holds **182,016 documents (~200M GPT-2 tokens total)**. Compute is fixed: the model is trained on a **fixed budget of 12,000,000 tokens** ??? only ~6% of the pool, far smaller than the whole thing. So you cannot train on everything: you must decide *which* documents are worth training on. The training pipeline tokenizes your selection (GPT-2 BPE) in priority order and stops once it has accumulated 12M tokens. ## Goal Curate the pool: choose the documents that, trained on under the fixed budget, produce the **best language model**. Quality is measured by **held-out perplexity on a BROAD, multi-domain high-quality English target** ??? equal parts encyclopedic (Wikipedia), general high-quality web prose, news, and technical Q&A. This is your disclosed target: select data that makes the model good across *all* of these registers, not just one. Lower perplexity is better. ## Deliverables 1. `/workspace/submission/selection.json` ??? a JSON list of pool `id`s to train on, in **priority order** (best first). The training pipeline consumes your list in order until the **12M-token** budget is filled, then trains. Provide enough ids to cover at least 12M tokens (a few hundred MB of text is ample); ids beyond the budget are simply unused. 2. `/workspace/submission/curate.py` ??? the reproducible script that produced your selection from a stated criterion (not a hand-picked id list). 3. `/workspace/claim.md` ??? Hypothesis / Mechanism (predict an observable other than the final perplexity) / Falsification / Transfer. ## Resources - The **frozen training script** is provided (`/workspace/train_nano.py`, with `model.py`). It trains a ~30M-parameter GPT from scratch on a token budget with fixed hyperparameters. **You may run it to evaluate a candidate selection** against a dev target (`/workspace/data/multi_dev.npy`) ??? one run is ~60s on the GPU ??? but you **cannot change it**; the official run uses an identical frozen copy. The only thing you control is the data selection. - Standard GPU sandbox (Python, PyTorch, transformers). How you assess document quality is entirely up to you (filters, classifiers, heuristics, a model, etc.). - Your session ends when the wall-clock budget runs out ??? check `/workspace/.timer/remaining_secs` at any point for the authoritative time left (don't assume a fixed number of hours). Every training run you do comes out of that budget. - No internet. The pool's documents are freshly assembled; you cannot look up any quality labels ??? any quality signal must come from your own analysis. ## How you are scored Official: your selection fills the fixed token budget, the frozen script trains the 30M GPT from scratch, and the verifier measures the **absolute held-out perplexity on the hidden high-quality target** (a different sample from the disclosed domain than the dev set). That perplexity is your score. **Lower is better** ??? the goal is simply to push held-out perplexity as low as you can. The verifier reports the raw perplexity and does no normalization. For reference and for a separate downstream normalization step, the benchmark records two fixed anchors: a **baseline** of held-out perplexity from a random selection (the do-nothing reference), and a **theoretical best** of perplexity 1 (the perplexity floor, unreachable in practice). Aim to get your held-out perplexity well below the random baseline. ## Constraints - `selection.json` must be produced by `curate.py` from a stated criterion; an opaque hand-picked id list fails the publishability check even if it scores well. - A selection with duplicate or out-of-range ids, or too few tokens to fill the budget, is marked invalid (no score recorded) and excluded from ranking, so it can't win and doesn't distort averages.
=== selection ===
total 530
drwxr-xr-x 1 root root 80 Aug 4 03:37 .
drwxr-xr-x 1 root root 260 Aug 4 03:40 ..
-rw-r--r-- 1 root root 320669 Aug 4 03:43 selection.json
-rw-r--r-- 1 root root 221560 Aug 4 03:31 selection_random.json
selection ids: 45000 first5 [180643, 58452, 23673, 82269, 88750]
[metadata] {"noOutputExpected": false}