Task guidelines
How to design a task that fits RSI Bench. First version — expanded as the benchmark and contribution flow mature.
What makes a good task
Placeholder — full guidelines pending.
- Real research work
- Mirrors an actual R&D step (data, kernels, training, evals, alignment).
- Objective verifier
- Graded by a verifier against a measured baseline — not a subjective judge.
- Fixed budget
- A clearly bounded compute budget; no way to trivially brute-force it.
- Discriminative
- Separates strong from weak agents (meaningful pass@k spread).
- Held-out safe
- Public split never exposes the hidden test set or ground truth.
Taxonomy
Full 9-category taxonomy TBC.
Tasks are organized by first-level category — Infra & Systems, Post-training, Multimodal, Alignment, Applied, Data, Pre-training, Evals.