Skip to content

Task guidelines

How to design a task that fits RSI Bench. First version — expanded as the benchmark and contribution flow mature.

What makes a good task

Placeholder — full guidelines pending.
Real research work
Mirrors an actual R&D step (data, kernels, training, evals, alignment).
Objective verifier
Graded by a verifier against a measured baseline — not a subjective judge.
Fixed budget
A clearly bounded compute budget; no way to trivially brute-force it.
Discriminative
Separates strong from weak agents (meaningful pass@k spread).
Held-out safe
Public split never exposes the hidden test set or ground truth.

Taxonomy

Full 9-category taxonomy TBC.

Tasks are organized by first-level category — Infra & Systems, Post-training, Multimodal, Alignment, Applied, Data, Pre-training, Evals.