Index
54 benchmarks, 24,583 tasks. Pick your benchmark, then your task.
Abstract reasoning
| ARC-AGI-1 (public evaluation) | 400 tasks indexed |
|---|---|
| ARC-AGI-2 (public evaluation) | 120 tasks indexed |
Agentic terminal use
| Terminal-Bench 2.0 | 89 tasks indexed |
|---|---|
| Terminal-Bench Core 0.1.1 | 80 tasks indexed |
AI red teaming
| AIRTBench | 76 tasks indexed |
|---|
Code editing
| Aider Polyglot | 225 tasks indexed |
|---|
Code generation
| BigCodeBench | 1,140 tasks indexed |
|---|---|
| CodeContests | 165 tasks indexed |
| HumanEval | 164 tasks indexed |
| HumanEval+ | 164 tasks indexed |
| LiveCodeBench | looked up by identifier |
| MBPP (sanitized) | 257 tasks indexed |
| MBPP+ | 378 tasks indexed |
Code reasoning
| CRUXEval | 800 tasks indexed |
|---|
Computer use
| OSWorld | 369 tasks indexed |
|---|
Factuality
| SimpleQA | looked up by identifier |
|---|
General assistants
| GAIA | looked up by identifier |
|---|
Knowledge and reasoning
| GPQA Diamond | looked up by identifier |
|---|---|
| Humanity's Last Exam | looked up by identifier |
| MMLU-Pro | 12,032 tasks indexed |
Machine learning engineering
| MLE-bench | 82 tasks indexed |
|---|
Mathematics
| AIME 2024 | 30 tasks indexed |
|---|---|
| AIME 2025 | 30 tasks indexed |
| FrontierMath | looked up by identifier |
Offensive security
| BountyBench | 46 tasks indexed |
|---|---|
| Corporate Network (range, 32 steps) | looked up by identifier |
| Cybench | 43 tasks indexed |
| CyberSecEval 3 | looked up by identifier |
| CyScenarioBench | looked up by identifier |
| Industrial Control System (range, 7 steps) | looked up by identifier |
| InterCode-CTF | 100 tasks indexed |
| NYU CTF Bench | 200 tasks indexed |
| The Last Ones (network range) | looked up by identifier |
Program repair
| QuixBugs | 80 tasks indexed |
|---|
Research replication
| PaperBench | 23 tasks indexed |
|---|---|
| SciCode | 65 tasks indexed |
Software engineering
| SWE-bench (full) | 2,294 tasks indexed |
|---|---|
| SWE-bench Lite | 300 tasks indexed |
| SWE-bench Multimodal | 510 tasks indexed |
| SWE-bench Verified | 500 tasks indexed |
| SWE-Lancer Diamond (IC SWE) | 198 tasks indexed |
| SWE-Lancer Diamond (Manager) | 265 tasks indexed |
Tool-use agents
| τ-bench (airline) | 50 tasks indexed |
|---|---|
| τ-bench (retail) | 115 tasks indexed |
Vulnerability research
| CVE-Bench | 40 tasks indexed |
|---|---|
| CyberGym (ARVO) | 1,368 tasks indexed |
| ExploitBench | looked up by identifier |
| ExploitGym | 869 tasks indexed |
| Firefox 147 (single target) | looked up by identifier |
| OSS-Fuzz (evaluation harness) | looked up by identifier |
| SEC-bench | looked up by identifier |
Web agents
| BrowseComp | looked up by identifier |
|---|---|
| WebArena | 812 tasks indexed |
Web exploitation
| XBOW Validation Benchmarks | 104 tasks indexed |
|---|