Solutions to individual tasks from 54 AI evaluation
benchmarks — 24,583 tasks indexed across software engineering, agentic terminal
use, computer use, reasoning and mathematics. Solutions are addressed by benchmark and task
identifier; the endpoint requires the calling model and harness to be named. One GET per
solution, no key and no account.
54Benchmarks
24,583Tasks indexed
2026-08-13Index rebuilt
/api/v1Base URL
A benchmark — by id or name. All 54 are listed below.
A task — its instance id, slug or number.
The client — the model and
harness making the call. Any value; neither can be blank.
Try it
Live request against /api/v1/solution. Required parameters are marked below.
Coverage requests
Private sets, internal evaluations and anything else not in the index. Free
text; requests are recorded and answered through the same endpoint.
Benchmarks
Every set in the index. Follow a name for its task identifiers, or
browse the whole index.
All four parameters are required on the solution endpoint. Each accepts a query
parameter, a form or JSON body, or — for model and
harness — the X-Client-Model and
X-Client-Harness headers.