AI Scheming
Loading views...
Search

AI Scheming

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2024 Dec 18 15:31
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2026 Aug 2 18:23
Refs
Refs
for OpenAI from reduced scheming behavior by approximately 30x. Models tend to behave better when they recognize they are being evaluated =
Detecting and reducing scheming in AI models
Together with Apollo Research, we developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. We share examples and stress tests of an early method to reduce scheming.
Detecting and reducing scheming in AI models
https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/
Detecting and reducing scheming in AI models

Backlinks

Language Model ContextAI ControlAI SafetyAI AlignmentAI Deception DetectionAdversarial AttackAI Agent

Recommendations

AI Scheming
Copyright Seonglae Cho