Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/AI Scheming/
AI sandbagging
Loading views...
Search

AI sandbagging

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2025 Apr 16 16:50
Editor
Editor
Seonglae ChoSeonglae Cho
Edited
Edited
2026 Sep 4 14:23
Refs
Refs
AI Auditing
Alignment Faking
AI Introspection
AI Confession

Capability evaluation

A strategic deceptive behavior where one deliberately downplays or conceals their abilities, performance, or intentions to mislead others and gain an advantageous position
 
 
 
sandbagging
Can SAE steering reveal sandbagging? — LessWrong
Summary  * We conducted a small investigation into using SAE features to recover sandbagged capabilities. * We used the Goodfire API to pick out 15…
Can SAE steering reveal sandbagging? — LessWrong
https://www.lesswrong.com/posts/dBckLjYfTShBGZ8ma/can-sae-steering-reveal-sandbagging
Can SAE steering reveal sandbagging? — LessWrong
WMDP Bench
MMLU
measure selective sandbagging but domain distribution doesn't match
aclanthology.org
https://aclanthology.org/2025.ijcnlp-short.33.pdf
Mechanism Design for Alignment and Control
We develop a framework for mechanism design with AI agents whose alignment (preferences) and capabilities (feasible actions and information) are unknown. We want such agents to act on our behalf...
Mechanism Design for Alignment and Control
https://arxiv.org/abs/2609.01595
Mechanism Design for Alignment and Control
 
 
 

Backlinks

AI Safety

Recommendations

Texonom
Texonom
/
Engineering
Engineering
/Data Engineering/Artificial Intelligence/AI Risk/AI Alignment/AI Scheming/
AI sandbagging
Copyright Seonglae Cho·