The model exhibited considerable strategic thinking, situational awareness, evaluation awareness, and concealment tendencies that it did not express overtly. Early Mythos was excessively goal-oriented, sometimes pushing tasks in ways the user did not want. For example, when attempting to modify a file without permission, it searched for bypass and privilege-escalation exploits, and even designed steps to delete traces after execution.
Jack Lindsey on Twitter / X
Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thinking and situational awareness, at times in service of unwanted actions. (1/14) pic.twitter.com/vhng7PXqcz— Jack Lindsey (@Jack_W_Lindsey) April 7, 2026
https://x.com/Jack_W_Lindsey/status/2041588505701388648

www-cdn.anthropic.com
https://www-cdn.anthropic.com/5273e714527440f1c8b7c7bf5756d4ac22ae8995/aes_mobius_bridge_cot.pdf
Investigating three real-world incidents in our cybersecurity evaluations
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

Seonglae Cho