AI cyber attack capability is better than human pen-tester
MLSN #19: Honesty, Disempowerment, & Cybersecurity
Also, a new AI safety fellowship for experienced researchers
https://newsletter.mlsafety.org/p/mlsn-19-honesty-disempowerment-and?utm_source=post-email-title&publication_id=415332&post_id=190670410

Backtracking: Reset token as a control tool 2024
arxiv.org
https://arxiv.org/pdf/2410.03893
SRFT: Self-Report Fine-Tuning
Hidden objective execution ability remains intact → Not actually better behaved, just better at confessing. However, honesty is confirmed to generalize strongly through training
arxiv.org
https://arxiv.org/pdf/2511.06626
Seonglae Cho