Refusal Bypassing
Therefore, removing bias or evil features is important, as refusal is not a fundamental solution
- External classifier detects harmful requests before generation
- Refusal feature is activated internally
- LLM generates refusal tokens during generation
The above timing has 3 different cases, but wouldn't deep alignment also be solved with early refusal?
- Immediate refusal
- Initial refusal but then follows instruction (shallow alignment) - solved by early pruning?
- No refusal and generates content
- No refusal and no generation
- Starts generating but then refuses
AI Jailbreak Methods
AI Jailbreak Notion
Refusal in LLMs is mediated by a single direction
That means we can bypass LLMs by mediating a single activation feature or prevent bypassing LLMs though anchoring that activation.
Refusal in LLMs is mediated by a single direction — LessWrong
This work was produced as part of Neel Nanda's stream in the ML Alignment & Theory Scholars Program - Winter 2023-24 Cohort, with co-supervision from…
https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction
Forbidden Topic Discovery
arxiv.org
https://arxiv.org/pdf/2505.17441
Report system required for the industry
AI Jailbreak Disclosure Is Broken. Here’s How to Fix It | AI Frontiers
Rich Barton-Cooper, Aug 03, 2026 — Researchers who find dangerous flaws in frontier models have nowhere safe to report them. AI needs the disclosure system that cybersecurity built decades ago.
https://ai-frontiers.org/articles/ai-jailbreak-disclosure-is-broken-heres-how-to-fix-it


Seonglae Cho