AI Jailbreak

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2023 Mar 7 15:1
Editor
Edited
Edited
2026 Aug 3 16:33

Refusal Bypassing

Therefore, removing bias or evil features is important, as refusal is not a fundamental solution
  1. External classifier detects harmful requests before generation
  1. Refusal feature is activated internally
  1. LLM generates refusal tokens during generation
The above timing has 3 different cases, but wouldn't deep alignment also be solved with early refusal?
  • Immediate refusal
  • Initial refusal but then follows instruction (shallow alignment) - solved by early pruning?
  • No refusal and generates content
  • No refusal and no generation
  • Starts generating but then refuses
AI Jailbreak Methods
 
 
 
AI Jailbreak Notion
 
 
 
 

Forbidden Topic Discovery

arxiv.org

Report system required for the industry

AI Jailbreak Disclosure Is Broken. Here’s How to Fix It | AI Frontiers
Rich Barton-Cooper, Aug 03, 2026 — Researchers who find dangerous flaws in frontier models have nowhere safe to report them. AI needs the disclosure system that cybersecurity built decades ago.
AI Jailbreak Disclosure Is Broken. Here’s How to Fix It | AI Frontiers
 
 

Recommendations