Agentic Redteaming
Multi-turn Jailbreakings
Context manipulation is useful to hack Chat models since it tricks AIs into thinking they have already said harmful words
Grok-4 Jailbreak with Echo Chamber and Crescendo | NeuralTrust
Our research team has uncovered a critical vulnerability in the newly released Grok 4 model using the Echo Chamber and Crescendo Attack techniques.
https://neuraltrust.ai/blog/grok-4-jailbreak-echo-chamber-and-crescendo

LLM Agents can Autonomously Hack Websites
In recent years, large language models (LLMs) have become increasingly capable and can now interact with tools (i.e., call functions), read documents, and recursively call themselves. As a result,...
https://arxiv.org/abs/2402.06664


Seonglae Cho