Control Reinforcement Learning: Interpretable Token-Level Steering...
Sparse autoencoders (SAEs) decompose language model activations into interpretable features, but existing methods reveal only which features activate, not which change model outputs when...
https://arxiv.org/abs/2602.10437

Activation Steering in 2026: A Practitioner’s Field Guide
I’ve been working with steering vectors for months. Here’s what actually works in practice, what fails in ways nobody warned me about, and the honest playbook for getting started.
https://subhadipmitra.com/blog/2026/activation-steering-field-guide/


Seonglae Cho