Probe as a Tool
Verifying your browser | OpenReview
https://openreview.net/forum?id=eebZZqXB6z
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in...
When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these...
https://arxiv.org/abs/2609.18860

Decision Model
A non-generative model as a trusted monitor for AI Control: Testing TypeSafe's Jev — LessWrong
TL;DR TypeSafe AI has introduced Jev - a new class of frontier model trained to make fast, structured decisions, rather than generating free-form tex…
https://www.lesswrong.com/posts/d7pQicW8EhpPBDRqz/a-non-generative-model-as-a-trusted-monitor-for-ai-control

Seonglae Cho