AI Toxicity

Creator
Creator
Seonglae ChoSeonglae Cho
Created
Created
2026 Jul 21 9:37
Editor
Edited
Edited
2026 Jul 21 9:43
Refs
Refs
 
 
 
 
 
 
  • Toxicity is not a single axis; sub-directions like violence, hate, and privacy violations exist separately within a shared structure.
  • Toxicity and refusal are not the same direction.
  • The magnitude of the toxicity direction stabilizes quickly early in training; afterward, what changes is mainly the direction itself (rotation / re-alignment).
Harmfulness Directions in OLMo — LessWrong
Introduction This work was conducted as part of the MARS 4.0 program, supervised by Lorenzo Pacchiardi, with Hannes Whittingham and Mikhail Mironov a…
Harmfulness Directions in OLMo — LessWrong
 

Recommendations