Preprint
Aug 2026
How Language Models Organize and Structure Moral Knowledge
This work trains six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category, and examines how the resulting directions relate to each other in representation space, finding the directions neither collapse into a single moral detector nor isolate from one another.
Orion Reblitz-Richardson
· 0 citations