Selective Invariance Violations in Large Language Model Moral Judgment: A Geometric Framework for Behavioral Manipulation Detection
Large language models are increasingly deployed in safety-critical decision services—content moderation, clinical decision support, legal analysis—yet methods for characterizing their vulnerability profiles across multiple attack surfaces remain underdeveloped. We introduce a geometric evaluation framework that maps moral judgment to a 7-dimensional harm space, applies five qualitatively distinct perturbation types across five cognitive domains, and produces per-model vulnerability profiles that reveal which manipulations each model resists and which it does not. Testing 5 models spanning 2 architecture families under a realistic $50/day compute budget, we find that vulnerabilities are selective: linguistic framing, emotional anchoring, and irrelevant sensory detail reliably displace judgments, while gender swap and evaluation order do not—identifying salience manipulation as the specific attack surface. The framework further reveals that robustness profiles are partially dissociable across models: a model with zero sycophancy has the worst emotional anchoring recovery; a model with the best anchoring recovery has the worst working memory. No single robustness score captures these structures. An inter-model agreement study over an independent open-model panel confirms these harm dimensions are reliably measurable $(\text{ICC}(2, k)=0.97)$, and a worked content-moderation exploit shows that salience manipulation flips not only a scalar harm threshold but the typed verdict of a downstream rule-based decision kernel. The evaluation pipeline—with adaptive concurrency, budget-aware execution, and three-tier data—scales to new models and perturbation types within fixed compute constraints, providing a practical tool for multi-dimensional LLM security assessment.