It is argued that demonstrating internal incoherence is a necessary precursor to AI alignment as well as a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs.
Abstract
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.
The argument further holds that AI does not possess moral agency in the classical sense but functions as an infrastructural precondition for the reconfiguration of normative hierarchies—in an empirical rather than transcendental sense.
Moral disagreement is often invoked to support skeptical conclusions about moral judgment, typically by modeling it on the epistemology of peer disagreement. On the dominant picture, disagreement with a peer supplies higher‐order evidence of error and thereby rationally requires suspension of judgment or a substantial reduction in confidence. We argue that this evidential picture is problematic, especially in the moral domain. First, the conditions for epistemic peerhood are difficult to specify and apply without arbitrariness, and this undermines the assumption that disagreement is a uniform sign of error. Second, even where peerhood conditions are plausible, moral disagreement need not play the specific evidential role presupposed by the skeptical argument. Drawing on resources from social epistemology, we develop an alternative framework on which persistent moral disagreement can function as a structured form of collective inquiry, regulated by argumentative practices and epistemic institutions. On this view, steadfastness is not automatically a sign of epistemic arrogance; under appropriate conditions, it is compatible with epistemic humility, and it can contribute to both epistemic and moral improvement.
Mario Gensollen, Marc Jiménez‐Rolland, Alejandro Mosqueda· Ratio· 0 citations
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.
Political science is undergoing a pronounced methodological shift toward causal identification and design-based inference, increasingly marginalizing qualitative and observational approaches. We argue that this shift rests on the flawed premise that methods can be ranked in the abstract, independently of the research question and the theories at stake. Drawing on a Popperian understanding of scientific progress, we develop a framework of comparative theory testing in which method choice is derived from the competing theories themselves. When rival theories generate observationally non-equivalent implications—causal, correlational, or descriptive—any form of evidence capable of adjudicating between them has epistemic standing. This grounds methodological pluralism not in a normative appeal for inclusivity, but in the internal logic of rigorous theory testing. We illustrate the framework through canonical examples, propose criteria for evaluating the rigor of comparative theory tests, and show that design-based methods earn their place within—not above—this framework. The result is a unified perspective that preserves our capacity to engage big-picture questions while maintaining the standards of falsifiability and theoretical precision that scientific inquiry requires.
H. Bulutgil, Harris Mylonas· International Political Scie...· 1 citation
The paper’s central argumentative shift is to change the narrative from bias mitigation to bias management—treating bias not as a defect to be corrected but as an ongoing condition to be governed.
Gabriela Arriagada-Bruneau· Science and Engineering Ethi...· 0 citations
The literature on feasibility in political philosophy has been dominated by non-moralized accounts, but also moralized variants have been defended. Recently, however, it has been argued that unless reduced to mere possibility, also allegedly non-moralized accounts are in fact moralized, and therefore cannot do the theoretical work assigned to them. In this paper, we reject that conclusion. Whether moralization is understood in the commonsensical sense—according to which an instantiation of the concept itself provides a moral reason—or in the weaker sense of value-ladenness—according to which correct application depends on background evaluative judgments—the concept of feasibility does not collapse into moralization. The paper furthermore argues that feasibility should not be moralized, since doing so obscures the distinct role of feasibility in practical deliberation. A non-moralized concept of feasibility preserves the difference between what is infeasible and what is merely costly, demanding, or morally unattractive, and thereby retains its value as an analytical and practical tool in political philosophy.
Eva Erman, Niklas Möller· Res Publica· 0 citations