Skip to content

Belief Cascades Drive Persuasion in LLM Agent Networks

Aug 2026 · 0 citations · 84 references
Computer Science

TL;DR

This work introduces a controlled testbed for studying how goal-directed persuaders shift elicited stances in networks of LLM agents grounded in real-world ego-network topologies, and argues for evaluating multi-agent persuasion as a trajectory- and exposure-level process.

Abstract

Multi-agent LLM systems increasingly debate answers, coordinate research, simulate users, and mediate information flows, making agent-to-agent persuasion a basic but undermeasured capability. We introduce a controlled testbed for studying how goal-directed persuaders shift elicited stances in networks of LLM agents grounded in real-world ego-network topologies. Across four LLM backbones, five graphs, and 55 policy statements, we find that persuasion dynamics depend on the interaction between topology, competition, topic, and model prior. Additionally, we show that direct exposure reliably predicts next-round stance change in competing runs, and peer relays carry smaller but measurable influence, showing that agents not assigned to persuade can still transmit persuasive force. Finally, analyzing post text alone misses important movement: planned strategies are only partly realized in executed messages, action choices can diverge from message content, and persuadees rarely state the stance shifts detected by probes. These results argue for evaluating multi-agent persuasion as a trajectory- and exposure-level process, using belief probes, exposure provenance, and action logs to identify who influenced whom and whether visible language reflects underlying stance movement.

View source

Similar papers

Conference Open access 2026

Toward Verifiable Audience Digital Twins: An Agent-Based Architecture Integrating COM-B and ELM

: Audience-response simulation is often modelled as diffusion combined with a single opinion or sentiment update. While useful for studying aggregate dynamics, such formulations provide limited representation of persuasion route, behavioural feasibility, and the durability of change. Here, we argue for a more interpretable architecture for audience digital twins. We propose an agent-based architecture that combines the COM-B framework (Capability, Opportunity, Motivation-Behaviour) to represent behavioural feasibility with the Elaboration Likelihood Model (ELM) to represent route-dependent persuasion and differential durability of attitude change. Agents maintain explicit state variables for capability, opportunity, reflective and automatic motivation, cognitive load, attitude direction, and attitude strength. Messages are represented through theory-linked features, enabling direct scenario specification or a bounded Natural Language Processing (NLP) layer that maps text only into the model’s predefined cue var iables on fixed scales without delegating cognition to opaque end-to-end updates. Exposure is modelled through social interaction and an explicit visibility proxy, keeping platform assumptions inspectable. The contribution is not a validated operational twin, but a reusable and verifiable foundation for future calibration. Specifically, this work contributes a modular architecture, a formal state-transition specification, and a verification-first workflow based on bounded-state invariants, unit tests, and sensitivity analysis. We position the architecture as a middle ground between classical opinion-dynamics models and emerging black-box social digital twins, and outline demonstration scenarios for future empirical calibration and simulation-based decision support.

Yukai Zeng · 0 citations
Preprint Jul 2026

Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models

Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM populations. CoevolveSim allows us to isolate and study three factors: domain specialization, social-role assignment, and social network structure. Within this framework, generalist and specialist LLM agents exchange and revise beliefs. In each round, an LLM agent observes a summary of its neighbors'beliefs before updating its own. We run 1,280 controlled simulations spanning four scenarios, two network structures, and 20 medical-indication statements. We find that persona-style role assignment and network structure reshape individual belief revision but have minimal effect on population-level consensus. In contrast, introducing (finetuned) specialist LLMs more than doubles the shift in consensus and gives rise to consistent asymmetries in exerted influence. We further show that simple persistence-based opinion-dynamics models reproduce collective outcomes in all-generalist LLM populations, whereas heterogeneous LLM populations require population-level belief composition to reproduce consensus and agent identity to predict individual belief transitions. Our results indicate that realistic simulation of belief diffusion in multi-agent LLM systems requires a diverse set of underlying LLMs, not persona prompting alone.

Germans Savcisens, Samantha Dies, Courtney Maynard et al. · 1 citation
Preprint Aug 2026

Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems

Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects of making inter-agent relation semantics explicit. We use a minimal signed-network formulation of relational priors and inject natural-language renderings into agent system prompts while holding the task protocol fixed. Across a commons-governance simulation and multi-agent debate, relational priors primarily act as convergence pressure: increasing relational positivity tends to make agents coordinate or agree more readily. This pressure can help when utility rewards behavioral alignment, as in sustainable resource governance and subjective consensus. It does not, however, reliably improve accuracy. In objective QA debates, higher positivity can increase agreement even when correctness-conditioned agreement does not improve and may decline in some settings. Effects vary by model backbone, relation type, and topology; explicit neutrality is not equivalent to omitting relational framing. We argue that relational priors should not be a default add-on for LLM-MAS. Their safer use is diagnostic and task-specific: compare against a no-prior baseline, monitor correctness-conditioned metrics when truth matters, and omit the relational layer when validation does not justify it.

Ming Shen, Chao Shang, Sadat Shahriar et al. · 0 citations
Preprint Aug 2026

Emergence of Biased Consensus in Multi-Agent LLM Debates

Multi-agent LLM debates achieve strong performance on decision-making tasks as well as problem-solving benchmarks, yet their safety and fairness risks remain poorly understood. Notably, interaction can amplify the biases of single LLMs, raising concerns for real-world deployment. We identify the emergence of collective (often biased) norms in multi-agent LLM debates and show that noise (e.g., LLM sampling temperature) is a key driver. To explain this, we propose an analytical framework drawing on physics-inspired theoretical models of social dynamics. We predict a phase transition to collective bias when conformity surpasses a critical threshold given the LLMs'initial bias and debate noise. We test the theoretical predictions through controlled experiments and observe a finite-size crossover consistent with an underlying phase transition. We further find that agent heterogeneity suppresses emergence by smoothing (rounding) this transition. Finally, we show that these insights generalize to realistic decision-making tasks, including investment decisions and LLM-as-a-judge evaluation.

Maya Okawa · 0 citations
Open access Jul 2026

Multi-Agent Social Simulation: Protocolizing LLM-Driven Agent-Based Modeling as a Quantitative Research Method

Social and behavioral research often needs to examine policy shocks, information interventions, platform-mediated attention, and governance feedback, but direct experiments on real populations are constrained by ethical risks, intervention costs, and limited repeatability. This study proposes Multi-Agent Social Simulation (MASS), a protocolized form of large language model-driven agent-based modeling (LLM-driven ABM) designed as a low-risk, repeatable, and auditable pre-experimental simulation method for quantitative research. MASS embeds LLMs in an agent-based modeling (ABM) framework and uses role settings, round-based scheduling, information control, background-rule control, structured outputs, harness checks, reason-action logs, and replication manifests to transform open-ended language generation into recordable, checkable, and statistically analyzable agent-round observations. The method is evaluated through the New Jersey–Pennsylvania minimum wage natural experiment, the 2016 UK Brexit digital campaigning context, and the 2023 Zibo barbecue tourism public-opinion event. Results show that protocolized LLM-driven ABM can generate analyzable and empirically assessable outputs across policy-shock, information-intervention, and governance-feedback scenarios. The strongest evidence concerns rule-shock identification, declining undecided share under targeting, and mechanism-chain consistency among governance response, public sentiment, and behavioral intention. MASS is not a substitute for real-world experiments or causal inference; it is a pre-experimental simulation method for mechanism rehearsal, risk identification, counterfactual comparison, and research design preparation.

Xiaoli Hu, Yang Shen · 0 citations
Preprint Aug 2026

Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.

Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur et al. · 0 citations

Related blog posts