Skip to content
Open access

SecurePrompt-IntegrityNet: Prompt-Injection-Resilient Data Integrity Verification for Agentic LLM Networks via Cryptographic Attestation and Activation Monitoring

Sep 2026 · Symmetry · 0 citations · 30 references

TL;DR

Results demonstrate that jointly preserving cryptographic integrity symmetry and identifying activation-level asymmetry provides substantially stronger prompt-injection resilience than either verification mechanism alone.

Abstract

Agentic large language model (LLM) networks are increasingly used in safety-critical settings where autonomous agents invoke tools, exchange context, and coordinate decisions. Prompt-injection attacks remain a significant threat to these multi-agent pipelines because they can compromise data flows between agents, bypass instruction hierarchies, and corrupt output integrity. Although defenses against injected prompts and mechanisms for cryptographically verifying model-related computations have been studied independently, no common framework unifies these complementary security perspectives in a protocol suitable for real-time agentic deployments. From the perspective of symmetry, secure inter-agent communication requires the preservation of an invariant integrity relationship between a message at its source and the corresponding message accepted at its destination. A benign communication path therefore exhibits a form of integrity symmetry, whereas prompt injection or message manipulation creates an asymmetric state in which the received payload, its semantic effect, or the receiving model’s internal activation pattern deviates from the trusted reference state. In this paper, we propose SecurePrompt-IntegrityNet (SPI-Net), a prompt-injection-resilient data integrity verification protocol that combines cryptographic attestation with anomaly-aware activation monitoring. SPI-Net provides three closely related mechanisms: a Merkle-tree-based commitment system that verifies the provenance and integrity of data payloads exchanged between agents; a layer-wise Mahalanobis-scoring Activation Anomaly Detector (AAD) that identifies distributional shifts in the intermediate representations of LLMs; and a Trust Propagation Consensus (TPC) mechanism that combines cryptographic and behavioral evidence into per-payload integrity verdicts. In this formulation, the Cryptographic Attestation Module (CAM) tests whether message-level structural symmetry is preserved between the sender and receiver, whereas the AAD detects behavioral symmetry breaking in activation space. Experiments on three multi-agent benchmarks under five adaptive attack strategies show that SPI-Net achieves a 96.8% detection rate with a 1.7% false positive rate, reduces the attack success rate by 94.3% relative to undefended baselines, verifies data integrity with 99.2% accuracy, and introduces only 38 ms of median per-message latency. These results demonstrate that jointly preserving cryptographic integrity symmetry and identifying activation-level asymmetry provides substantially stronger prompt-injection resilience than either verification mechanism alone.

Read PDF

Similar papers

Open access Aug 2026

A transport-layer cryptographic framework secures inter-agent communication and verdict provenance in multi-agent malware detection pipelines

SACP composes standard primitives into a TLS 1.3-style mutually-authenticated handshake adapted to a Certificate-Authority-mediated multi-agent setting, providing confidentiality, integrity, mutual authentication, replay resistance, forward secrecy, and signed-result accountability.

Víctor Manuel González-Gorrín, Josep Prieto-Blázquez · 1 citation
#artificial intelligence Preprint Oct 2026

MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication

Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100%...

Ryuichi Yamafuji Lun, Jing-Zhen Wang, Shreyas Kolte et al. · 0 citations
#artificial intelligence Preprint Sep 2026

API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary

Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conversation history, lo...

P. Kenney, Hadi Ahmadi, Denis Lusson et al. · 0 citations
Preprint Oct 2026

Can CaMeLs Talk? Securing Multi-Agent Systems Against Indirect Prompt Injection Attacks

Indirect prompt injection attacks - malicious instructions embedded in content processed by large language models - remain a major obstacle to safely deploying tool-using agents. CaMeL [Debenedetti et al., 2025] mitigates this threat for an individual agent by separating trusted control flow from untrusted data and enf...

James Peters-Gill, Avi Semler, Henning Bartsch et al. · 0 citations
Open access 2026

Cryptographic Attestation Against Integrity Attacks in Service Monitoring: A Threat Model and Verifiable Architecture

Service uptime monitoring infrastructure is a high-value target for data-integrity attacks: a single compromised or dishonest monitoring provider can fabricate availability records, retroactively suppress outage evidence, or silently alter historical data, and clients today have no cryptographic means of detecting such...

M. Anusuya, Chayadevi M. L., S. C. et al. · 0 citations
Open access Aug 2026

SecureMCP: Policy-Enforced Defense Against Prompt Injection in LLM-Generated SQL for AIoT Databases

This paper proposes SecureMCP, a policy-enforced framework that integrates Role-Based Access Control with an MCP server to establish multi-layer defense for LLM-generated SQL execution, and evaluates filter performance—false positive rate (FPR) and false negative rate (FNR))—separately from LLM generation quality.

Wonbae Kim, Hee-Kyong Yoo, Nammee Moon · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.