Aug 2026· 2026 IEEE/CIC International Conference on Communications in China (ICCC)· pp. 473-478· 0 citations· 18 references
Abstract
The rapid deployment of Large Language Model (LLM)-based agents in wireless networks has introduced severe communication bottlenecks due to the continuous exchange of massive prompt sequences. Existing prompt compression methods either cause semantic degradation or introduce prohibitive computational delays, and largely fail to jointly optimize network variables and compression strategies for ultra-low latency multi-agent collaboration. To address these challenges, we propose the Adaptive Prompting Agent-to-Agent Communication (APAC) framework, a dynamic prompt compression mechanism tailored for edge networks. By modeling the joint compute-communicate process as a highly constrained optimization problem, APAC dynamically toggles between lightweight extractive token pruning and high-fidelity generative semantic compression, while continuously adapting the compression ratio based on real-time bandwidth, prompt lengths, and hardware pipelining capabilities. We develop a custom Proximal Policy Optimization (PPO) algorithm tailored for hybrid action spaces to seamlessly balance local computational delay, transmission overhead, and task-conditioned semantic utility. Extensive evaluations demonstrate that APAC achieves a Pareto optimal trade-off between latency and semantic preservation, exhibiting remarkable resilience and strictly circumventing catastrophic latency violations under network congestion and scaling prompt lengths.
Diffusion language models (DLMs) offer a non-autoregressive alternative for mobile edge agentic artificial intelligence (AI) by refining tokens through iterative denoising rather than left-to-right decoding. Compared with autoregressive Transformer-based large language models (LLMs), DLMs can update multiple uncertain...
Chen-Qi Li, Ming-Hui Min, D. Niyato et al.· 0 citations
This study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents and shows decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.
A. Rajalakshmi, D. Saveetha, S. V. Manikanthan et al.· International Journal of Int...· 0 citations
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice in...
A ComBERT-driven service-based RAN UP decoupling method, specifically targeting the functional coupling and redundancy between the PDCP and RLC sublayers is proposed, providing a crucial foundation for on-demand service orchestration in 6G networks tailored to agent services.
Hai-Yu Ding, Shang-Yuan Du, Xin Sun et al.· Italian National Conference...· 0 citations
Analytical findings indicate that decentralized agent specialization, shared contextual state, adaptive workload redistribution, and failure-aware coordination can provide a stronger basis for resilient streaming than static pipelines.
Nethmi Perera, K. Fernando· American Journal Of Applied...· 0 citations
A compact reasoning model trained with verifier-based self-verification and periodically refined online via shadow updates is deployed, showing manageable, near-linear control-plane overhead as domains scale and during domain joins, and robust decision quality, including recovery after objective changes.
Masoud Shokrnezhad, T. Taleb· IEEE Network· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.