Skip to content
Book Open access

A YANG-Grounded LLM Agent Supporting Multi-Vendor OpenROADM Optical Transport Networks

Aug 2026 · Proceedings of the ACM SIGCOMM 2026 Conference · pp. 2109-2115 · 0 citations · 14 references

TL;DR

This paper presents an LLM agent that extends the Optical Network Robotic Automation Platform with YANG-Grounded Agentic Retrieval-Augmented Generation (RAG), and serves as the first instantiation of a pipeline whose YANG-aware components apply to other open standards built on YANG and NETCONF.

Abstract

The growing role of optical networks in Artificial Intelligence (AI) infrastructure exposes AI researchers to a field still gated by fluency with vendor-specific Yet Another Next Generation (YANG) models and Network Configuration Protocols (NETCONF). Large Language Models (LLMs) can help bridge this gap if applied properly. Specifically, LLMs do not know the YANG data models adopted by commercial Open Reconfigurable Optical Add-Drop Multiplexer (OpenROADM) equipment, and retrieving YANG as flat text loses the hierarchical structure that gives each constraint its meaning. This paper presents an LLM agent that extends the Optical Network Robotic Automation Platform (ON-RAP) with YANG-Grounded Agentic Retrieval-Augmented Generation (RAG). The agent parses 154 OpenROADM YANG files into hierarchically structured chunks, drives retrieval on demand through a Reasoning-and-Acting loop, and routes YANG constraints to the model through three complementary paths spanning prompt-time injection, tool-call retrieval, and post-generation validation. The system is evaluated against retrieval baselines on an open-source Lightynode simulated testbed and was demonstrated live on a commercial multi-vendor testbed at the Optical Fiber Communication Conference 2026. The same pipeline drives both testbeds without modification. OpenROADM serves as the first instantiation of a pipeline whose YANG-aware components apply to other open standards built on YANG and NETCONF.

Read PDF

Similar papers

2026

LLM-Empowered MAPPO for Embodied Cognitive Satellite–Terrestrial Networks With RSMA

This paper introduces an embodied agentic AI framework that integrates large language models (LLMs) with multi-agent reinforcement learning (MADRL) to enable adaptive control in cognitive satellite-terrestrial networks (CSTNs). The framework embeds LLM-based cognitive modules into network entities, transforming them into autonomous agents capable of semantic perception, reasoning, and collaborative decision-making. To address key CSTN challenges such as dynamic interference, complex resource allocation, and heterogeneous quality-of-service (QoS) demands, we employ LLMs to interpret high-level operational intents, augmented by retrieval-augmented generation (RAG) for accessing domain knowledge. This enables each agent to adaptively configure rate-splitting multiple access (RSMA)-based protocols, derive key performance metrics (e.g., outage probability, age of information), and formulate a constrained long-term energy efficiency optimization problem. To solve this problem, we propose an LLM-enhanced multi-agent proximal policy optimization (LEMAPPO) algorithm for joint power and rate allocation. The LLM enhances MAPPO through action guidance and reward function design, thereby improving learning efficiency and policy robustness. Simulations demonstrate that the proposed algorithm achieves substantial energy efficiency gains while satisfying reliability and timeliness constraints, outperforming existing benchmarks. Specifically, it outperforms standard MAPPO by up to 28.5% in energy efficiency under stringent outage constraints, and achieves 27.3% higher efficiency than MAPPO in multi-user scenarios.

Chenbo Hu, Hongjuan Yang, Bo Li et al. · 0 citations
Open access

Multi-agent systems for optical networks

(English) Unlike earlier mobile generations, 6G is expected to support a wide range of applications such as immersive communications, remote healthcare, autonomous transportation, and smart cities. These use cases will significantly increase the number of connected devices and impose stringent requirements on bandwidth, latency, reliability, and energy efficiency. As a result, the networks supporting these services will face major challenges in scalability, resource management, and control. In this context, this doctoral thesis investigates the use of Multi-Agent Systems (MAS) as a foundation for next-generation network control. The goal of this thesis is to design and evaluate MAS-based solutions that improve the intelligence, scalability, and energy efficiency of optical networks across both the optical and packet layers. The first objective addresses the optical layer by investigating centralized and distributed MAS-based approaches for dynamic spectrum control in point-to-multipoint (P2MP) connections. A centralized solution based on traffic prediction and integer linear programming computes optimal allocations under near-real-time constraints, achieving high spectrum utilization but introducing synchronization and scalability limitations. To overcome these issues, distributed architectures are proposed in which transponder agents perform decision-making locally. Three strategies are studied: a mixed-strategy gaming model, a distributed deterministic algorithm, and a multi-agent reinforcement learning (MARL) approach. The MARL solution achieves the best overall performance by anticipating traffic variations and allocating capacity proactively, while distributed methods significantly improve scalability and robustness. Communication efficiency is also studied with the MARL approach allowing for asynchronous operation and reducing inter-agent messaging. Results show that distributed MAS can approach centralized performance while avoiding bottlenecks and single points of failure. The second objective focuses on the packet layer, where an extended MAS architecture enables end-to-end near-real-time control of network services (NS) through autonomous flow operation. Routing decisions are driven by telemetry and optimized using Deep Reinforcement Learning (DRL) to minimize delay and operational cost, while agents monitor performance and coordinate with the software-defined networking (SDN) controller. The architecture supports the full lifecycle of a NS, including deployment, dynamic reconfiguration, and handover scenarios. A model-selection approach based on offline training and real-time telemetry is proposed, together with an active probe-testing mechanism and long short-term memory (LSTM) based traffic prediction trained online by flow agents. Simulations demonstrate that transferring of trained models between agents enables accurate predictions and knowledge generation allowing for fast reconfiguration decisions while maintaining QoS over the NS. This MAS architecture provides the foundation for the third objective where experimental results demonstrate reliable QoS maintenance and effective MAS reconfiguration during operation. In conclusion, this thesis shows that MAS combined with learning-based decision-making, predictive analytics, and distributed control provide a flexible and effective framework for managing future networks. The proposed solutions improve scalability, adaptability, and energy efficiency while maintaining strict performance guarantees, establishing MAS as a key enabler for intelligent and autonomous 6G networks. (Català) A diferència de les generacions mòbils anteriors, s'espera que el 6G admeti una àmplia gamma d'aplicacions com ara comunicacions immersives, atenció mèdica remota, transport autònom i ciutats intel·ligents. Aquests casos d'ús augmentaran significativament el nombre de dispositius connectats i imposaran requisits estrictes sobre l'amplada de banda, la latència, la fiabilitat i l'eficiència energètica. Com a resultat, les xarxes que donen suport a aquests serveis s'enfrontaran a grans reptes en escalabilitat, gestió de recursos i control. En aquest context, aquesta tesi doctoral investiga l'ús de sistemes multiagent (MAS) com a base per al control de xarxa de nova generació. L'objectiu d'aquesta tesi és dissenyar i avaluar solucions basades en MAS que millorin la intel·ligència, l'escalabilitat i l'eficiència energètica de les xarxes òptiques tant a la capa òptica com a la de paquets. El primer objectiu aborda la capa òptica investigant enfocaments centralitzats i distribuïts basats en MAS per al control dinàmic de l'espectre en connexions punt a multipunt (P2MP). Una solució centralitzada basada en la predicció de trànsit i la programació lineal entera calcula assignacions òptimes sota restriccions gairebé en temps real, aconseguint una alta utilització de l'espectre però introduint limitacions de sincronització i escalabilitat. Per superar aquests problemes, es proposen arquitectures distribuïdes en què els agents transponedors prenen decisions localment. S'estudien tres estratègies: un model de joc d'estratègia mixta, un algoritme determinista distribuït i un enfocament d'aprenentatge per reforç multiagent (MARL). La solució MARL aconsegueix el millor rendiment general anticipant les variacions del trànsit i assignant la capacitat de manera proactiva, mentre que els mètodes distribuïts milloren significativament l'escalabilitat i la robustesa. També s'estudia l'eficiència de la comunicació amb l'enfocament MARL que permet el funcionament asíncron i redueix la missatgeria interagent. Els resultats mostren que el MAS distribuït pot aproximar-se al rendiment centralitzat evitant els colls d'ampolla i els punts únics de fallada. El segon objectiu se centra en la capa de paquets, on una arquitectura MAS estesa permet el control de punta a punta en temps gairebé real dels serveis de xarxa (NS) mitjançant el funcionament autònom del flux. Les decisions d'encaminament es controlen mitjançant telemetria i s'optimitzen mitjançant l'aprenentatge per reforç profund (DRL) per minimitzar el retard i el cost operatiu, mentre que els agents supervisen el rendiment i es coordinen amb el controlador de xarxa definida per programari (SDN). L'arquitectura dóna suport al cicle de vida complet d'una xarxa de xarxa (NS), incloent-hi el desplegament, la reconfiguració dinàmica i els escenaris de traspàs. Es proposa un enfocament de selecció de models basat en l'entrenament fora de línia i la telemetria en temps real, juntament amb un mecanisme actiu de proves de sondes i una predicció de trànsit basada en memòria a curt termini (LSTM) entrenada en línia per agents de flux. Les simulacions demostren que la transferència de models entrenats entre agents permet prediccions precises i generació de coneixement que permeten prendre decisions de reconfiguració ràpides mentre es manté la QoS sobre la NS. Aquesta arquitectura MAS proporciona la base per al tercer objectiu, on els resultats experimentals demostren un manteniment fiable de la QoS i una reconfiguració MAS eficaç durant el funcionament. En conclusió, aquesta tesi demostra que el MAS combinat amb la presa de decisions basada en l'aprenentatge, l'anàlisi predictiva i el control distribuït proporciona un marc flexible i eficaç per a la gestió de les xarxes futures. Les solucions proposades milloren l'escalabilitat, l'adaptabilitat i l'eficiència energètica, mantenint alhora garanties de rendiment estrictes, establint el MAS com un factor clau per a les xarxes 6G intel·ligents i autònomes. (Español) A diferencia de las generaciones móviles anteriores, se espera que 6G sea compatible con una amplia gama de aplicaciones, como las comunicaciones inmersivas, la atención médica remota, el transporte autónomo y las ciudades inteligentes. Estos casos de uso aumentarán significativamente el número de dispositivos conectados e impondrán requisitos estrictos de ancho de banda, latencia, fiabilidad y eficiencia energética. Como resultado, las redes que soportan estos servicios se enfrentarán a importantes retos de escalabilidad, gestión de recursos y control. En este contexto, esta tesis doctoral investiga el uso de Sistemas Multiagente (MAS) como base para el control de red de próxima generación. El objetivo de esta tesis es diseñar y evaluar soluciones basadas en MAS que mejoren la inteligencia, la escalabilidad y la eficiencia energética de las redes ópticas en las capas óptica y de paquetes. El primer objetivo aborda la capa óptica mediante la investigación de enfoques centralizados y distribuidos basados en MAS para el control dinámico del espectro en conexiones punto a multipunto (P2MP). Una solución centralizada basada en la predicción de tráfico y la programación lineal entera calcula asignaciones óptimas con restricciones casi en tiempo real, logrando una alta utilización del espectro, pero introduciendo limitaciones de sincronización y escalabilidad. Para superar estos problemas, se proponen arquitecturas distribuidas en las que los agentes transpondedores toman decisiones localmente. Se estudian tres estrategias: un modelo de juego de estrategia mixta, un algoritmo determinista distribuido y un enfoque de aprendizaje por refuerzo multiagente (MARL). La solución MARL logra el mejor rendimiento general al anticipar las variaciones de tráfico y asignar capacidad de forma proactiva, mientras que los métodos distribuidos mejoran significativamente la escalabilidad y la robustez. También se estudia la eficiencia de la comunicación con el enfoque MARL, que permite la operación asíncrona y reduce la mensajería entre agentes. Los resultados muestran que el MAS distribuido puede aproximarse al rendimiento centralizado, evitando cuellos de botella y puntos únicos de fallo. El segundo objetivo se centra en la capa de paquetes, donde una arquitectura MAS extendida permite el control integral y casi en tiempo real de los servicios de red (NS) mediante

Hailey Josephine Shakespear Miles · 0 citations
Preprint Jul 2026

From Intent to Infrastructure: LLM-Driven Agent Compilers for ISAC Networks

Integrated sensing and communications (ISAC) is moving from proof-of-concept demonstrations to system-level deployment in sixth-generation (6G) networks. Because sensing and communication share hardware, spectrum, and waveform resources, ISAC design now involves many tightly coupled choices, including waveform selection, sensing algorithm setup, resource scheduling, and deployment planning. This design space is already too large to manage well through manual tuning or isolated optimizers. This article introduces the \textit{Agent Compiler}, a large language model (LLM)-enabled compilation layer that translates high-level engineering intent into complete and executable ISAC system configurations. The Agent Compiler works in four stages: intent parsing, task decomposition, policy graph synthesis, and infrastructure mapping. It produces a verifiable intermediate representation called the ISAC Policy Graph (IPG). A runtime engine then deploys the compiled configuration and supports closed-loop adaptation at three levels: fast parameter updates, partial recompilation of affected subgraphs, and full workflow recompilation. The core design principle is strict time-scale separation: the LLM handles slow-loop strategic decisions, while proven algorithms retain real-time control in the fast loop. A UAV-assisted disaster rescue example illustrates the full compilation process. We also discuss open issues, including compilation latency, output reliability, constraint verification, and pipeline security, to guide future research.

Lijie Zheng, Xudong Zhong, Baoquan Ren et al. · 0 citations
Open access 2026

An LLM-Agent-Based Framework for Age of Information Optimization in Heterogeneous Multiple Access Networks

With the rapid expansion of the Internet of Things (IoT) and heterogeneous wireless networks, Age of Information (AoI) has emerged as a critical metric for evaluating information freshness in real-time systems. AoI-oriented access optimization in heterogeneous multiple access networks is challenging because legacy access mechanisms, such as TDMA and ALOHA, may coexist over a shared channel, while conventional rule-based and learning-based methods often suffer from limited adaptability, slow convergence, and poor interpretability. In this paper, we propose Reflex-Core, an LLM-agent-based framework for AoI-oriented adaptive access in heterogeneous wireless networks. Reflex-Core adopts an “Observe-Reflect-Decide-Execute” closed-loop mechanism to refine transmission strategies through semantic feedback and historical memory. To provide an analytical foundation for reflection-guided strategy refinement, we derive a drift-plus-penalty design principle and construct a reflection-cycle-level reward target that jointly captures weighted AoI reduction and collision cost. This reward target guides reflection selection, reward model training, and PPO-based post-training. Based on Reflex-Core, we develop the Reflexive Multiple Access (RMA) protocol and a priority-aware RMA variant for differentiated freshness requirements. We further discuss an asynchronous edge-assisted implementation, where LLM-based reflection can be offloaded without blocking slot-level random access. Simulation results show that RMA reduces AoI by up to 14.9% compared with representative baselines and maintains robust performance in dynamic and priority-aware scenarios. Additional scalability and backbone-sensitivity experiments further confirm that Reflex-Core remains effective in a 20-node heterogeneous scenario with varied ALOHA loads and is robust when LongChat-7B-16k is replaced by Qwen2.5-7B-Instruct.

Fang Liu, Erchao Zhu, Jiedan Tan et al. · 0 citations
Open access Jul 2026

JW-ASTClaw: A Generalizable Multi-agent Framework for Autonomous Solar Telescope and Its Implementation within Chinese Meridian Project

We present the first deployment of an end-to-end autonomous control system driven by a large language model (LLM) on an operational solar telescope—the Solar Full-disk Multi-layer Magnetograph, named JW-ASTClaw. This system employs a multi-agent framework adopting a decoupled three-layer architecture (perception–decision–execution) interconnected through the Model Context Protocol, which addresses real-time adaptive scheduling under complex environmental conditions while achieving high portability: the perception and decision logic are reused unchanged across instruments, requiring only telescope-specific command interfaces to be adapted. Three perception agents—data-quality-agent, cloud-analyzer-agent, and flare-detector-agent—encode senior observer expertise, including wind jitter detection via limb-ring standard deviation, projected-circle zonal cloud analysis, and multi-band active region identification, as LLM-callable rules, while a central reasoning engine performs multi-source fusion and conflict resolution. The system supports graceful degradation from cloud LLM to local inference and finally to rule-based fallback, designed for remote field stations with unstable connectivity. Cross-season validation on archival data demonstrates 100% cloud detection with zero false positives across 10 distinct observation dates, with active-region counts and positions closely matching the NOAA Solar Region Summary (SRS) reports (102 versus 100 across 10 separate validation dates). These capabilities significantly improve scientific-intent-driven observation accessibility, enable rapid flare response for space weather monitoring, enhance data usability under adverse conditions, and increase observability during partially cloudy periods. This work represents the first concrete engineering step toward the embodied intelligent solar telescope concept, providing a validated foundation for the transition from automated scheduling to AI-driven autonomous observation.

Liyue Tong, Jiaben Lin, Yuanyong Deng et al. · 1 citation