Skip to content

HSD-PA: Hierarchical Semantic Decomposition and Predictive Adaptive Scheduling for Distributed Intelligent Agent Systems

Jul 2026 · International journal of pattern recognition and artificial intelligence · 0 citations

TL;DR

HSD-PA introduces a semantic-aware task graph to model dynamically evolving task dependencies, enabling flexible and context-aware decomposition and a predictive scheduling strategy estimates agent workload, communication latency, and execution reliability to support look-ahead decision-making, thereby mitigating bottlenecks and improving coordination efficiency.

Abstract

Distributed intelligent agent systems have become a fundamental paradigm for large-scale autonomous decision-making in dynamic environments characterized by fluctuating workloads, time-varying communication delays, heterogeneous computational resources, and evolving task dependencies and resource-constrained environments. However, existing approaches to task decomposition and scheduling often rely on static structures or reactive strategies, which fail to capture evolving task semantics and lack predictive adaptability. To address these limitations, we propose HSD-PA, a hierarchical semantic decomposition and predictive adaptive scheduling framework for distributed task execution. HSD-PA introduces a semantic-aware task graph to model dynamically evolving task dependencies, enabling flexible and context-aware decomposition. A hierarchical mechanism further refines sub-task representations based on real-time system states. In addition, a predictive scheduling strategy estimates agent workload, communication latency, and execution reliability to support look-ahead decision-making, thereby mitigating bottlenecks and improving coordination efficiency. A distributed consistency mechanism is also incorporated to achieve scalable coordination with partial local information. Experimental results on multiple datasets demonstrate that HSD-PA consistently improves task completion time, system throughput, and robustness compared with state-of-the-art methods, particularly under workload-dynamic, communication-uncertain, and heterogeneous multi-agent environments.

View source

Similar papers

Open access Aug 2026

Adaptive Multi-Agent AI Framework for Real-Time Data Streaming with Enhanced Scalability and Resilience

Real-time data streaming systems increasingly operate under highly variable workloads, heterogeneous data sources, latency constraints, and frequent service disruptions. Conventional stream-processing architectures generally depend on predefined routing, static resource allocation, and centralized coordination, which can limit their ability to adapt when event rates, computational requirements, or infrastructure conditions change rapidly. This paper proposes an Adaptive Multi-Agent AI Framework for Real-Time Data Streaming with Enhanced Scalability and Resilience, in which autonomous AI agents collaboratively perform stream monitoring, workload classification, task allocation, resource adaptation, anomaly detection, and recovery. The theoretical foundation combines multi-agent coordination with contextual representation, long-document processing, memory management, and adaptive decision-making. Prior work on aspect-controllable summarization demonstrates the value of controlling computational objectives according to task requirements, while studies of coreference, lexical chains, and entity-based coherence emphasize the importance of preserving relationships across distributed information units (Amplayo, Angelidis, & Lapata, 2021; Baldwin & Morton, 1998; Barzilay & Elhadad, 1997; Barzilay & Lapata, 2005). Long-context language modeling further motivates mechanisms capable of retaining relevant information over extended streaming windows (Beltagy, Peters, & Cohan, 2020). The proposed framework extends these principles to adaptive streaming environments and aligns with recent multi-agent event-streaming research emphasizing resiliency and scalability (Reddy et al., 2026). Analytical findings indicate that decentralized agent specialization, shared contextual state, adaptive workload redistribution, and failure-aware coordination can provide a stronger basis for resilient streaming than static pipelines. The paper also identifies trade-offs involving coordination overhead, state consistency, model complexity, and resource consumption.

Nethmi Perera, Kasun Fernando · 0 citations
Conference Jul 2026

Autonomous Multi-Step Workflow Orchestration using an Agentic AI Framework in Cloud-Edge Enterprises

Cloud-edge computing environments are evolving rapidly, requiring orchestration mechanisms that may automatically construct and manage complex multi-step workflows with little human intervention. We introduce a framework for the agentic AI and how it should be able to orchestrate an autonomous end-to-end workload of cloud-edge enterprise infrastructures in general. The proposed framework relies on large language model (LLM)-driven agents capable of dynamic task decomposition, real-time decision-making, and self-correcting execution pipelines to manage heterogeneous workloads. Through the incorporation of multi-agent coordination protocols, context-aware scheduling algorithms, and feedback-driven optimization loops, the system facilitates seamless task delegation throughout edge nodes and cloud backend systems while managing latency, resource allocation, and compliance constraints. Experimental evaluations show up to percentage improvements in workflow completion rates, resource utilization, and fault tolerance over traditional static-command Rule-based orchestration approaches. Additionally, the framework features explainability modules and audit trails to promote transparency and accountability in autonomous operations. The results provide evidence that agentic AI architectures can serve as a scalable, resilient and intelligent control mechanism for next generation enterprise workflow management across hybrid cloud-edge settings. This has laid a foundation and is to our best of knowledge, the first systematic pioneers work that lays down a roadmap for production-grade autonomous orchestration deployed in analytics and enterprise domains.

Shiza Arshad, Anusha Joodala, A. Agade et al. · 0 citations
Preprint Aug 2026

Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems

A semantic-uncertainty-guided orchestration approach, HASSUM is introduced as a general framework for uncertainty-aware coordination in multi-agent systems and suggests that semantic uncertainty is a practical and general-purpose signal for improving robustness and trustworthiness in agentic AI systems.

John Knowlton, Aritra Guha, Risto Miikkulainen · 0 citations
Review 2026

The Systems Architecture of LLM Multi-Agent Systems: Routing, Memory, and Resource Optimisation

This survey presents a systematic taxonomy and technical review of dynamic orchestration strategies designed to address communication overhead, KV cache management challenges, and increased token consumption within large Language Model-based Multi-Agent Systems.

Heet Nagoriya, H. Raithatha · 0 citations
Conference Jul 2026

Topology-Aware Multi-Agent Reinforcement Learning for Efficient Resource Allocation in Cloud-Native Stream Processing

High-velocity workloads and intricate task dependencies inherent in distributed stream-processing systems pose a fundamental challenge to efficient resource allocation. Traditional heuristic and single-agent reinforcement learning (RL) schedulers frequently fail to recognize these complex network and data-flow interactions, leading to severe resource fragmentation and catastrophic tail latency spikes. In order to accomplish coordinated, low-latency scheduling, we propose a Topology-Aware Multi-Agent Reinforcement Learning (TAMARL) framework utilizing a Centralized Training and Decentralized Execution (CTDE) architecture. TAMARL allows distributed agents to optimize task placement across heterogeneous cluster nodes and prevent backpressure cascades by integrating topology-aware state representations. We evaluate TAMARL on a production-grade cloud-native stack leveraging Apache Flink and Kubernetes. Compared to state-of-the-art baselines across six demanding stress-test scenarios, experimental evaluations demonstrate that TAMARL improves Service Level Objective (SLO) attainment by 27% while reducing P99 tail latency by up to 68%. Additionally, TAMARL maintains stable, resilient performance under 90% cluster utilization while securing 95% network locality.

Sunday J. Awine, Jinwei Liu · 0 citations
Aug 2026

An Adaptive AI Multi-Agent Model for Optimizing Real-Time Data Streaming and System Resilience

Real-time data-streaming environments increasingly require adaptive decision mechanisms capable of coordinating heterogeneous computational resources while maintaining throughput, resilience, and service continuity under dynamic workloads. This paper proposes a conceptual adaptive artificial intelligence (AI) multi-agent model for optimizing real-time data streaming and system resilience. The proposed model combines autonomous agents for stream monitoring, workload allocation, resource coordination, anomaly response, and resilience management within a decentralized decision architecture. Its theoretical foundation is informed by research on distributed allocation, fairness, efficiency, optimization, and computational complexity in multi-agent decision environments. In particular, studies of fair and efficient allocation provide useful principles for balancing competing resource demands, while work on Nash social welfare and allocation algorithms demonstrates the value of optimization objectives that consider collective system utility. The proposed architecture extends these principles from indivisible-resource allocation toward dynamic streaming-resource management. The methodology defines agent roles, state representation, utility functions, adaptive allocation policies, coordination mechanisms, resilience procedures, and evaluation criteria. Analytical findings indicate that adaptive multi-agent coordination can improve resource utilization, reduce the impact of localized failures, and support scalable stream processing when compared conceptually with rigid centralized allocation. The model is particularly relevant to event-streaming environments in which workload intensity, resource availability, and service conditions change continuously. The paper further identifies limitations concerning coordination overhead, convergence, observability, and the absence of empirical benchmarking in the present conceptual study. The framework therefore provides a research foundation for implementing resilient AI-driven streaming systems and for future experimental validation.

Chinedu Okafor, Amara R. Eze · 0 citations