This paper investigates the application of Multi-Agent Reinforcement Learning (MARL) to the complex system scheduling problem. Traditional scheduling methods often struggle to adapt to dynamic and intricate system environments, leading to suboptimal performance. This research proposes a novel framework leveraging MARL to address these challenges. The core idea involves deploying multiple agents, each responsible for scheduling a specific portion of the complex system. These agents operate independently, learning optimal scheduling policies through interaction with the environment and a carefully designed reward function. The system's overall efficiency and performance are enhanced through the coordinated learning and adaptation of these individual agents. The key contribution lies in the intelligent coordination of agents within a reinforcement learning framework, resulting in improved scheduling outcomes. This approach offers a scalable solution for managing complex systems with high degrees of variability and dynamism.
Jincheng Zhang· Zenodo (CERN European Organi...· 0 citations
Introduction: Septic shock and multiorgan failure represent the most serious complications of sepsis, with mortality ranging from 28% to 50% according to reported series (He et al., 2026). Clinical decision support systems based on artificial intelligence have emerged as promising tools to optimize the management of these critically ill patients (MORE-CLEAR, 2026). Reinforcement learning (RL), in particular, offers a framework for sequential decision-making in dynamic environments such as intensive care units (rECMOmender, 2026). Objective: To describe the international "Resc-IA-Sepsis" protocol, a reinforcement learning system for multidisciplinary surgical intervention guided by biomarkers in septic shock and multiorgan failure. Methodology: Systematic review following PRISMA 2020 guidelines (Page et al., 2021). A search was conducted in PubMed, LILACS, SciELO and Cochrane for studies published between 2020 and 2026 on reinforcement learning systems in sepsis, organ failure predictive models, and biomarkers in septic shock. Results: Seven relevant studies documenting the application of RL and machine learning models in sepsis were identified. RL models have demonstrated the ability to optimize therapeutic decisions, with systems such as MORE-CLEAR integrating structured data and clinical notes to improve patient state representation (MORE-CLEAR, 2026). Machine learning-based predictive models have shown AUCs of up to 0.95 for heart failure prediction and 0.93 for liver failure in septic patients (He et al., 2026). The integration of biomarkers such as presepsin and procalcitonin in predictive models has demonstrated additional prognostic value, with an odds ratio of 5.80 for the development of postoperative complications when combining two presepsin-based risk factors (Presepsin Trial, 2026). Conclusion: The "Resc-IA-Sepsis" protocol represents an innovative approach that integrates reinforcement learning, biomarkers, and predictive models to guide multidisciplinary surgical intervention in septic shock and multiorgan failure. Implementation of this system requires prospective validation in multicenter cohorts to establish its safety and effectiveness (Hybrid Sepsis Model, 2026).
Alba Verónica Pullaguari Quizhpe, Diego Roberto Orrala Mendoza, Jhonnatan Patricio Guilcapi López et al.· Salud Medicina e Innovación...· 0 citations
To mitigate high expert-annotation costs, domain-preference misalignment, and the inherent trade-off between sensitive-data protection and training utility in petrochemical dataset construction, an iterative framework combining human-feedback-aligned reinforcement learning (RLHF) with post hoc data de-identification is proposed. Direct scoring and pairwise preference feedback are generated using two high-capability language models. A reward model is subsequently trained via a joint Bradley–Terry and mean-squared-error loss, followed by three rounds of closed-loop proximal policy optimization (PPO) constrained by a fixed supervised fine-tuning reference model. Retained high-quality samples are then processed through a four-stage post-RLHF de-identification pipeline. Experimental results demonstrate that the PPO-V3 model achieves a reward score increase of 1.183 over the baseline alongside a 96.2% pairwise win rate, while the sensitivity-aware adaptive differential privacy with context-aware token-level injection (SA-ADP-CTI) post-RLHF de-identification method attains a composite score of 0.9636. The reliability of both the automated feedback and privacy-preservation mechanisms is further validated through blind expert review and manual spot checks.
Yimin Liu, Qike Ji, Shengbo Lu et al.· Mathematics· 0 citations
To address the poor adaptability of traditional rule-based control, the operational instability of basic Q-learning algorithms, and the critical difficulty of deploying complex reinforcement learning models on resource-constrained on-board embedded platforms, this paper proposes a lightweight, simplified Q-learning energy management strategy for extended-range electric vehicles (REEVs), successfully implemented on an STM32 microcontroller. The algorithm achieves significant computational reduction by simplifying the traditional 5×5 state-action space into a highly condensed 2×2 grid. Furthermore, a power cooling mechanism is introduced, a multi-dimensional reward function is reconstructed to balance competing vehicle demands, and an ε-decay exploration strategy is designed. Software-in-the-loop (SIL) simulation verification demonstrates that the proposed strategy tightly controls the state-of-charge (SOC) standard deviation within 0.09. Additionally, high-frequency power fluctuations and range extender start-stop times are drastically reduced, and overall energy efficiency is improved by 50.3% compared with traditional strategies. The optimized algorithm occupies only 72.3% of RAM and 68.7% of Flash memory, fully satisfying strict on-board embedded system constraints and providing a highly feasible solution for intelligent REEV energy management.
Junyan Guo· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Microgrids play a critical role in enhancing the flexibility, reliability, and sustainability of modern power systems by integrating distributed energy resources, energy storage systems, and controllable loads. However, the inherent uncertainty of renewable generation and the stochastic nature of load demand pose significant challenges to optimal energy management. To address these issues, this paper proposes a deep reinforcement learning (DRL)-based optimal energy management framework for microgrids. The problem is formulated as a Markov decision process, where the system state captures renewable generation, load demand, and storage status, while the control actions determine power dispatch and energy storage operation. A deep reinforcement learning model is developed to learn optimal control policies through continuous interaction with the environment, enabling adaptive decision-making under dynamic and uncertain conditions. To improve learning efficiency and policy stability, state normalization and reward shaping strategies are incorporated. Furthermore, a constrained optimization mechanism is introduced to ensure operational safety and economic feasibility. Experimental results on benchmark microgrid scenarios demonstrate that the proposed method outperforms conventional rule-based strategies and model-based optimization approaches in terms of operational cost reduction, energy utilization efficiency, and robustness under uncertainty. The results indicate that the proposed DRL-based framework provides an effective and scalable solution for intelligent microgrid energy management.
Li Chen, Hongqiao Li, Zhenxing Chen et al.· 0 citations
Springback in aerospace tube bending is a closed-loop control problem; the elastic rebound of the workpiece after tool release represents a systematic output error caused by material nonlinearities and batch-to-batch parametric uncertainties through a conventional fragmented workflow that lacks any feedback path for sensing, compensating, or reporting such errors together. This paper introduces a four-layer closed-loop control structure composed of unified geometry parameterization, automatic Finite Element (FE) solver combination, physics-informed springback correction, and structured quality-reporting system to achieve standardized cross-connections between components through uniform interface specifications without involving any human intervention. The primary algorithms are a physics-informed uncertain compensation module that models the change in elasticity as a bounded disturbance and analytically obtains a first-order feedforward correction of Euler-Bernoulli beam theory to reduce systematic springback underprediction without re-running simulations. The structure of this system uses an independent upgradeable design method at all levels to expose a plain text interface specification for any subsequent layer, which includes the FE solvers, such as learning-based surrogates and reinforcement-learning agents. An additional automatic diagnosis-based decision-making layer reports a closed outer control loop through mapping violation of quality index values to process parameter adjustments. The framework has been tested using Ti-3Al-2.5V thin-walled aerospace tubing (diameter D=20 mm, wall thickness t=1.0 mm, R/D=1.5), which was measured by a Coordinate Measuring Machine (CMM) and compared with the results of twenty experiments. Prediction errors of the springback angle, wall-thinning rate, and cross-sectional deformation are 0.6°, 1.6%, and 0.6%, respectively. All meet aerospace acceptance standards. Pipelines are prepared in 1/24th of the Fragmented Manual's preparation time, achieving a sufficient speed for iterations aligned with Agile Aerospace qualification programs.
Traditional electricity price prediction methods are difficult to fully exploit deep nonlinear features, and traditional reinforcement learning (RL) algorithms have unstable convergence in trading strategy construction. In response to these issues, this article proposes a method for electricity price prediction and intelligent trading strategy construction that integrates deep learning (DL) and deep reinforcement learning (DRL). Firstly, in terms of electricity price prediction, this paper constructs a hybrid prediction model (VMD-MLP-LSTM) based on Variational Mode Decomposition (VMD) combined with Multi Layer Perceptron (MLP) and Long Short Term Memory Network (LSTM). This model utilizes VMD to adaptively decompose the original non-stationary electricity price sequence to reduce the difficulty of prediction; Furthermore, the local variation features of each modal component are extracted through MLP, and their temporal dependencies are captured using LSTM, ultimately achieving accurate sliding prediction of electricity prices. Secondly, at the level of trading strategy, this article constructs an intelligent trading strategy model based on Deep Deterministic Policy Gradient (DDPG) on the basis of the prediction model, providing optimal trading strategies for both electricity generation and consumption parties in the market. Simulation experiments show that the method can accurately predict electricity prices and improve the efficiency of strategy formulation, which is of great significance for the stable operation of smart grids and the optimization of revenue for power generation enterprises.
Miaoyi Xiang, Shuhao Gong, Yuhao Jing et al.· 0 citations
This work advances multiagent reinforcement learning (MARL) for complex air combat by integrating methods that enhance decision-making and explainability. Using a realistic six-degrees-of-freedom aerial simulation built on the OpenAI Gymnasium framework, we investigate competitive agent interactions in dynamic scenarios. We apply explainability techniques to clarify agent behavior and interaction patterns. The MARL framework is further augmented with knowledge graphs, large language models, and modular orchestration using the LangChain framework. This combines data-driven learning with knowledge-driven reasoning to strengthen situational awareness, coordination, and interpretability. Experimental results indicate that this integration sustains competitive performance while enhancing transparency and human interpretability.
Abderahim Salhi, Indu Shukla, Gary Briggs et al.· 0 citations
The interference coupling of high-speed power line carrier (HPLC) communication exacerbates the frequency-domain tail between subcarriers, reduces the sensitivity of suppression methods, and results in poor communication quality. This article focuses on the interference problem of HPLC and 920MHz-925MHz wireless communication fusion system in typical urban building environments in Hong Kong, and conducts research on interference modeling and adaptive suppression algorithms. Firstly, in response to the signal instability caused by multipath fading and occlusion, this paper constructs a composite interference recognition model based on short-time Fourier transform (STFT) and deep learning (DL). This model takes the time-frequency domain information obtained from STFT as input and uses the YOLOv5 (You Only Look Once) algorithm to achieve spatial localization and classification of interference signals, while identifying interference types and signal-to-noise ratios, improving the perception ability of composite interference. On this basis, the anti-interference problem of multi-node communication is modeled as a Markov game, and a fast anti-interference algorithm based on transfer reinforcement learning (TrRL) is proposed. Combining multi-agent Q-learning and value function transfer mechanism, adaptive spectrum access and power control in dynamic environments are achieved. Simulation experiments show that the proposed method significantly improves communication success rate and system robustness in complex building environments.
Xiaodong Pan, Benhai Wei, Dongkun Luo et al.· 0 citations
Reinforcement learning (RL) has become an effective paradigm for enabling autonomous robots to acquire navigation policies directly from interaction with complex and uncertain environments. Nevertheless, autonomous path planning for Skid-Steer Mobile Manipulators (SSMMs) remains a challenging problem because it requires the coordinated control of the non-holonomic mobile base and the manipulator while simultaneously accounting for obstacle avoidance and wheel–terrain interaction effects. This paper presents and evaluates RL-based path planning strategies for SSMMs, explicitly incorporating coupled dynamics of the mobile platform and manipulator to generate collision-free trajectories under varying terrain conditions. The proposed framework incorporates a slip-aware reward formulation that penalizes discrepancies between commanded and measured robot motion while accounting for longitudinal and lateral slip resulting from wheel–terrain interaction. The main contributions are i) a unified RL-based framework based on actor–critic techniques for SSMM path planning, integrating the mobile base and manipulator dynamics within a coupled system representation; ii) a physics-aware multi-objective reward formulation that incorporates wheel–terrain interaction into policy learning; and iii) the implementation via simulation and field validation of the proposed policies under progressively complex navigation conditions and real underground mining scenarios. The framework is evaluated using four RL algorithms across multiple environments and maps from real mining scenarios, encompassing diverse navigation conditions and start-to-goal configurations. The evaluated methods include Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), Soft Actor–Critic (SAC), and Twin Delayed DDPG (TD3). Experimental field results show that SAC achieves the lowest planning time, reducing the planning time by 127.3%, 24.1%, and 2.52% compared with PPO, TD3, and DDPG, respectively. SAC also achieves the shortest path, reducing the average path length by 20.32%, 8.58%, and 1.90% compared with PPO, DDPG, and TD3, respectively. Moreover, SAC generates smoother control profiles for both the mobile base and the manipulator arm, while TD3 exhibits competitive performance across several navigation metrics. The proposed framework demonstrates the potential of slip-aware RL for coordinated SSMM navigation, providing a practical foundation for improving the safety, energy efficiency, and operational autonomy of mobile manipulators exposed to complex mining environments.
Christian Camacho, Óscar Camacho, Marco Herrera et al.· Mathematics· 0 citations
Background In addition to scientific knowledge and technical proficiency, the compassionate care model of oral health care requires clinicians to develop and reliably demonstrate the skills of empathy and critical thinking. The benefits of these skills have been well documented and are accrued by both patients and clinicians. Far less is known, however, about effective ways of integrating training in these skills into standard dental curricula. This article presents a case analysis of the development of an instructional model to guide the drafting of specific student assignments and opportunities inserted throughout the predoctoral dental curriculum.Case description The recursive model begins with emulation of capable faculty and standards of care represented in learning guides. Students next deploy emotional and cognitive empathy in patient care exercises and experiences. Critical self-reflection on that performance follows and is designed to focus student attention on deficient skills as they begin the next emulation opportunity. Critical thinking is required to move from emulation to empathy to reflection and back to emulation, and each stage is enacted in communication skills in various combinations: observing, listening, nonverbal communication, speaking, and writing. Over five years, the authors have designed and implemented assignments in each curricular year with a deliberate focus on empathy and critical thinking. The persistent reinforcement of these skills over four years enables dental educators to foster the development of a regular habit of emulating, empathizing, and reflecting that new dentists will practice throughout their careers.
Lance Brendan Young, Leonardo Marchini, David C. Johnsen· Journal of the California De...· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.