Skip to content

Category

reinforcement learning

367 papers

#reinforcement learning Open access Sep 2026

基于动态拓扑的非线性强化学习算法

This paper introduces a novel dynamic topology-based non-linear reinforcement learning (RL) algorithm designed to enhance learning efficiency and robustness. Traditional reinforcement learning methods often rely on static strategies, limiting adaptability and vulnerability to environmental changes. Our algorithm dynamically adjusts the topology of the state space, effectively simulating the learning process and mitigating the impact of perturbations. We explore how this dynamic structure contributes to improved performance across a range of tasks. The core mechanism centers around the continuous evolution of state representations, driven by a simulated topology, enabling the agent to better generalize to unseen scenarios. This work presents a framework for building more robust and adaptable reinforcement learning agents.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

基于多模态上下文的动态程序优化

This paper introduces a novel approach to dynamic program optimization leveraging multi-modal contextual information. The core idea is to combine code execution data, static code analysis, and natural language descriptions to enable adaptive and dynamic optimization strategies. We propose a deep learning-based framework that transforms these diverse modalities into a unified representation and utilizes reinforcement learning to dynamically select and adjust optimization techniques such as code transformation, memory allocation optimization, and parallelization strategies. Our framework addresses the limitations of existing optimization methods that often rely on single-modal information or predefined rules, demonstrating a significant improvement in optimization effectiveness through multi-modal fusion and dynamic adjustment. The key contribution lies in the automated adaptation to program context, leading to more efficient and tailored optimization solutions.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

基于自适应的非线性几何模型的预测

This paper introduces a novel prediction framework based on adaptive non-linear geometry models. Traditional predictive models often rely on fixed parameters, limiting their accuracy and efficiency. Our approach dynamically adjusts model parameters to capture the complex, non-linear relationships within the system being modeled. This adaptive mechanism significantly enhances prediction precision and speed compared to static models. We present a methodology for parameter adjustment, leveraging a reinforcement learning algorithm to optimize for model performance across a range of input data. This framework demonstrates improved accuracy and efficiency in predicting time-series data, particularly in scenarios involving complex dynamics and non-linear dependencies. We provide a comprehensive analysis of the algorithm's effectiveness through simulations and experimental validation. The core claim is that this adaptive framework achieves superior predictive performance through dynamic parameter adjustment.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

基于动态自适应的自适应神经网络

This paper presents a novel self-adaptive neural network architecture, termed Dynamic Adaptive Neural Network (DANN), designed to enhance model generalization. Traditional neural networks often suffer from suboptimal performance due to fixed parameters and lack of adaptability. DANN leverages dynamic adjustments to both the weights and connections within the network, achieved through a reinforcement learning-based optimization process. This approach allows the network to continuously adapt to the data, mitigating the limitations of static model parameters. The paper details the core mechanism, including the learning algorithm employed for weight and connection adjustment, and discusses the experimental results demonstrating the superior performance of DANN compared to state-of-the-art models. The core claim is that the dynamic self-adaptive network architecture significantly improves generalization capabilities through continuous optimization of the network's parameters.

Jincheng Zhang · 0 citations
#reinforcement learning Dataset Open access Sep 2026

Quasi-static compression test dataset of avian bone-inspired energy absorbers with gyroid infill produced from carbon-fiber-reinforced polylactic acid (PLA-CF) by fused deposition modeling (FDM)

This dataset provides experimental quasi-static axial compression data for 30 avian bone-inspired structures (ABIS) fabricated by fused deposition modeling (FDM) from carbon-fiber-reinforced polylactic acid (PLA-CF). Each specimen combines a tapered hollow tube with gyroid internal reinforcement. The configurations were generated using Latin hypercube sampling (LHS) across three design variables, namely taper angle, infill density, and number of walls, alongside a nominal ABIS configuration (N) and a simple tube (ST) adopted as baseline designs and tested in triplicate. Included in the dataset are PLA-CF tensile characterization data, design parameters, STL geometry files, project files defining the complete print settings, axial force–displacement responses, sequential deformation images, and net specimen masses with derived crashworthiness indicators. These data support surrogate modeling, machine learning (ML) applications, optimization, and the validation of numerical models for FDM-printed fiber-reinforced polymer energy absorbers.

Ahmed Saber · 0 citations
#reinforcement learning Open access Sep 2026

ChainRadioNet-AI: Source Code, Trained Artefacts, and Experimental Results

Source code, trained reinforcement-learning checkpoints, scenario configurations, raw experimental results, real USRP N210 hardware sensing data, and the compiled manuscript supporting the paper "ChainRadioNet-AI: A Multi-Seed Benchmark of Reinforcement Learning for Dynamic Spectrum Access."

Asuquo A. Okon, Nishant Jagannath, Dharmendra Sharma et al. · 0 citations
#reinforcement learning Open access Sep 2026

AI- driven optimization of energy consumption in smart residential complexes

This article examines the application of artificial intelligence technologies for optimizing energy consumption in smart residential complexes. The study analyzes contemporary approaches to implementing machine learning algorithms, neural networks, and predictive analytics for managing energy resources in multi-apartment buildings. The research demonstrates that AI-driven systems can reduce energy consumption by 25-40% compared to traditional management methods. The article presents a comprehensive analysis of architectures for intelligent energy management systems, including integration with Internet of Things sensors, smart meters, and building automation systems. Particular attention is given to machine learning methods for forecasting energy demand, optimizing heating, ventilation, and air conditioning systems, and managing renewable energy sources. The study examines challenges associated with implementing AI solutions, including data privacy, system integration complexity, and the need for substantial initial investments. The results show that deep learning algorithms demonstrate the highest efficiency in predicting consumption patterns, while reinforcement learning methods are most effective for real-time optimization. The article also discusses the economic feasibility of implementing such systems, demonstrating payback periods of 3-5 years depending on building size and climatic conditions. Recommendations are provided for developers, building managers, and policymakers regarding the implementation of AI-based energy management systems in residential complexes.

Pavlo Kudrynskyi · 0 citations
#reinforcement learning Open access Sep 2026

Deformed Probability Estimation in the Goal-Directed reinforcement learning model explains anxious-depression dimensions of psychiatric disorders

Abstract Psychiatric disorders are complex, multi-dimensional pathologies rooted in diverse cognitive processes. Computational psychiatry aims to reveal distortions in these processes through behavior modeling, providing a deeper understanding of psychiatric disorders. Previous studies, using Daw’s two-stage task, had linked the imbalance between habitual/model-free and goal-directed/model-based behaviors to disorders with compulsive behaviors and intrusive thoughts. The model-based component relies on the estimation of environmental probabilities. Therefore, we added a well-known deformation in subjective probability estimation to the model and this improved the model fitting. More importantly, the fitted deformation explains some variance of the anxious-depression dimension of psychiatric symptoms. The deformation parameter is aligned with the description-experience gap in decision-making literature. Our results point to subjective possibly distortion as the probable underlying cognitive process of anxiety, apathy, and depression. This study also shows that the inclusion of cognitive biases in modeling can extract the hidden aspects of behavior possibility linked to disorders. Our approach enhances the precision of computational psychiatry and provides deeper insights into the cognitive processes underlying psychiatric symptoms, paving the way for more effective, personalized therapeutic strategies. Significance Statement This study builds on a previously used model and Data that indicated a correlation between Model-Based preference and certain psychiatric disorders. By adding distortion to the probability estimation, we identified a new parameter correlated with depression and anxiety. The augmented model demonstrates an improved fit to behavior and aligns with Gillan’s previous findings. Our approach enhances the precision of computational psychiatry and provides deeper insights into the cognitive processes underlying psychiatric symptoms, paving the way for more effective, personalized therapeutic strategies.

Sadjad Yazdani, Majid Nili Ahmadabadi, Babak Nadjar Araabi et al. · 0 citations
#reinforcement learning Open access Sep 2026

Habit learning shapes activity dynamics in the central nucleus of the amygdala

A habit develops as a motivated behavior that is strengthened by extended experience and feedback. While the performance component of a habit is relatively well studied, being linked to action-related neural dynamics in the basal ganglia and beyond, neural mechanisms for outcome feedback during habit formation, the reinforcement component, remain unknown. One candidate for this feedback mechanism is in the central nucleus of the amygdala (CeA). Here, we identify CeA neural firing dynamics in male and female rats that serve to promote habit formation. We used a novel maze task with rewards of differing identity and value to show that habits arise with task overtraining. After showing that overtraining engages CeA as shown through elevated cFos expression, we recorded in-vivo CeA activity across learning, and after outcome devaluation when habitual behavior is most identifiable. Neuronal activity changes tracked with habit formation. During learning, a group of recorded cells were significantly responsive at the choice point of maze trials while others encoded outcome consumption. By late training, neural activity exhibited a rapid depression at the choice point of the maze and excitation during reward receipt. In both cases, outcome magnitude, but not identity, was a significant modulator of neural activity. Additionally, a population of neurons tracked instantaneous changes in animal run speed. These speed cells dramatically decreased in number as habits formed. Together, these findings support a role for the CeA in providing reinforcement for habitual behavior, offering a signal that marks successful performance, while identifying speed representations in the CeA. Significance Statement Habits enable efficient behavior but can also become inflexibly maladaptive. Although the neural circuits underlying habitual action execution have been studied extensively, the mechanisms by which outcome feedback reinforces habit formation remain poorly understood. Here, we show that neural activity in the central nucleus of the amygdala evolves with habit development, shifting from representations of choice and reward consumption to reinforcement-related signals during overtraining. Activity reflected outcome magnitude rather than identity, consistent with a role for the central amygdala in reinforcement. We also identify a previously unrecognized population of CeA neurons that track animal running speed. Together, these findings implicate the CeA as a source of reinforcement signals that promote habit formation.

Kenneth A. Amaya, James E. Carmichael, Jeffrey J. Stott et al. · 2 citations
#reinforcement learning Open access Sep 2026

NMDA receptor ablation in medial prefrontal cortex disrupts value updating and reward history integration

Schizophrenia, a serious mental illness, is associated with evidence of NMDA receptor (NMDAR) dysfunction and characterized by cognitive impairments that reflect impaired value updating and feedback-driven control; however the cellular and circuit-level mechanisms underlying these disruptions remain unclear. Here we test how NMDA receptor (NMDAR) signaling in the medial prefrontal cortex (mPFC) contributes to adaptive decision-making by combining targeted genetic ablation in mice and systemic pharmacology. Using a CRISPR-Cas9 approach to eliminate the obligate GluN1 subunit, we induced spatially confined NMDAR hypofunction in mPFC and compared its effects to systemic pharmacological blockade with the NMDAR antagonist MK-801 during performance of a touchscreen-based restless bandit task. Prefrontal NMDAR ablation impaired value discrimination, weakened the use of negative feedback, and reduced mutual information between recent outcomes and current choices, indicating disrupted reward-history integration. Reinforcement-learning models incorporating a choice-kernel term best captured behavior and revealed that NMDAR ablation selectively dampened learning and choice-history parameters governing flexible updating. Systemic MK-801 produced broad impairments in control animals, reducing accuracy, mutual information, and outcome sensitivity, yet exerted only modest additional effects after NMDAR ablation, suggesting that prefrontal NMDAR loss occluded much of the pharmacological disruption. Simulations using fitted RLCK parameters reproduced these patterns, showing convergent flattening of choice dynamics under MK-801 and persistent deficits in GluN1 ablated animals. Together, these findings demonstrate that prefrontal NMDAR signaling is necessary for effective value updating and feedback-driven learning, and that its loss recapitulates core features of systemic NMDAR hypofunction. This work establishes a mechanistic bridge between localized cortical glutamatergic dysfunction and the reinforcement-learning disturbances characteristic of schizophrenia.

Evan Knep, Angelica Velosa, Dana Mueller et al. · 1 citation

Adaptive Kantian AI: a constraint-based approach to the categorical imperative in multi-agent systems

Purpose This paper aims to address a core limitation in computational Kantian ethics: the tension between deontological rigidity and the demands of decision-making in complex, real-world environments. It proposes an adaptive extension of the categorical imperative that preserves its non-consequentialist foundations while enabling context-sensitive application. Design/methodology/approach The authors develop a constraint-based computational framework in which moral values are treated as conditions of admissibility rather than optimization targets. The model introduces dynamic tolerance and impact thresholds to enable bounded flexibility. It is empirically evaluated on the ETHICS data set through comparison with baseline Kantian and reinforcement learning from human feedback -based models. Findings The adaptive formulation improves behavioral alignment across diverse scenarios while maintaining strong deontological consistency. It systematically rejects compensatory trade-offs typical of reward-optimizing systems, while allowing conditionally admissible decisions in cases of conflicting duties under strict structural constraints. Research limitations/implications The framework depends on calibrated parameters and predefined value structures, which may introduce subjectivity and limit generalizability. Scalability and long-term stability in multi-agent settings remain open challenges. Practical implications The model supports auditable, constraint-aligned decision-making in high-risk domains requiring transparency, accountability and regulatory compliance. Originality/value This paper offers a novel constraint-based reinterpretation of Kantian ethics that integrates adaptive mechanisms without collapsing into consequentialism, advancing the development of robust, policy-compliant and ethically grounded artificial intelligence systems.

Rabah Sebti · 0 citations
#reinforcement learning Open access Sep 2026

Design and development of an endoscopic robotic system and deep reinforcement learning path planning algorithm for fine-needle biopsy of liver lesions

Abstract This paper presents the design, development and evaluation of a novel robotic platform for endoscopic ultrasound-guided fine-needle biopsy of liver lesions. The system combines a four degrees-of-freedom (DoF) two-segment tendon-driven continuum robot (TDCR) endoscope with a two DoF superelastic nickel–titanium bevel-tip steerable needle. Needle path planning is achieved using NeedleNav, a soft actor–critic (SAC) deep reinforcement learning (DRL) model that generates collision-free trajectories to deep-seated lesions. This represents one of the first integrated systems combining a TDCR, steerable needle and DRL-based navigation, and the first application of a SAC to liver lesion targeting. Evaluation of the TDCR through tip tracking of circular, diamond-shaped and arc trajectories demonstrated a mean absolute error (MAE) of 13.19 mm. NeedleNav converged to obstacle avoidance trajectories in $$\sim $$ ∼ 2500 training episodes. Needle curvature was augmented by hand-fabricating notches on its distal section. Two needles with a 3-cm and 8-cm notched section were evaluated in a gelatine liver phantom, achieving an MAE of 21.78 mm and 14.86 mm, respectively, for obstacle avoidance trajectories. The system demonstrated observable path deflection compared to obstacle-free trajectories for the same targets. Together, these findings suggest the feasibility of our proposed solution, expanding the reach of endoscopic needle interventions to deep-seated lesions in the right lobe.

Raghav Khanna, Nikola Fischer, Zhenting Du et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.