Skip to content

Category

reinforcement learning

365 papers

#reinforcement learning Open access Sep 2026

Dynamic Resource Allocation via Relational Reinforcement Learning

This paper introduces a novel Relational Reinforcement Learning (RRL) framework designed to address the challenges of dynamic resource allocation in complex environments. Traditional Reinforcement Learning (RL) approaches often fall short when dealing with environments that constantly evolve and demand adaptive resource management. The core innovation lies in learning relational embeddings that capture the intricate dependencies between resources. The agent learns a state representation where resources are represented as points in a relational embedding space, allowing it to optimize allocation strategies based on these relationships. The reward function is explicitly designed to encourage the flow of resources between related entities, promoting efficiency and robustness. This approach represents a shift from purely individual resource optimization towards a systemic view, ultimately leading to more effective and resilient resource allocation systems. The framework is presented with detailed mathematical formulations and is intended to serve as a foundation for future research in dynamic resource management.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

基于量子机器学习的动态拓扑优化算法

This paper introduces a novel dynamic topology optimization algorithm powered by quantum machine learning. Traditional topology optimization methods often require manual design of the topology, leading to limited flexibility and computational expense. Our algorithm leverages quantum machine learning to automatically learn and optimize the dynamic topology structure of data, resulting in significantly improved optimization efficiency. We propose a framework that utilizes quantum algorithms to represent and manipulate the data distribution, enabling adaptive topology adjustments throughout the optimization process. The core mechanism involves a quantum-enhanced representation of the data's topology, coupled with a reinforcement learning loop to dynamically adjust the topology based on feedback. We demonstrate the effectiveness of this approach through a series of benchmark problems, showcasing enhanced optimization speeds and improved solution quality compared to existing methods. The paper concludes with a discussion of the potential applications of this technology across diverse fields.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

Title: Adaptive Reinforcement Learning Reward Adjustment

This paper investigates a novel reinforcement learning algorithm utilizing adaptive reward function adjustment to enhance learning performance. Traditional reinforcement learning often relies on fixed reward functions, which can struggle to adapt to complex environments and unexpected situations. We propose a method that dynamically adjusts the reward function based on the learning process's inherent noise and uncertainty. This is achieved through a novel "adaptive adjustment" mechanism that continuously monitors and modifies the reward function in response to these factors. The core mechanism aims to mitigate the limitations of static reward functions, leading to improved sample efficiency and robustness in reinforcement learning. This work addresses the critical gap in current approaches by introducing a mechanism to dynamically tune the reward landscape, thereby improving the generalization capabilities of reinforcement learning agents.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

Industrial waste valorization for sustainable self-healing epoxy vitrimer hybrid composites: Taguchi optimization and machine learning-based performance prediction

This study investigates the development of sustainable self-healing vitrimer composites reinforced with recycled carbon fibers (rCF), waste cotton textile fibers (WCTF), and graphene nanoplatelets (GNP). The rCF, WCTF and GNP were in the range of 20–30%, 10–20%, and 0.5–1%, respectively, and an orthogonal design (L9) was used to create the optimal composition and curing temperature of the reinforcement. The formulation with 30% rCF, 10% WCTF, and 1% GNP, cured at 160°C exhibits high tensile strength of 138 MPa, flexural strength of 208 MPa, and an impact strength of 42 kJ/m² with a strength retention of 92% after saltwater aging. Dynamic mechanical analysis revealed a storage modulus of 5.2 GPa and a glass transition temperature of 132°C, indicating enhanced thermomechanical stability. FESEM, optical microscopy, and XRD analyses confirmed improved fiber–matrix interfacial bonding, reduced void content, and effective crack closure after healing. Taguchi and ANOVA analyses identified recycled carbon fiber content as the dominant factor affecting tensile (50.65%), flexural (49.82%), impact (36.52%), and strength retention (46.87%) properties, whereas GNP content primarily governed healing efficiency (42.63%). Machine learning models verified excellent predictive performance, with average R² values of 0.979 and 0.989 for XGBoost analysis. The results determine the potential of waste-derived reinforcements and vitrimer technology for high-performance, self-healing, and environmentally durable composite applications.

Vinod B, P. Venkataramana, Nitla Stanley Ebenezer et al. · 0 citations
#reinforcement learning Open access Sep 2026

Title: Dynamically Generated Geometric Patterns for Computational Geometry

This research investigates the application of reinforcement learning to automatically generate and optimize geometric patterns, focusing on identifying novel patterns exhibiting complex, self-organizing behavior. Traditional geometric pattern generation often relies on predefined rules, limiting the potential for truly creative and responsive designs. We propose a novel approach leveraging reinforcement learning to explore the vast space of possible patterns, rewarding patterns that demonstrate emergent complexity and self-organization. The core mechanism centers on using a reinforcement learning agent to iteratively refine patterns based on feedback, leading to the discovery of aesthetically pleasing and functionally relevant geometric forms. This work aims to advance computational geometry by providing a method for automated pattern design, pushing the boundaries of what is possible with algorithmic creativity.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Sep 2026

Progressive Dribbling Learning Model to Improve Dribbling Skills in Elementary School Children Aged 9–12 Years at SD IT Mutiara

This research aims to develop a feasible, practical, and effective progressive dribbling learning model to improve dribbling skills in elementary school children aged 9–12 years. The research was motivated by the low dribbling skills of students due to learning models that are still conventional, lack variety, and do not consider the characteristics of children's motor development according to age. The product developed is called the Progressive Islamic Soccer Dribbling Model (PISDM), a progressive training model that integrates gradual dribbling technique learning, a multimedia approach, game activities, and reinforcement of Islamic values. This research used the Research and Development (R&D) method with the ADDIE (Analysis, Design, Development, Implementation, Evaluation) development model. The research subjects were 34 students in grades IV–VI of SD IT Mutiara Duri aged 9–12 years. Data collection techniques included observation, interviews, documentation, dribbling skills tests, and expert validation. Data were analyzed using quantitative and qualitative descriptive analysis through the results of the dribbling skills pretest and posttest. The results showed that the product developed was a progressive learning model consisting of 3 meeting sessions covering ball mastery, zig-zag cone, figure-8 turn, and samba passing-dribbling. The implementation of the model provided an increase in dribbling skills in all measurement instruments. The average Boomerang Dribble score increased from 32.56 to 34.53, Triangle Dribble increased from 13.82 to 15.40, Modified Slalom Dribble increased from 21.51 to 23.94, while Illinois Football Dribble showed the highest average increase with a difference of 3.33 points after the implementation of the model. These results indicate that the progressive dribbling model can gradually improve students' ball control, speed, agility, and maneuverability, in accordance with the motor development characteristics of elementary school-aged children. Based on these results, it can be concluded that the Progressive Islamic Soccer Dribbling Model (PISDM) is suitable for use as an alternative soccer learning and training model for elementary school students aged 9–12 years because it can improve dribbling skills in a more systematic, engaging, and developmentally appropriate manner.

Atthariq Alwin, Rizki Mulyawan · 0 citations
#reinforcement learning Open access Sep 2026

The structure of infrastructure adoption: competition of hydrogen and electric for long-haul freight decarbonization

Long-haul freight decarbonization requires deploying hydrogen refueling stations (HRSs) across networks designed around incumbent diesel supply. Existing planning tools treat node-level demand as uniform, ignoring how it emerges from the coupling between freight flows and technology adoption dynamics. Understanding this coupling is essential for capital-constrained deployment decisions. This paper applies a network-coupled deployment framework—integrating Bass diffusion with reinforcement-learning-based hydrogen refueling station (HRS) sequencing under capital constraints—to a southeastern U.S. freight network derived from Freight Analysis Framework data, systematically varying hydrogen cost trajectory and deployment pace. Three findings emerge. First, deployment pace yields non-linear benefits: the adoption midpoint decreases sharply with the deployment rate but saturates quickly. Second, network position drives order-of-magnitude demand differences—a high-centrality hub generates over ten times the refueling volume of peripheral nodes at saturation—challenging uniform sizing assumptions. Third, the long-run hydrogen fuel-cell electric truck (FCET) versus battery-electric truck (BET) market split is well approximated by a linear function of the fraction of tonne-hour capacity on corridors exceeding BET range; this relation is largely insensitive to fuel-cost dynamics once hydrogen reaches diesel parity. These results show that the competition between fuel-cell electric trucks (FCETs) and battery-electric trucks (BETs) is governed by two mechanisms—the pace at which stations are deployed and the heterogeneity of demand across network positions—yielding transferable scaling relationships for infrastructure planning under capital constraints.

Mina Kim, Valerie Thomas, Benoit Montreuil · 0 citations
#large language models Open access Sep 2026

A Review of Neural Question Generation: Approaches, Challenges, and Future Directions

The goal of question generation is to automatically produce relevant and meaningful questions from diverse inputs such as knowledge bases, natural language texts, and images. With the rapid advancement of neural architectures, neural question generation (NQG) has attracted growing attention across both academia and industry. In this survey, we provide a comprehensive review of developments in NQG, spanning traditional neural approaches to the latest paradigms driven by large language models (LLMs) and multimodal large language models (MLLMs). We begin by outlining the fundamental components of NQG, including its problem formulation, benchmark datasets, evaluation metrics, and representative applications. Next, we categorize existing methods into three main types: structured NQG , which relies on structured data sources; unstructured NQG , which handles loosely structured inputs such as texts or images; and hybrid NQG , which integrates multiple modalities. For each category, we review representative neural models and synthesize the problems addressed by successive generations of methods, their remaining limitations, and the motivations behind major methodological transitions. Furthermore, we trace the progression of NQG from supervised neural approaches and pre-trained models to prompting, retrieval-augmented generation, reinforcement learning, and emerging tool-augmented and agent-based paradigms. We also discuss how recent LLMs and MLLMs have enabled more contextually aligned, knowledge-grounded, and reasoning-enhanced question generation, together with emerging concerns such as hallucination, bias, and evaluation reliability. Finally, we outline open challenges and emerging research trends, offering a forward-looking perspective on the evolution of NQG. This survey presents a meticulously curated compilation of related papers, datasets, and code, serving as a comprehensive resource for anyone studying NQG.

Shasha Guo, Liang Pang, Jing Zhang et al. · 0 citations
#large language models Open access Sep 2026

VietLegalLM: Progressive Legal Expertise Through Synthetic Comprehension and Reinforcement Learning

Legal AI systems require accuracy and verifiable reasoning, yet under-resourced languages lack the specialized models needed to meet these standards. This challenge is particularly acute for statute-based civil law systems like Vietnam’s, where the core task is interpreting and applying codified statutes rather than matching legal precedents. We address this gap by introducing a comprehensive framework for developing reliable legal AI under resource constraints. First, we present vilaw-bench, a novel evaluation benchmark tailored for Vietnamese statute-based legal reasoning that assesses cognitive abilities from basic knowledge retrieval to complex legal interpretation. Second, we propose a multi-stage training framework that progressively builds legal expertise through four phases: focused foundational learning on core legal texts, intensive comprehension practice using large-scale synthetic question-answer data, targeted skill acquisition through supervised fine-tuning, and response quality refinement using Group Relative Policy Optimization (GRPO). We implement this framework to develop VietLegalLM, training on Qwen3-1.7B-Base and Qwen3-4B-Base models. Our systematic evaluation reveals that synthetic comprehension practice produces the largest single-phase improvements in legal reasoning capabilities, while GRPO efficiently refines reasoning structure with minimal training steps. The complete sequential training approach achieves significant performance gains on vilaw-bench compared with baseline models. Our ablation studies demonstrate that practitioners can make informed tradeoffs between comprehensive training and computational efficiency: direct GRPO after pre-training offers a viable alternative to full instruction tuning when resources are limited. We release vilaw-bench, VietLegalLM, and our training framework as open-source resources, providing a reproducible roadmap for developing legal AI in other under-resourced, statute-based legal systems.

Thang Van Le, Anh-Cuong Le, Nguyen Viet Hà et al. · 0 citations
#reinforcement learning Open access Sep 2026

基于多智能体强化学习的软件开发流程优化

This paper investigates the application of Multi-Agent Reinforcement Learning (MARL) to optimize software development processes. Traditional software development methodologies often struggle with adaptability and efficiency, particularly in complex projects. This research proposes a novel approach leveraging MARL to dynamically adjust and refine the development workflow. The core idea involves modeling the software development process as a multi-agent system, where each agent represents a distinct stage or activity. These agents learn optimal strategies through interaction and reward signals, leading to improved development efficiency and quality. We present a framework for formulating this problem, detailing the agent architecture, state space, action space, and reward function. The effectiveness of the MARL approach is demonstrated through a theoretical analysis and conceptual design, highlighting its potential to overcome limitations of static, rule-based methodologies. Future work will focus on implementing and testing this framework within a simulated software development environment.

Jincheng Zhang · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.