Skip to content

Category

reinforcement learning

415 papers

Context Before Code: An Experience Report on Vibe Coding in Practice

Code-generating tools are increasingly used in software development, yet experience reports on conversational"vibe coding"under production constraints remain limited. This paper presents an experience report from a small full-stack team that applied contextual prompting and explicit architectural constraints to build (i) a multi-project agent learning platform designed for sustained, production-oriented use and (ii) an academic retrieval-augmented generation system. The agent platform supports multiple isolated projects, each with structured memory and background processing, thereby enforcing project-level isolation. The RAG system provides citation-grounded answers, role-based access control, and evaluation tracking. Across both systems, vibe coding accelerated scaffolding and integration. However, the generated code often under-specified isolation rules and infrastructure constraints when these were not explicitly defined. Consequently, aspects such as multi-tenancy, access control, memory policies, and asynchronous processing required deliberate architectural design and verification. We observe a shift in engineering effort from boilerplate implementation toward constraint specification and enforcement auditing. We also identify recurring architectural"non-delegation zones"where conversational code generation remains insufficient for production reliability.

Md Nasir Uddin Shuvo, M. Islam, Mahade Hasan et al. · 0 citations
#machine learning Open access Jun 2025

Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation

Retrieval-Augmented Generation (RAG) systems are emerging as a key approach for grounding Large Language Models (LLMs) in external knowledge, addressing limitations in factual accuracy and contextual relevance. However, there is a lack of empirical studies that report on the development of RAG-based implementations grounded in real-world use cases, evaluated through general user involvement, and accompanied by systematic documentation of lessons learned. This paper presents five domain-specific RAG applications developed for real-world scenarios across governance, cybersecurity, agriculture, industrial research, and medical diagnostics. Each system incorporates multilingual OCR, semantic retrieval via vector embeddings, and domain-adapted LLMs, deployed through local servers or cloud APIs to meet distinct user needs. A web-based evaluation involving a total of 100 participants assessed the systems across six dimensions: (i) Ease of Use, (ii) Relevance, (iii) Transparency, (iv) Responsiveness, (v) Accuracy, and (vi) Likelihood of Recommendation. Based on user feedback and our development experience, we documented twelve key lessons learned, highlighting technical, operational, and ethical challenges affecting the reliability and usability of RAG systems in practice.

M. Hasan, Muhammad Waseem, Kai-Kristian Kemell et al. · 10 citations · ⚡1
#machine learning Open access Feb 2025

Anomaly detection in smart power grids with graph-regularized MS-SVDD: a multimodal subspace learning approach

Anomaly detection in smart power grids is a critical challenge due to the complexity, heterogeneity, and dynamic nature of sensor data streams. Existing one-class classification methods, particularly Subspace Support Vector Data Description (SVDD), have been extended to multimodal scenarios but often fail to fully exploit the structural dependencies across modalities, limiting their robustness in real-world applications. In this paper, we address this gap by proposing a generalized Multimodal Subspace Support Vector Data Description (MS-SVDD) model with graph-embedded regularization. The method projects data from multiple modalities into a shared low-dimensional subspace while preserving modality-specific structure through Laplacian regularizers. Our approach is evaluated on a three-modality dataset derived from smart grid event time series, using a dedicated preprocessing pipeline for constructing one-class classification training samples. The results demonstrate that our graph-embedded MS-SVDD improves robustness of event detection compared to conventional approaches, highlighting the potential of integrating graph priors with multimodal subspace learning for advancing anomaly detection in critical infrastructure. More broadly, this work contributes to the wider field of AI by illustrating how relational and structural information can be systematically embedded into one-class models, enabling robust learning under complex, high-dimensional, and multimodal conditions.

Thomas Debelle, F. Sohrab, Pekka Abrahamsson et al. · 1 citation
#reinforcement learning Open access Aug 2026

DIArc Foundational Note v0.1 — Minimum Claim Edition

Abstract The rapid development of artificial intelligence has significantly increased the availability of information, analytical capability, and machine-assisted reasoning. However, greater access to information does not necessarily produce better decisions. In many organizational contexts, the emerging bottleneck is no longer information acquisition, but the human and organizational capacity to determine what information is sufficient, when analysis should stop, when a decision should be made, and how outcomes should improve future judgment. This Foundational Note introduces Decision Intelligence Architecture (DIArc) as an architectural framework for Human–AI collaborative decision systems. DIArc is based on a central proposition: in the AI era, competitive advantage increasingly depends not on maximizing information, but on maximizing the rate at which high-quality decisions generate learning and improve judgment, under explicit constraints on information consumption and decision cycles. The architecture is organized into four theoretical layers. First, the Capability Inversion Hypothesis describes a structural shift in which information, knowledge, and analysis become increasingly abundant while judgment, commitment, execution, and learning become comparatively scarce capabilities. Second, Identity-driven Information Consumption (IDIC) describes a decision failure mechanism in which continued information consumption may serve identity reinforcement rather than decision improvement. Third, the Decision Constraint Architecture, comprising Decision Information Budget (DIB) and Decision Cycle Budget (DCB), introduces explicit constraints on information consumption and analytical iteration. Fourth, High-quality Decision Velocity (HQDV) describes the performance objective of accelerating completed high-quality decision loops, while Judgment Evolution Rate (JER) represents the longer-term evolutionary objective of improving judgment through outcome-based learning. This note constitutes the initial public disclosure of the DIArc architecture and establishes its theoretical baseline for subsequent research and branch concepts.

Lucas Xiaochun Xu · 0 citations
#reinforcement learning Open access Aug 2026

Evaluating a School Waste Bank Strategy for Strengthening Students’ Environmental Responsibility: A Case Study at SDN 4 Tanggungharjo

This study aims to evaluate the strategy for strengthening students’ environmental care character through the Waste Bank program at SDN 4 Tanggungharjo, Grobogan Regency. This study employed a qualitative approach with a case study design. Data were collected through observation, interviews, and documentation involving the principal, teachers, students, Waste Bank management team, and other relevant school stakeholders. Data analysis was conducted through data condensation, data display, and conclusion drawing, while data credibility was established through source and technique triangulation. The findings indicate that the evaluation of the Waste Bank strategy was conducted through four interconnected mechanisms: periodic evaluation and collective reflection, financial transparency and program accountability, adaptive responses to problems, and program sustainability. Periodic evaluation enabled the school to identify problems related to student participation, waste sorting, and program management and to formulate corrective actions collaboratively. Financial transparency was maintained through individual student savings records and the main Waste Bank financial records, strengthening accountability and trust. Adaptive responses transformed students’ mistakes in waste sorting into learning opportunities through additional explanation, guidance, and repeated practice. Program sustainability was supported by institutional planning and budgeting, adequate facilities, continued student participation, and cooperation with external waste-management partners. Overall, the evaluation process demonstrated that the Waste Bank had developed beyond a waste-collection activity into a school-based mechanism for strengthening environmental care character. The evaluation functioned as a feedback mechanism connecting reflection, corrective action, behavioral reinforcement, accountability, and institutional sustainability. The study concludes that the effectiveness of a school-based Waste Bank strategy depends not only on the implementation of environmental activities but also on the school’s capacity to continuously evaluate, adapt, and institutionalize the program to support students’ environmental responsibility.

Nur Solikin, Endang Wuryandini, Widya Kusumaningsih · 0 citations
#reinforcement learning Open access Aug 2026

Adaptive Controllable Emergence in Multi-Task Air and Space Defense Systems: A Framework for Mission Reconfiguration, Resilient Coordination, Intelligent Decision-Making, and Dynamic Resource Allocation

Emergent collective intelligence provides an important theoretical and computational perspective for understanding how locally interacting agents can generate coordinated global behaviors that cannot be explained by the behavior of individual agents alone. In large-scale air and space defense systems, this property is particularly relevant because heterogeneous sensing, decision-making, communication, and execution resources must operate under dynamic environments, incomplete information, changing mission requirements, limited resources, and potentially degraded communication conditions. However, conventional controllable-emergence models generally assume relatively stable task structures and predefined interaction rules, which limits their adaptability when multiple tasks arrive concurrently or when the network topology and available resources change over time.This study proposes an Adaptive Controllable Emergence (ACE) framework for multi-task air and space defense systems. The proposed framework extends graph-based multi-agent modeling and multi-agent reinforcement learning by introducing four coupled mechanisms: dynamic mission reconfiguration, resilient coordination, intelligent distributed decision-making, and dynamic resource allocation. The system is represented as a time-varying interaction graph in which sensing, decision, and execution agents dynamically modify their relationships according to mission requirements and resource availability. A decentralized partially observable Markov decision process is employed to formulate local decision-making under incomplete information. A multi-objective reward function jointly considers mission completion, coordination quality, resource utilization, network resilience, adaptation cost, and decision latency. Furthermore, a mission-reconfiguration mechanism is introduced to enable the system to modify task-agent assignments when the task set, network topology, or resource state changes. The resulting framework transforms controllable emergence from a static rule-design problem into an adaptive optimization process in which microscopic policies continuously modify macroscopic system behavior. The proposed mathematical formulation provides a basis for analyzing emergence quality, adaptation speed, coordination robustness, resource efficiency, and convergence. A simulation framework is also developed for evaluating the proposed architecture under static, dynamic, multi-task, and communication-degradation scenarios. The framework is intended as a general computational model for studying adaptive coordination in large-scale multi-agent systems rather than as a platform-specific operational defense procedure.

Nor Ahmed Gujar · 0 citations
#reinforcement learning Dataset Open access Aug 2026

A Twin Delayed Deep Deterministic-based control method for a full vehicle semi-active suspension system

Through this system, users can input parameters for a vehicle semi-active suspension using a magneto-rheological damper (sprung mass, sprung mass centroid position parameters, unsprung mass, magneto-rheological damper model parameters, suspension spring stiffness, wheel equivalent spring stiffness, balance bar torsion spring stiffness, and ground excitation). Through this program's calculations, a vehicle suspension deep reinforcement learning controller model can be obtained to optimize the shock absorption effect at the vehicle's center of mass. Development hardware environment: CPU Intel Core i5 11600k, Memory: 32GB, Hard disk space: 4TB; Runtime hardware environment: CPU Intel Core i5 10400, Memory: 8GB, Hard disk space: 500GB. Development software environment: Windows 10; Runtime software environment: Windows 10.

Yongjun Wang, Xiaoming Wang, Gang Zhi et al. · 0 citations
#reinforcement learning Open access Aug 2026

智能电动汽车品牌声音DNA的动态生成——基于深度学习的主动声音设计系统研究

(ASD) widespread adoption of intelligent electric vehicles (IEVs) has eliminated traditional engine noise, raising pedestrian safety concerns and accelerating product homogeneity, which severely weakens brand auditory identity. Active sound design (ASD) has thus become a core technology for reshaping brand sound DNA and enhancing in‑cabin immersion and interaction. However, existing ASD systems largely rely on static concatenation of audio samples or fixed rule‑based parameter mapping, struggling to cope with complex driving scenarios and failing to deliver dynamic evolution or personalized expression of brand sound DNA. To address this limitation, this paper presents a deep learning‑based system for dynamic brand sound DNA generation and active sound design. We construct a multidimensional acoustic feature corpus of brand sound DNA and quantitatively decode the deep mapping between acoustic parameters and brand emotional semantics (e.g., sense of technology, sportiness, and luxury). A sequential audio generation model is designed by fusing a conditional generative adversarial network (cGAN) with a long short‑term memory (LSTM) network. An online adaptation mechanism built on deep reinforcement learning is introduced, which collects real‑time physiological feedback and subjective evaluations to dynamically fine‑tune the sound generation policy, enabling personalized evolution of the brand sound. This work provides an innovative technical pathway for auditory interaction design in IEVs and advances automotive acoustic engineering from static presets toward dynamic intelligent generation.

瑞珅 张 · 0 citations
#reinforcement learning Open access Aug 2026

Implementasi Nilai Karakter Pada Cerita Rakyat Lampung dalam Pembelajaran Bahasa Lampung SDN 3 Tegineneng Kelas V

This study aims to describe the implementation of character values through Lampung folklore in Lampung language learning among fifth-grade students at SDN 3 Tegineneng, Pesawaran. A descriptive qualitative approach was employed, involving three purposively selected informants: one Lampung language teacher as the main informant, supported by one school principal and one fifth-grade student. Data were collected through non-participant observation, semi-structured interviews, and documentation and analyzed using data condensation, data display, and conclusion drawing and verification. The findings show that character values were integrated into the introductory, core, and closing stages of learning. Nine character values were identified and actualized: religiosity, honesty, discipline, responsibility, hard work, caring, cooperation, patriotism, and respect for diversity. The teacher strengthened these values through exemplary behavior, habituation, discussion, motivation, appreciation, and reflection. Implementation was supported by school commitment, teacher competence, learning resources, a conducive classroom environment, and student participation, while limited instructional time, differences in Lampung language proficiency, limited learning media, and students' confidence constituted constraints. The main contribution of this study is the conceptual pattern of value internalization–actualization–reinforcement, which explains how moral values represented in folklore are understood by students, practiced through classroom activities, and reinforced through teachers' pedagogical actions.

Mela Santika, Chairul Amriyah, Yudesta Erfayliana · 0 citations
#reinforcement learning Review Open access Aug 2026

Reinforcement Learning in Wearable Robotic Systems for Orthopedic Rehabilitation: An Elbow-Focused Narrative Review

Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.

Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al. · 0 citations
#reinforcement learning Review Open access Aug 2026

Reinforcement Learning in Wearable Robotic Systems for Orthopedic Rehabilitation: An Elbow-Focused Narrative Review

Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.

Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.