Skip to content

Category

reinforcement learning

461 papers

#machine learning Book Open access May 2019

An Empirical Study on Female Participation in Software Project Courses

Gender issues in software engineering education are gaining research attention due to the desire to promote female participation in the field. The objective of this work is to enhance the understanding of female students' participation in software engineering projects to support gender-aware course optimization. Since 2015, we have investigated the participation of female students in terms of software engineering activities and team dynamics in a software project course that involves a real customer. We found that female students are more active with project management and requirement engineering, while they remain under-represented in highly complex or specific tasks, i.e. architecture work, and user experience design. We found no statistically significant difference in perceived team dynamics between male and female students. Insights on female project activities would facilitate the arrangement of project teams so that learning can be distributed equally across genders

Anh Nguyen-Duc, M. L. Jaccheri, P. Abrahamsson · 9 citations
#machine learning Review Open access Apr 2023

StartCards - A method for early-stage software startups

course of 4 AR cycles. During the AR process, the method was used by 44 student startup teams in a practical course setting. Data from the use of the method was collected through self-reporting in the form of modified learning diaries, mentoring meetings with the startup teams, and a qualitative survey. Results: We consider the current version of StartCards useful for early-stage startups based on the data we have collected. The method can also be used as a pedagogical tool in startup education. Conclusions: The paper presents the first published version of the method. While work on the method continues, the method is deemed ready for use.

Kai-Kristian Kemell, Anh Nguyen-Duc, Mari Suoranta et al. · 25 citations · ⚡1

VAPU: System for Autonomous Legacy Code Modernization

In this study, we present a solution for the modernization of legacy applications, an area of code generation where LLM-based multi-agent systems are proving essential for complex multi-phased tasks. Legacy applications often contain deprecated components that create compatibility, security, and reliability risks, but high resource costs make companies hesitate to update. We take a step forward to integrate an LLM-based multi-agent system as part of a legacy web application update to provide a cost-effective solution to update legacy applications autonomously. We propose a multi-agent system named a Verifying Agent Pipeline Updater (VAPU), which is designed to update code files in phases while simulating different roles in a software development team. In our previous study, we evaluated the system for legacy version updates by using six legacy web application view files by resulting errors and accomplished requirements. This study extends the previous evaluation of a multi-agent pipeline system by extending the evaluation of VAPU from a single LLM to five LLMs and using the temperature parameter in both 0 to 1 settings. Additionally, we tested the system with 20 open-source Python GitHub projects. The results of the evaluation were compared to Zero-Shot Learning (ZSL) and One-Shot Learning (OSL) prompts. The extended evaluation of VAPU showed that particularly in a low-temperature VAPU can get similar level of error count compared to the ZSL/OSL prompts but with a higher level of fulfilled requirements, depending on the LLM. VAPU showed up to 22.5% increase in the succeeding Python file update requirements compared to ZSL/OSL prompts. The study indicates that an LLM-based multi-agent system is a capable solution to update components of a legacy application autonomously.

Valtteri Ala-Salmi, Z. Rasheed, Malik Abdul Sami et al. · 3 citations

Autonomous Legacy Web Application Upgrades Using a Multi-Agent System

The use of Large Language Models (LLMs) for autonomous code generation is gaining attention in emerging technologies. As LLM capabilities expand, they offer new possibilities such as code refactoring, security enhancements, and legacy application upgrades. Many outdated web applications pose security and reliability challenges, yet companies continue using them due to the complexity and cost of upgrades. To address this, we propose an LLM-based multi-agent system that autonomously upgrades legacy web applications to the latest versions. The system distributes tasks across multiple phases, updating all relevant files. To evaluate its effectiveness, we employed Zero-Shot Learning (ZSL) and One-Shot Learning (OSL) prompts, applying identical instructions in both cases. The evaluation involved updating view files and measuring the number and types of errors in the output. For complex tasks, we counted the successfully met requirements. The experiments compared the proposed system with standalone LLM execution, repeated multiple times to account for stochastic behavior. Results indicate that our system maintains context across tasks and agents, improving solution quality over the base model in some cases. This study provides a foundation for future model implementations in legacy code updates. Additionally, findings highlight LLMs' ability to update small outdated files with high precision, even with basic prompts. The source code is publicly available on GitHub: https://github.com/alasalm1/Multi-agent-pipeline.

Valtteri Ala-Salmi, Z. Rasheed, Malik Abdul Sami et al. · 4 citations

Context Before Code: An Experience Report on Vibe Coding in Practice

Code-generating tools are increasingly used in software development, yet experience reports on conversational"vibe coding"under production constraints remain limited. This paper presents an experience report from a small full-stack team that applied contextual prompting and explicit architectural constraints to build (i) a multi-project agent learning platform designed for sustained, production-oriented use and (ii) an academic retrieval-augmented generation system. The agent platform supports multiple isolated projects, each with structured memory and background processing, thereby enforcing project-level isolation. The RAG system provides citation-grounded answers, role-based access control, and evaluation tracking. Across both systems, vibe coding accelerated scaffolding and integration. However, the generated code often under-specified isolation rules and infrastructure constraints when these were not explicitly defined. Consequently, aspects such as multi-tenancy, access control, memory policies, and asynchronous processing required deliberate architectural design and verification. We observe a shift in engineering effort from boilerplate implementation toward constraint specification and enforcement auditing. We also identify recurring architectural"non-delegation zones"where conversational code generation remains insufficient for production reliability.

Md Nasir Uddin Shuvo, M. Islam, Mahade Hasan et al. · 0 citations
#machine learning Open access Jun 2025

Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation

Retrieval-Augmented Generation (RAG) systems are emerging as a key approach for grounding Large Language Models (LLMs) in external knowledge, addressing limitations in factual accuracy and contextual relevance. However, there is a lack of empirical studies that report on the development of RAG-based implementations grounded in real-world use cases, evaluated through general user involvement, and accompanied by systematic documentation of lessons learned. This paper presents five domain-specific RAG applications developed for real-world scenarios across governance, cybersecurity, agriculture, industrial research, and medical diagnostics. Each system incorporates multilingual OCR, semantic retrieval via vector embeddings, and domain-adapted LLMs, deployed through local servers or cloud APIs to meet distinct user needs. A web-based evaluation involving a total of 100 participants assessed the systems across six dimensions: (i) Ease of Use, (ii) Relevance, (iii) Transparency, (iv) Responsiveness, (v) Accuracy, and (vi) Likelihood of Recommendation. Based on user feedback and our development experience, we documented twelve key lessons learned, highlighting technical, operational, and ethical challenges affecting the reliability and usability of RAG systems in practice.

M. Hasan, Muhammad Waseem, Kai-Kristian Kemell et al. · 10 citations · ⚡1
#machine learning Open access Feb 2025

Anomaly detection in smart power grids with graph-regularized MS-SVDD: a multimodal subspace learning approach

Anomaly detection in smart power grids is a critical challenge due to the complexity, heterogeneity, and dynamic nature of sensor data streams. Existing one-class classification methods, particularly Subspace Support Vector Data Description (SVDD), have been extended to multimodal scenarios but often fail to fully exploit the structural dependencies across modalities, limiting their robustness in real-world applications. In this paper, we address this gap by proposing a generalized Multimodal Subspace Support Vector Data Description (MS-SVDD) model with graph-embedded regularization. The method projects data from multiple modalities into a shared low-dimensional subspace while preserving modality-specific structure through Laplacian regularizers. Our approach is evaluated on a three-modality dataset derived from smart grid event time series, using a dedicated preprocessing pipeline for constructing one-class classification training samples. The results demonstrate that our graph-embedded MS-SVDD improves robustness of event detection compared to conventional approaches, highlighting the potential of integrating graph priors with multimodal subspace learning for advancing anomaly detection in critical infrastructure. More broadly, this work contributes to the wider field of AI by illustrating how relational and structural information can be systematically embedded into one-class models, enabling robust learning under complex, high-dimensional, and multimodal conditions.

Thomas Debelle, F. Sohrab, Pekka Abrahamsson et al. · 1 citation
#reinforcement learning Preprint Aug 2026

Guided Riemannian Optimization (GuRO): Bridging Model Predictive Control and Decision Transformers

Decision-making in high-dimensional, nonlinear systems remains a central challenge in robotics. While model-based methods like Model Predictive Control (MPC) offer sample efficiency and interpretability, their performance degrades when the dynamics model is inaccurate or long-horizon predictions are required. Conversely, model-free reinforcement learning (RL) learns policies directly from interaction but suffers from high sample complexity and unstable optimization. Recent advances in sequence modeling have inspired transformer-based decision-making frameworks that can unify MPC and RL, but their training typically faces significant optimization challenges due to highly non-convex loss landscapes. In this work, we propose a novel framework that integrates MPC with RL in a sequence decision-making framework and leverages a curvature-aware optimization to efficiently tackle non-convex loss landscapes. MPC provides predictions of locally optimal trajectories that guide the decision transformer, removing the need for extensive offline pretraining. To address the slow and unstable convergence of traditional optimizers, we train the policy in a Riemannian parameter space using an efficient Riemannian (curvature-aware) method, leading to faster and more robust optimization. We evaluate our framework on high-dimensional quadruped control tasks and demonstrate consistent improvements over strong baselines, including TRPO, SAC, and Online Decision Transformer, achieving higher returns and faster convergence.

Hossein Abdi, Satya Dash, Mingfei Sun · 0 citations
#reinforcement learning Open access Aug 2026

DIArc Foundational Note v0.1 — Minimum Claim Edition

Abstract The rapid development of artificial intelligence has significantly increased the availability of information, analytical capability, and machine-assisted reasoning. However, greater access to information does not necessarily produce better decisions. In many organizational contexts, the emerging bottleneck is no longer information acquisition, but the human and organizational capacity to determine what information is sufficient, when analysis should stop, when a decision should be made, and how outcomes should improve future judgment. This Foundational Note introduces Decision Intelligence Architecture (DIArc) as an architectural framework for Human–AI collaborative decision systems. DIArc is based on a central proposition: in the AI era, competitive advantage increasingly depends not on maximizing information, but on maximizing the rate at which high-quality decisions generate learning and improve judgment, under explicit constraints on information consumption and decision cycles. The architecture is organized into four theoretical layers. First, the Capability Inversion Hypothesis describes a structural shift in which information, knowledge, and analysis become increasingly abundant while judgment, commitment, execution, and learning become comparatively scarce capabilities. Second, Identity-driven Information Consumption (IDIC) describes a decision failure mechanism in which continued information consumption may serve identity reinforcement rather than decision improvement. Third, the Decision Constraint Architecture, comprising Decision Information Budget (DIB) and Decision Cycle Budget (DCB), introduces explicit constraints on information consumption and analytical iteration. Fourth, High-quality Decision Velocity (HQDV) describes the performance objective of accelerating completed high-quality decision loops, while Judgment Evolution Rate (JER) represents the longer-term evolutionary objective of improving judgment through outcome-based learning. This note constitutes the initial public disclosure of the DIArc architecture and establishes its theoretical baseline for subsequent research and branch concepts.

Lucas Xiaochun Xu · 0 citations
#reinforcement learning Open access Aug 2026

Evaluating a School Waste Bank Strategy for Strengthening Students’ Environmental Responsibility: A Case Study at SDN 4 Tanggungharjo

This study aims to evaluate the strategy for strengthening students’ environmental care character through the Waste Bank program at SDN 4 Tanggungharjo, Grobogan Regency. This study employed a qualitative approach with a case study design. Data were collected through observation, interviews, and documentation involving the principal, teachers, students, Waste Bank management team, and other relevant school stakeholders. Data analysis was conducted through data condensation, data display, and conclusion drawing, while data credibility was established through source and technique triangulation. The findings indicate that the evaluation of the Waste Bank strategy was conducted through four interconnected mechanisms: periodic evaluation and collective reflection, financial transparency and program accountability, adaptive responses to problems, and program sustainability. Periodic evaluation enabled the school to identify problems related to student participation, waste sorting, and program management and to formulate corrective actions collaboratively. Financial transparency was maintained through individual student savings records and the main Waste Bank financial records, strengthening accountability and trust. Adaptive responses transformed students’ mistakes in waste sorting into learning opportunities through additional explanation, guidance, and repeated practice. Program sustainability was supported by institutional planning and budgeting, adequate facilities, continued student participation, and cooperation with external waste-management partners. Overall, the evaluation process demonstrated that the Waste Bank had developed beyond a waste-collection activity into a school-based mechanism for strengthening environmental care character. The evaluation functioned as a feedback mechanism connecting reflection, corrective action, behavioral reinforcement, accountability, and institutional sustainability. The study concludes that the effectiveness of a school-based Waste Bank strategy depends not only on the implementation of environmental activities but also on the school’s capacity to continuously evaluate, adapt, and institutionalize the program to support students’ environmental responsibility.

Nur Solikin, Endang Wuryandini, Widya Kusumaningsih · 0 citations
#reinforcement learning Open access Aug 2026

Adaptive Controllable Emergence in Multi-Task Air and Space Defense Systems: A Framework for Mission Reconfiguration, Resilient Coordination, Intelligent Decision-Making, and Dynamic Resource Allocation

Emergent collective intelligence provides an important theoretical and computational perspective for understanding how locally interacting agents can generate coordinated global behaviors that cannot be explained by the behavior of individual agents alone. In large-scale air and space defense systems, this property is particularly relevant because heterogeneous sensing, decision-making, communication, and execution resources must operate under dynamic environments, incomplete information, changing mission requirements, limited resources, and potentially degraded communication conditions. However, conventional controllable-emergence models generally assume relatively stable task structures and predefined interaction rules, which limits their adaptability when multiple tasks arrive concurrently or when the network topology and available resources change over time.This study proposes an Adaptive Controllable Emergence (ACE) framework for multi-task air and space defense systems. The proposed framework extends graph-based multi-agent modeling and multi-agent reinforcement learning by introducing four coupled mechanisms: dynamic mission reconfiguration, resilient coordination, intelligent distributed decision-making, and dynamic resource allocation. The system is represented as a time-varying interaction graph in which sensing, decision, and execution agents dynamically modify their relationships according to mission requirements and resource availability. A decentralized partially observable Markov decision process is employed to formulate local decision-making under incomplete information. A multi-objective reward function jointly considers mission completion, coordination quality, resource utilization, network resilience, adaptation cost, and decision latency. Furthermore, a mission-reconfiguration mechanism is introduced to enable the system to modify task-agent assignments when the task set, network topology, or resource state changes. The resulting framework transforms controllable emergence from a static rule-design problem into an adaptive optimization process in which microscopic policies continuously modify macroscopic system behavior. The proposed mathematical formulation provides a basis for analyzing emergence quality, adaptation speed, coordination robustness, resource efficiency, and convergence. A simulation framework is also developed for evaluating the proposed architecture under static, dynamic, multi-task, and communication-degradation scenarios. The framework is intended as a general computational model for studying adaptive coordination in large-scale multi-agent systems rather than as a platform-specific operational defense procedure.

Nor Ahmed Gujar · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.