Short-term Heart Rate Variability (HRV) forecasting could provide clinicians with actionable lead time for detecting autonomic dysfunction and adverse cardiac events. Consumer wearable devices generate fragmented, artifact-rich HRV signals that challenge conventional forecasting approaches. In this study, we evaluated the forecasting ability of three Time Series Foundation Models (TSFMs), TimesFM, Chronos, and MOIRAI, against traditional baselines (Mean, Exponential Smoothing, and Exponentially Weighted Moving Average) on real-world wearable data collected from 49 healthy individuals. To address data fragmentation, we introduce a variability-preserving imputation method that augments linear interpolation with locally adaptive stochastic noise, retaining physiological dynamics essential for accurate forecasting. The results show that TSFMs outperformed all baselines without fine-tuning, achieving average Mean Absolute Scaled Error (MASE) between 0.81 and 0.87 across TSFMs and both context lengths (32 and 64 time steps), with Chronos and TimesFM as the top models, though MOIRAI showed limited gains over baselines. With up to a 2-hour forecast horizon, the results establish a baseline for TSFMs'performance on a real-world dataset, highlighting domain-specific fine-tuning as a promising direction for clinical deployment.
Luukas Peräkylä, F. Sohrab, Ville Hautamäki et al.· 0 citations
Automated authoring of Gherkin Behavior-Driven Development (BDD) acceptance criteria remains a manual bottleneck in requirements engineering. This study investigates whether epic-organized LLM-generated Gherkin produces higher quality and coverage than requirement-aligned generation. We compare our Timeless (an epic-organized LLM pipeline) approach against a naive large language model (LLM) baseline on four requirements documents (107 requirements) from the PURE dataset. Evaluation covers structural metrics, automated requirement coverage via TF-IDF and dense embeddings, and blind expert assessment by four researchers. In our evaluation, the JSON-constrained pipeline produced structurally valid scenarios across all generated outputs, while the zero-shot baseline achieved 99% structural validity. Semantic coverage was comparable to the baseline, with Timeless achieving 94.3% semantic Requirement Coverage Rate compared with 92.9% for the baseline. TF-IDF produced lower coverage scores for the epic-organized output, suggesting that lexical metrics may miss coverage when scenarios paraphrase requirements at a higher level of abstraction. Expert raters prefer the epic-organized strategy on Correctness (4.61 vs 4.14), Executability (4.61 vs 4.07), and Completeness (4.31 vs 3.50). Overall, the results suggest that epic-organized generation can improve perceived Gherkin quality while maintaining comparable semantic coverage, although broader replication is needed before generalizing this finding.
Shahbaz Siddeeq, M. Abbasi, Jussi Rasku et al.· 0 citations
In agile software development, maintaining high-quality user stories is crucial, but also challenging. This study explores the use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams. We developed a reference model for an Autonomous LLM-based Agent System and implemented it at the company. The quality of user stories in the study and the effectiveness of these agents for user story quality improvement was assessed by 11 participants across six agile teams. Our findings demonstrate the potential of LLMs in improving user story quality, contributing to the research on AI role in agile development, and providing a practical example of the transformative impact of AI in an industry setting.
Zheying Zhang, M. Rayhan, Tomas Herda et al.· International Conference on...· 48 citations· ⚡4
Retrieval-Augmented Generation (RAG) systems are emerging as a key approach for grounding Large Language Models (LLMs) in external knowledge, addressing limitations in factual accuracy and contextual relevance. However, there is a lack of empirical studies that report on the development of RAG-based implementations grounded in real-world use cases, evaluated through general user involvement, and accompanied by systematic documentation of lessons learned. This paper presents five domain-specific RAG applications developed for real-world scenarios across governance, cybersecurity, agriculture, industrial research, and medical diagnostics. Each system incorporates multilingual OCR, semantic retrieval via vector embeddings, and domain-adapted LLMs, deployed through local servers or cloud APIs to meet distinct user needs. A web-based evaluation involving a total of 100 participants assessed the systems across six dimensions: (i) Ease of Use, (ii) Relevance, (iii) Transparency, (iv) Responsiveness, (v) Accuracy, and (vi) Likelihood of Recommendation. Based on user feedback and our development experience, we documented twelve key lessons learned, highlighting technical, operational, and ethical challenges affecting the reliability and usability of RAG systems in practice.
M. Hasan, Muhammad Waseem, Kai-Kristian Kemell et al.· EUROMICRO Conference on Soft...· 10 citations· ⚡1
Reach audiences
Advertise in front of researchers, engineers, and readers.
Refactoring is a constant activity in software development and maintenance. Scale and maintain software systems are based on code refactoring. However, this process is still labor intensive, as it requires programmers to analyze the codebases in detail to avoid introducing new defects. In this research, we put forward a large language model (LLM)-based multi-agent system to automate the refactoring process on Haskell code. The objective of this research is to evaluate the effect of LLM-based agents in performing structured and semantically accurate refactoring on Haskell code. Our proposed multi-agent system based on specialized agents with distinct roles, including code analysis, refactoring execution, verification, and debugging. To test the effectiveness and practical applicability of the multi-agent system, we conducted evaluations using different open-source Haskell codebases. The results of the experiments carried out showed that the proposed LLM-based multi-agent system could average 11.03% decreased complexity in code, an improvement of 22.46% in overall code quality, and increase performance efficiency by an average of 13.27%. Furthermore, memory allocation was optimized by up to 14.57%. These results highlight the ability of LLM-based multi-agent in managing refactoring tasks targeted toward functional programming paradigms. Our findings hint that LLM-based multi-agent systems integration into the refactoring of functional programming languages can enhance maintainability and support automated development workflows.
Shahbaz Siddeeq, Muhammad Waseem, Z. Rasheed et al.· International Conference on...· 4 citations
Anomaly detection in smart power grids is a critical challenge due to the complexity, heterogeneity, and dynamic nature of sensor data streams. Existing one-class classification methods, particularly Subspace Support Vector Data Description (SVDD), have been extended to multimodal scenarios but often fail to fully exploit the structural dependencies across modalities, limiting their robustness in real-world applications. In this paper, we address this gap by proposing a generalized Multimodal Subspace Support Vector Data Description (MS-SVDD) model with graph-embedded regularization. The method projects data from multiple modalities into a shared low-dimensional subspace while preserving modality-specific structure through Laplacian regularizers. Our approach is evaluated on a three-modality dataset derived from smart grid event time series, using a dedicated preprocessing pipeline for constructing one-class classification training samples. The results demonstrate that our graph-embedded MS-SVDD improves robustness of event detection compared to conventional approaches, highlighting the potential of integrating graph priors with multimodal subspace learning for advancing anomaly detection in critical infrastructure. More broadly, this work contributes to the wider field of AI by illustrating how relational and structural information can be systematically embedded into one-class models, enabling robust learning under complex, high-dimensional, and multimodal conditions.
Thomas Debelle, F. Sohrab, Pekka Abrahamsson et al.· Scientific Reports· 1 citation
This paper presents MARARE, a real-time multi-agent system that transforms meeting dialogues into structured software requirements. One agent interacts with participants, while background agents extract and verify requirements collaboratively. Evaluation using the LLM-as-a-Judge method across five meetings (5–8 minutes each) shows a mean coverage of 80.0 ± 11.2 % (mean ± SD), semantic similarity of 0.86 ± 0.05, and hallucination rate of 14.3 ± 6.2 %. Preliminary results indicate performance differences across LLMs, suggesting that model choice influences coverage, consistency, and hallucination rates.
Malik Abdul Sami, Gessé Evangelista, Kai-Kristian Kemell et al.· AGENT@ICSE· 0 citations
Context: Organizations adopting Artificial Intelligence (AI) face challenges in eliciting and analyzing requirements that align with strategic objectives, especially when human oversight and iterative refinement are needed. Large Language Models (LLMs)-based Multi-agent systems provide a potential solution by supporting structured and collaborative Requirements Engineering (RE) processes for AI adoption planning.
Objective: The objective of this study is to investigate whether a multi-agent system, built on LLMs and supported by human input, can assist in requirements analysis for AI adoption. Method: We used a mixed-method approach: (i) designed and developed a multi-agent system to support the generation and prioritization of requirements for AI adoption, (ii) conducted multiple case studies with four companies to evaluate the system, and (iii) collected data through post-session questionnaires from nine participants and follow-up interviews, one per company.
Results: Questionnaire and interview findings together indicate that the system may assist in identifying relevant and goal-aligned requirements. Seven participants considered the generated requirements relevant, and six found them aligned with organizational goals. Participants noted that iterative feedback improved completeness and feasibility, often within two feedback rounds. Both data sources show that human input was essential to clarify technical details, ensure contextual accuracy, and validate prioritization results. Participants from all companies also identified usability, transparency, and scalability as areas requiring further refinement for broader organizational use.
Conclusions: LLM-based multi-agent systems can support strategic AI planning by enabling iterative refinement with human experts. Future work will include more interviews with stakeholders and adjustments to system features to improve transparency, usability, and scalability.
Malik Abdul Sami, Zheying Zhang, Muhammad Waseem et al.· e-Informatica Software Engin...· 6 citations
Large Language Models (LLMs) offer new opportunities for automated code refactoring. However, generated changes must reduce targeted quality problems without introducing new issues or altering behaviour-relevant code structures. We introduce REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution), a tool-agnostic, evidence-aware multi-agent approach for generating Java file-level refactoring candidates. REFINE combines static-analysis-guided smell identification, smell-informed planning, LLM-based transformation, automated re-analysis, preservation checks, and structured reporting. We evaluate REFINE on 450 Java files from 15 open-source systems, producing 1,350 model-pass outputs using OpenAI GPT-5.5, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.8. REFINE reduces detected code smells by 68.26%, 72.79%, and 68.49% across the three configurations, respectively, with the strongest reductions observed for major smells. A matched 150-file direct-prompt baseline shows that REFINE achieves a higher median code-smell reduction with smaller edits and fewer public-method removals. However, broader quality improvements are inconsistent, and preservation checks reveal residual risks, including assert/fail-call changes and public-method removal. Therefore, REFINE outputs should be treated as refactoring candidates requiring compilation, testing, dependency analysis, and human review before adoption in repository- or system-level settings.
Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson· 0 citations
This paper examines whether differences in the speed with which traded assets respond to a common market shock can predict subsequent relative returns. The framework combines a lagged rolling factor model with Absorption Gap (AG), which measures an asset’s response error, and Shock Coherence (SC), which characterizes the contemporaneous market state. The public specification is evaluated using executable next-open timing, explicit transaction costs, dependence-aware inference, randomized-signal benchmarks, chronological diagnostics, and machine-learning extensions. The study uses 24 ETFs from 4 January 2010 through 28 August 2026, with eight factor proxies excluded from the 16-asset traded cross-section. The corrected public baseline produces a combined Rank IC of -0.00592, an approximately flat zero-cost gross result, and materially negative performance after transaction costs. A within-date randomized-signal benchmark yields an empirical two-sided p-value of 0.299, while standalone Absorption Gap, coherence-conditioned tests, chronological subsamples, and machine-learning models provide no robust evidence of economically viable public alpha. The contribution is therefore methodological as much as empirical: the paper connects an economic hypothesis about heterogeneous information absorption to an executable trading test, documents why the disclosed implementation fails, separates diagnostic and exploratory analysis from confirmatory evidence, and establishes a reproducible public baseline while keeping the proprietary alpha layer outside the evidence package.
Khaybullina Alina· Zenodo (CERN European Organi...· 0 citations
Algal-bacterial biofilms function as stratified microecosystems in which photosynthetic oxygen production and heterotrophic respiration are coupled through counter-diffusive \(\:{\text{O}}_{2}-{\text{C}\text{O}}_{2}\:\) transport. Quantifying this internal carbon loop is critical for designing carbon-neutral and energy-efficient wastewater treatment systems. This study develops a mechanistic two-layer diffusion–reaction model representing an outer phototrophic layer, where CO 2 is assimilated and O 2 is produced, and an inner heterotrophic layer, where O 2 is consumed, and CO 2 is released. Steady-state mass balances with interfacial flux continuity are solved to obtain coupled O 2 and CO 2 profiles across the biofilm. A new dimensionless Carbon Loop Efficiency Index (CLEI) is introduced to quantify the degree of photosynthetic–respiratory coupling based on interfacial gas fluxes. By definition, CLEI = 0 represents a carbon-positive regime with no internal CO 2 recycling, CLEI = 1 denotes complete loop closure and carbon-neutral operation, and CLEI > 1 indicates a potentially carbon-negative regime in which CO 2 assimilation exceeds internally generated CO 2 under the modeled conditions. Model results suggest that, within the validated parameter ranges considered in this study, CLEI approaches unity only within a constrained design window characterized by intermediate total biofilm thickness (approximately 300–500 μm), where oxygen and carbon dioxide transport are balanced through coupled diffusion–reaction processes. Thinner biofilms are carbon-limited (CLEI > 1), whereas thicker biofilms develop oxygen diffusion limitation (CLEI < 1), preventing complete loop closure. Thinner biofilms are carbon-limited (CLEI > 1), whereas thicker biofilms develop oxygen diffusion limitation (CLEI < 1), preventing full loop closure. The framework demonstrates how stratified algal–bacterial biofilms can achieve self-oxygenation and intrinsic CO₂ mitigation through transport-controlled coupling, and provides a mechanistic basis for the rational design of carbon-neutral or potentially carbon-negative wastewater treatment and photobioreactor systems.
Deepak Sharma, Rakesh Choudhary, Y. Duraisamy et al.· Scientific Reports· 0 citations
The prevailing narrative of the AI race assumes that technological competition culminates in a single winner, a malformed rhetoric that rests on an under-specified concept of victory (Siegel, 2026). Unlike historical technological competitions, the AI race has no agreed endpoint, no universally accepted metric of success, and no consensus on whether achievement should be defined by artificial general intelligence (AGI), frontier-model capability, compute capacity, scientific productivity, economic competitiveness, military advantage, or global diffusion (Akula & Guest, 2026). Consequently, assertions that one state will ultimately win the AI race are conceptually incomplete. This working paper therefore shifts attention from identifying a presumed winner to examining an alternative question: whether strategic influence may depend not only on frontier capability, but also on how AI systems are adopted, trusted, and deployed across third-country markets. The following corpus outlines ten working parts.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.