A Systematic Mapping Study on the quality of AI-based software identifies six recurring challenge categories, with the most prominent being limitations in existing quality assessment models followed by issues in non-functional requirement management, quality-aware development, and quality assurance.
Abstract
Artificial Intelligence (AI) is increasingly embedded in modern software systems, raising important questions about how its quality should be defined, assessed, and assured. This paper presents a Systematic Mapping Study (SMS) on the quality of AI-based software. The study synthesizes primary studies published between January 2020 and January 2026 and selected from five electronic data sources. A total of 33 primary studies were included after automated search, screening, and snowballing. The results identify six recurring challenge categories, with the most prominent being limitations in existing quality assessment models, followed by issues in non-functional requirement management, quality-aware development, and quality assurance. The findings suggest a call for collaboration of researchers and industrial practitioners with standardization organizations, that could possibly devise comprehensive quality assessments and their measurement methods.
An enormous course in scholarly output from 2023 onwards is revealed by the findings, and this growth is driven by the industrial adoption of Large Language Models alongside autonomous agentic systems.
Abdullah A. H. Alzahrani· International Journal of Adv...· 0 citations
The growing complexity of software systems has increased demand for intelligent solutions in Software Quality Assurance (SQA). Agentic Artificial Intelligence (AI) offers autonomous decision-making, adaptive behavior, and proactive quality improvement, yet its applications across SQA remain fragmented. This systematic mapping study examines Agentic AI integration in SQA, analyzing application distribution, agent characteristics, and supported stakeholder roles to identify coverage, trends, benefits, limitations, and gaps. Following systematic mapping protocols and PRISMA, five digital libraries were searched, retrieving 37 primary studies. Data were classified through frequency analysis and qualitative categorization aligned with Agentic AI behavior, architecture, and autonomy dimensions. Applications concentrate in Product Assurance, mainly test generation and defect management, with fewer studies in Process or Planning activities. Most agents exhibit goal-oriented or diagnostic behaviors, moderate autonomy, and hybrid or multi-agent architectures. Benefits include efficiency, accuracy, and maintainability improvements; limitations involve scalability, generalizability, and transparency. Role coverage focuses on execution-level stakeholders, while managerial and leadership roles receive minimal support. Agentic AI shows potential for strengthening SQA intelligence and adaptability but remains unevenly distributed across activities and roles. Transparency, validation, and leadership support are priority research gaps for developing trustworthy, process-aware Agentic systems across the full SQA lifecycle.
The literature on vibe coding has grown rapidly; however, it remains fragmented and is largely dominated by industry reports, leaving its position relative to traditional manual programming and low-code development insufficiently examined. This gap makes it difficult for both researchers and practitioners to determine when vibe coding is appropriate and what risks should be anticipated. Purpose: This study aims to systematically map the current landscape of vibe coding, develop a comparative framework against manual and low-code software development approaches, and propose practical risk mitigation recommendations for software development practitioners. Methodology: A Systematic Literature Review (SLR) was conducted following the PRISMA protocol. Relevant publications from 2023 to 2026 were retrieved from IEEE Xplore, ACM Digital Library, Springer, ScienceDirect, and arXiv, resulting in 61 studies that were analyzed using thematic analysis. Findings: The results indicate that vibe coding can accelerate software prototyping by approximately 40–60% compared with manual development. However, it introduces a verification bottleneck by shifting developers' workload from code implementation to quality assurance and validation. Compared with low-code development, vibe coding provides greater flexibility in expressing user intent but exhibits lower output predictability. In comparison with manual development, it offers significant gains in development speed while sacrificing architectural control and code security, thereby increasing the risks of technical skill degradation, hidden security vulnerabilities, and accumulated technical debt. Implications: The findings provide practical guidance for software development teams in identifying project phases that are suitable for extensive adoption of vibe coding and those that still require manual architectural review. The study also emphasizes the importance of integrating security auditing and technical debt monitoring into AI-assisted software development workflows. Originality/Value: The novelty of this study lies in its explicit comparative framework, which systematically positions vibe coding alongside manual and low-code development across six technical dimensions, extending previous studies that have generally examined vibe coding in isolation.
A. Jauhari, Fahmi Fathullah, I. Permana et al.· International Journal for Sc...· 0 citations
The findings indicate that techniques such as artificial neural networks, optimization algorithms, machine learning models, and hybrid approaches consistently yield improvements in estimation accuracy, with average error reductions reported in the literature ranging approximately from 15% to 30% when compared with traditional methods.
Rodolfo Barbosa dos Santos, L. E. G. Martins· Journal of Software: Evoluti...· 0 citations
Artificial intelligence-based assistants have been developing at an incredible pace in recent years, facilitating fast information retrieval, automating repetitive tasks, and improving operational efficiency. In the field of software engineering, these tools now assist in activities such as testing, code generation, and documentation. Despite the increasing recognition of artificial intelligence (AI) as a valuable element of software development, its overall contributions remain insufficiently characterized. This research conducts a bibliometric analysis of AI-based assistants in the context of software development, identifying key publications, leading authors and organizations, and underexplored areas where these tools have a significant impact. As a result of a literature search conducted using the Scopus database, 83 papers (2014–2025) were analyzed using the Bibliometrix R-package for bibliometric evaluation. The collected documents revealed a sustained annual growth rate of 24.14%, with a peak in 2024 (31 papers), reflecting the surge in generative AI and LLM-based tools. The most cited article received 118 citations (39.33 citations per year), highlighting a strong impact in recent contributions. The United States led in publications with 26.1%, while Europe had the highest citation impact with 73 citations. Software design is the dominant theme with 41 papers and 19% of occurrences, while keyword trends focus on language models (LLM) with 19 papers, chatbots with 18 papers, and LLMs with 11 papers, confirming a clear shift toward LLM-centered research in software engineering. It provides insights into future research directions and opportunities for AI-driven innovation in software development .
Alvaro Fernández Del Carpio, Leonardo Bermón Angarita, Andrés Alberto Osorio Londoño· Informatica· 0 citations
Artificial intelligence (AI) is transforming requirements elicitation: machine learning, natural language processing (NLP), and large language models (LLMs) now identify software requirements automatically from the textual data that surrounds every project—user feedback, specifications, regulations, and stakeholder transcripts. This paper presents a systematic mapping study of 74 peer-reviewed primary studies on AI-based automated requirements elicitation published between 2021 and 2025, identified from five databases following PRISMA 2020 and classified by AI technique, textual source, elicitation activity, and application domain. The evidence is divided into two equally sized source families—user feedback and agile artefacts versus formal documentation—each coupled to the AI techniques that suit its signal profile. Fine-tuned transformer encoders set the performance ceiling and, task-for-task, still outperform far larger generative models, while LLMs extend elicitation to long regulatory documents, multilingual feedback, and structured outputs. The central finding concerns automation depth. AI identifies requirements with consistently high accuracy (routinely F1 0.8 and above), but automation thins at every subsequent step: 51% of approaches structure what they identify, 23% consolidate them, and only 8% engineer stakeholder validation into the loop. This leaves the steps that turn candidates into agreed requirements largely manual. Benchmark fragmentation (77% custom datasets), thin industrial validation (14%), and skewed non-functional coverage compound this gap. The resulting map gives researchers an evidence-derived agenda for deepening automation, and practitioners guidance on which techniques the evidence supports for each elicitation task and textual source.
Safaa Eltahier, S. Al-Ghuribi, Mawal A. Mohammed et al.· Information· 0 citations