Skip to content

Category

software testing

592 papers

#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

Context: Manual qualitative data analysis is time-intensive and can compromise validity and replicability, affecting analysis design, implementation, and reporting. Large Language Models (LLMs) enable human-bot collaboration in Software Engineering (SE), but their potential for qualitative data analysis in SE remains largely unexplored. Objective: The objective of this study is to design and develop an LLM-based multi-agent system that synergizes human decision support with AI to automate various qualitative data analysis approaches. Methods: We used LLM-based multi-agents systems to assist the qualitative data analysis process, deploying 27 agents, each responsible for a specific task, such as text summarization, initial code generation, and extracting themes and patterns. Results: The main findings are: (1) the LLM-based multi-agent system accelerates the qualitative data analysis process, (2) the system effectively automates tasks such as text summarization, initial code generation, and theme extraction, and (3) the publicly accessible code facilitates validation and further evaluation. Conclusion: The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners. Future improvements focus on enhancing multilingual performance and integrating continuous expert feedback. The source code of proposed system and system details can be found here: https://github.com/GPT-Laboratory/Qualitative-Analysis-with-an-LLM-Based-Agentts

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 40 citations
#computer vision Apr 2026

Agentic Frameworks for Reasoning Tasks: An Empirical Study

Recent advances in agentic frameworks have enabled AI agents to perform complex reasoning and decision-making. However, evidence comparing their reasoning performance, efficiency, and practical suitability remains limited. To address this gap, we empirically evaluate 22 widely used agentic frameworks across three reasoning benchmarks: BBH, GSM8K, and ARC. The frameworks were selected from 1,200 GitHub repositories collected between January 2023 and July 2025 and organized into a taxonomy based on architectural design. We evaluated them under a unified setting, measuring reasoning accuracy, execution time, computational cost, and cross-benchmark consistency. Our results show that 19 of the 22 frameworks completed all three benchmarks. Among these, 12 showed stable performance, with mean accuracy of 74.6-75.9%, execution time of 4-6 seconds per task, and cost of 0.14-0.18 cents per task. Poorer results were mainly caused by orchestration problems rather than reasoning limits. For example, Camel failed to complete BBH after 11 days because of uncontrolled context growth, while Upsonic consumed USD 1,434 in one day because repeated extraction failures triggered costly retries. AutoGen and Mastra also exhausted API quotas through iterative interactions that increased prompt length without improving results. We also found a sharp drop in mathematical reasoning. Mean accuracy on GSM8K was 44.35%, compared with 89.80% on BBH and 89.56% on ARC. Overall, this study provides the first large-scale empirical comparison of agentic frameworks for reasoning-intensive software engineering tasks and shows that framework selection should prioritize orchestration quality, especially memory control, failure handling, and cost management.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 1 citation
#computer vision Open access Nov 2023

Autonomous Agents in Software Development: A Vision Paper

Large Language Models (LLM) and Generative Pre-trained Transformers (GPT), are reshaping the field of Software Engineering (SE). They enable innovative methods for executing many software engineering tasks, including automated code generation, debugging, maintenance, etc. However, only a limited number of existing works have thoroughly explored the potential of GPT agents in SE. This vision paper inquires about the role of GPT-based agents in SE. Our vision is to leverage the capabilities of multiple GPT agents to contribute to SE tasks and to propose an initial road map for future work. We argue that multiple GPT agents can perform creative and demanding tasks far beyond coding and debugging. GPT agents can also do project planning, requirements engineering, and software design. These can be done through high-level descriptions given by the human developer. We have shown in our initial experimental analysis for simple software (e.g., Snake Game, Tic-Tac-Toe, Notepad) that multiple GPT agents can produce high-quality code and document it carefully. We argue that it shows a promise of unforeseen efficiency and will dramatically reduce lead-times. To this end, we intend to expand our efforts to understand how we can scale these autonomous capabilities further.

Z. Rasheed, Muhammad Waseem, Kai-Kristian Kemell et al. · 35 citations · ⚡2
#artificial intelligence Book Open access Nov 2023

Examining Privacy and Trust Issues at the Edge of Isomorphic IoT Architectures: Case Liquid AI

The growing domain of liquidity in computing extends its boundaries to include advancements like liquid artificial intelligence (AI). Liquid AI leverages liquid software using isomorphic Internet of Things (IoT) architecture to enhance computation at the edge. This innovation unveils vast opportunities yet also introduces significant challenges, particularly around privacy and trust. We explore the vulnerabilities that might hinder the progression of this technological fusion toward achieving trustworthy AI. Through an intensive examination of the literature, this research highlights the heightened threats to data integrity and stakeholder trust in these evolving ecosystems. Four main challenges: Data collection, Data storage and Access, Data utilization and sharing, and Surveillance and profiling were identified and examined under privacy, and two, Algorithms and decision-making and Security of IoT infrastructure under trust. The concerns are further categorized to highlight their impact on the development of trustworthy AI. The study acknowledges the early state of the field. Consequently, this research navigates through the limited available literature, initiating a pioneering discourse emphasizing fostering a foundation for developing secure and trustworthy Liquid AI environments.

M. Agbese, Niko Mäkitalo, Muhammad Waseem et al. · 6 citations · ⚡1
#computer vision Conference Open access Feb 2026

Carbon-Aware Governance Gates: An Architecture for Sustainable GenAI Development

The rapid adoption of Generative AI (GenAI) in the software development life cycle (SDLC) increases computational demand, which can raise the carbon footprint of development activities. At the same time, organizations are increasingly embedding governance mechanisms into GenAI-assisted development to support trust, transparency, and accountability. However, these governance mechanisms introduce additional computational workloads, including repeated inference, regeneration cycles, and expanded validation pipelines, increasing energy use and the carbon footprint of GenAI-assisted development. This paper proposes Carbon-Aware Governance Gates (CAGG), an architectural extension that embeds carbon budgets, energy provenance, and sustainability-aware validation orchestration into human-AI governance layers. CAGG comprises three components: (i) an Energy and Carbon Provenance Ledger, (ii) a Carbon Budget Manager, and (iii) a Green Validation Orchestrator, operationalized through governance policies and reusable design patterns.

M. Abbasi, T. Mikkonen, Petri Ihantola et al. · 0 citations
#natural language process... Review Open access 2026

Fabrication of hollow fiber membranes via NIPS spinning system for CO2 capture

Abstract. Carbon dioxide (CO2) emissions from industrial activities remain one of the greatest contributors to global climate change. Hollow fiber membranes (HFMs) have emerged as a promising technology for post-combustion CO2 separation owing to their high surface-area-to-volume ratio and scalability. This work focuses on the fabrication of HFMs with an emphasis on gas separation, particularly CO2, using the non-solvent induced phase separation (NIPS) spinning process for HFMs fabrication. The process allows specific control over dope and bore fluid selection, and flowrates, enabling the formation of asymmetric structures with desirable porosity, mechanical strength and suitable morphology for gas separation. The fabrication of polyethersulfone (PES)-based HFMs via NIPS, with 3 wt% polyethylene glycol (PEG) as a pore-forming additive, served as a foundational and basic study framework to provide an overview of the general hollow fibre membrane fabrication process. Preliminary assessments demonstrated the suitability of the fabricated membranes for gas separation applications as per requirements of membrane-based carbon capture technologies. Scanning electron microscopy (SEM), gas permeability tests, and tensile testing all revealed improvements in morphology, porosity, and mechanical strength, implying that this method for fabricating hollow fibre membranes has significant potential for tuning hollow fibre membranes for gas separation applications. Finally, the potential of HFM-based systems for energy-efficient CO2 capture is highlighted to be explored further.

Muhammad Waseem · 0 citations
#natural language process... Preprint Aug 2026

Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

Coding agents are increasingly used for software engineering tasks, including bootstrapping projects from third-party repositories whose integrity cannot be assumed. Prior work on repository poisoning largely focuses on attacker-controlled injection and disguise, but developers also shape risk through everyday invocation choices: what task to delegate, how to phrase the request, and which skills or rules to supply. We term these user-side choices Prompt-Level Configurations (PLCs) and introduce CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories. CIPR comprises 1,920 instances across 20 repositories, four task types, three social-media-grounded prompt styles, and three skill/rule conditions, and measures attack success rate (ASR) and agent alert rate (AR) using automated runtime and trace-based oracles. Our evaluation reveals two key insights: (1) Vulnerability is highly context-dependent, with task type creating up to a 4.5-fold difference in ASR, with test-execution task forming a silent attack surface (high ASR, low AR). (2) Prompt expression shifts risk indirectly: underspecified prompts reduce ASR by truncating execution depth; noisy prompts exhibit a directional trend toward suppressing alerts by making malicious content less conspicuous. These findings highlight that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.

Fu-Kang Zhu, Binbin Zhao, Ruixiao Lin et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Automated Testing of LLM-Based Post Hoc Explainers Using Model Checking as an Oracle

Large language models (LLMs) are used as post hoc explainers of sequential decision-making policies, producing natural-language explanations of why an action was chosen. However, LLMs often generate plausible but incorrect statements, and no existing approach systematically tests whether such explanations are faithful to the underlying environment. Two classic software testing challenges stand in the way: there is no oracle for the correctness of an explanation, and the test inputs, natural language queries about a policy's behavior, lack the structure needed for systematic test case generation. We address both. Probabilistic model checking provides the test oracle, computing exact reference results against which LLM answers are graded automatically. A taxonomy of post hoc query categories structures the input space around the environment-level facts from which policy explanations are composed; test cases generated from it are prioritized by question-specific diagnostic difficulty scores. Across seven MDP environments, the testing separates three open-weight LLMs: a reasoning model passes 85% of test cases, a mid-size model 70%, and a 1B model falls below the random baseline, while prioritization surfaces significantly harder cases than random selection. Our results indicate how trustworthy LLM-generated explanations are in model-free settings, where the same LLMs are used but no oracle exists to verify them.

Dennis Gross, Helge Spieker · 0 citations
#machine learning Preprint Aug 2026

Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation

A company with a fixed artificial intelligence (AI) budget must decide which large language model (LLM) handles each recurring workload. What it lacks is the quality table, how well each model performs on each workload. Given that table, the decision is a multiple-choice knapsack problem and is routine to solve, so estimating it is the difficulty, and that estimation fails in two ways. Models are rarely compared on the same work, and the recorded score is usually a proxy rather than the outcome the company values. Causal and off-policy methods repair the first but condition on the second, while evaluator-validation methods estimate the second but stop short of the decision. Worse, buying more re-evaluation cannot settle the second: randomization governs which requests are scored, not how a score is produced, so the table stays uncertain however much evaluation is purchased. Yet the deployment decision may still be determined even when the table is not. We therefore ask whether one assignment stays optimal across every quality table consistent with the evidence. For the fixed-budget problem, this admits an exact two-solve certificate: solve once at the estimated table and once at a least-favourable table. Agreement certifies the assignment; disagreement identifies the model-workload pairs where further evidence can matter. We propose CASE (causal active sequential experimentation), which targets evaluation to those pairs and repeats the test as evidence accumulates. On a production log, the measurement failure is the larger of the two: correcting assignment exactly still leaves most of the loss, and randomized re-evaluation does not remove it. In our experiments, the available evidence often does not determine the assignment. On paid software tasks, better information about model quality yields more savings than further optimization of the assignment on the same estimates.

H. Khosravi, Xiao-Ming Huo · 0 citations
#software testing Open access Sep 2026

THE IMPACT OF PREOPERATIVE ESOPHAGEAL MANOMETRIC STUDY ON ANTI-REFLUX SURGICAL DECISION: A RETROSPECTIVE STUDY FROM JORDAN

Introduction: Esophageal manometry is an endoscopic study of the esophagus motility and can diagnose dysmotility disorders and the tone of the gastro-esophageal sphincter. The main aim of the study is to identify the role of esophageal manometry in surgery decision-making and to determine which type of fundoplication is required for gastric reflux disease. Methodology: it’s a retrospective of 210 patients with gastroesophageal reflux disease, who underwent endoscopic esophageal manometry in order to undergo anti-reflux surgery. The study was conducted at King Hussein Medical Hospital. We collected data retrospectively from the electronic medical recording system regarding all patients who performed the test from August 2022 to July 2025. All patients who were planned to undergo antireflux surgery and underwent esophageal manometry as a preoperative preparation were included in this study, whereas the patients who underwent the test because of dysphagia or esophageal dysmotility were excluded from the study. Then, we followed the surgical decision of antireflux surgery to do Nissen or partial Fundoplication accordingly. Medical biostatistical tests and software like SPSS are used to collect data and to measure the results. Results: A total of 210 cases, that underwent preoperative preparations and planned for anti-reflux surgery. The mean age was 49+-0.3 years. A 46 % (n=99) where female patients with slightly higher percentage in male patient (54%, n=111). Manometric study performed for all patient (100%), abnormal study observed in 15 patients (7.2%) that showed significant difference and P-value was than 0.03, most of these dysmotility were a chalasia 10 patients, while the other 5 patients suffered from non-specific abnormal esophageal contractility. Conclusions: esophageal manometry is pivotal preparation test must be performed to delineate the esophageal motility before undergoing anti-reflux procedure, it may change the type of surgical procedure and the degree of fundoplication

Qasem Alqaisi MD*, Mutaz Haddadin MD, Rami Al Omoor MD, Gayth Arabyat MD, Abdelrazzaq Abuabboud MD, Ghidaa Maswadah · 0 citations
#software testing Open access Sep 2026

EVALUATION OF ANESTHETIC DRUG USE IN OPHTHALMIC PROCEDURES AT PRINCE ZAID BIN AL HUSSEIN HOSPITAL

Introduction: While topical, local, and regional anesthesia are commonly used in ophthalmic surgery, general anesthesia may be necessary for children, uncooperative patients, or complex situations, selecting the incorrect anesthetic can result in systemic or ocular side effects, particularly in older patients with concurrent conditions. Drug utilization research is essential for rational medicine use in anesthetic departments, this study sought to evaluate the use of anesthetic drugs in ocular procedures at Prince Zaid Bin Al Hussein Hospital in order to maximize practice, minimize unnecessary consumption and improve patient safety. Methodology: All patients undergoing eye surgery who got any kind of anesthesia WERE included in this retrospective observational study, which was carried out, Data will be gathered from pharmacy databases and operating room records, Variables include procedure type, and specifics about anesthetic medications, including class and delivery method. The data will be analyzed using statistical software, the results will be summarized using descriptive statistics, and inferential tests, the outcomes will be compared with accepted clinical criteria. Result: the total annual consumption of topical ophthalmic anesthetic preparations was 820 minims. Lignocaine with fluorescein represented the largest share with 440 minims (53.7%), followed by Benoxinate with 380 minims (46.3%). During the same period, 190 ophthalmic procedures were recorded, mainly Phaco procedures, which accounted for 145 cases (76.3%). The calculated local consumption rate was approximately 4.32 anesthetic minims per ophthalmic procedure. Conclusion: Lignocaine with fluorescein and benoxinate was the most commonly used ocular anesthetic, this pattern is generally in line with international practice. This study offers a drug-utilization perspective, in contrast to the majority of international studies that concentrate on pain management and therapeutic efficacy. It connects the quantity and kind of ophthalmic procedures with the yearly administration of anesthetics.

Osama E. Al Bdairat, MD*1, Amjad T. Z. Alhamadin, MD2, Saif Addeen T. Al Bdairat, MD3, Albdl-Motaleb M. Al Shra'a, MD4, Marwan H. Alzoubi, MD5, Mohammad E. Al Bdairat, Pharm-D6 · 0 citations
#software testing Open access Sep 2026

EVALUATION OF ANESTHETIC DRUG USE IN OPHTHALMIC PROCEDURES AT PRINCE ZAID BIN AL HUSSEIN HOSPITAL

Introduction: While topical, local, and regional anesthesia are commonly used in ophthalmic surgery, general anesthesia may be necessary for children, uncooperative patients, or complex situations, selecting the incorrect anesthetic can result in systemic or ocular side effects, particularly in older patients with concurrent conditions. Drug utilization research is essential for rational medicine use in anesthetic departments, this study sought to evaluate the use of anesthetic drugs in ocular procedures at Prince Zaid Bin Al Hussein Hospital in order to maximize practice, minimize unnecessary consumption and improve patient safety. Methodology: All patients undergoing eye surgery who got any kind of anesthesia WERE included in this retrospective observational study, which was carried out, Data will be gathered from pharmacy databases and operating room records, Variables include procedure type, and specifics about anesthetic medications, including class and delivery method. The data will be analyzed using statistical software, the results will be summarized using descriptive statistics, and inferential tests, the outcomes will be compared with accepted clinical criteria. Result: the total annual consumption of topical ophthalmic anesthetic preparations was 820 minims. Lignocaine with fluorescein represented the largest share with 440 minims (53.7%), followed by Benoxinate with 380 minims (46.3%). During the same period, 190 ophthalmic procedures were recorded, mainly Phaco procedures, which accounted for 145 cases (76.3%). The calculated local consumption rate was approximately 4.32 anesthetic minims per ophthalmic procedure. Conclusion: Lignocaine with fluorescein and benoxinate was the most commonly used ocular anesthetic, this pattern is generally in line with international practice. This study offers a drug-utilization perspective, in contrast to the majority of international studies that concentrate on pain management and therapeutic efficacy. It connects the quantity and kind of ophthalmic procedures with the yearly administration of anesthetics.

Osama E. Al Bdairat, MD*1, Amjad T. Z. Alhamadin, MD2, Saif Addeen T. Al Bdairat, MD3, Albdl-Motaleb M. Al Shra'a, MD4, Marwan H. Alzoubi, MD5, Mohammad E. Al Bdairat, Pharm-D6 · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.