Skip to content

Category

software testing

633 papers

#small language model Open access Aug 2026

Let’s read the log: root cause analysis of railway test execution logs with large language models

Results showed that long-context LLMs tended to achieve higher accuracy than smaller models, suggesting that LLMs are currently better suited to support human-in-the-loop root cause analysis than to fully automate it, and motivating further work to improve prediction accuracy for log-based RCA.

Rahmanu Hermawan, Alessio Bucaioni, Eduard Paul Enoiu et al. · 0 citations
#software testing Open access Aug 2026

Serum Albumin and Uric Acid Levels Among Hypertensive Patients in Northwest Ethiopia: A Comparative Cross‐Sectional Study

It is demonstrated that the prevalence of hypoalbuminemia and hyperuricemia was higher in hypertensive patients than the controls, and albuminuria was a significant predictor of hypoalbuminemia and hyperuricemia.

Arega Zenaw, B. Biadgo, Getnet Fetene et al. · 0 citations
#small language model Review Open access Feb 2026

An Ontology for Workplace Violence: Protocol to support Cross-sector Incident Reporting in Public Services.

BACKGROUND Underreporting is a defining problem in the registration of workplace violence across public services. This is reinforced by inconsistent, often poor incident descriptions and the absence of a clear, shared nomenclature. Without an ontology-backed structure, similar events are recorded differently across settings or remain uncodable and effectively invisible. Moreover, WPV is typically registered and studied within sector- and even organization-specific categories and local reporting logics rather than through a shared semantic framework, limiting comparability and cumulative understanding across sectors. To address this, we propose creating and publicizing a workplace violence ontology (WPV-ONTO) to unify the representation of WPV events across public services. OBJECTIVE The objective of this protocol is to describe and justify a research methodology for developing a cross-sector ontology and reference nomenclature for WPV in public services (WPV-ONTO), including the scoping review, expert-consensus, and evaluation procedures used to build it. Promoting openness and high standards during its creation and encouraging its uptake once available are broader project goals that this protocol is designed to support, rather than measurable objectives that this protocol itself tests. METHODS Using the Protégé ontology editor and the METHONTOLOGY ontology development life-cycle guidelines, we will create an ontology that captures the cross-sector WPV domain in Web Ontology Language (OWL). In order to find common WPV concepts, definitions, and synonyms, the modeling process will employ a methodologically defined, iterative workflow that combines (1) focused scoping searches (Web of Science/PubMed/APA PsycInfo/ERIC/Sociological Abstracts); (2) structured extraction from industry-standard incident reporting tools and code lists; and (3) expert consensus. Based on iterative rounds of scientific literature reviews and industry-standard code lists, a team of domain experts will use a hybrid top-down and bottom-up approach to define and identify key ideas and relationships. A small pilot with 6-8 frontline workers from two sectors will test the prototype reporting form against usability criteria before version 1.0 release. RESULTS The primary output will be a comprehensive, versioned WPV-ONTO accommodating key WPV concepts relevant to public services, augmented with synonyms, definitions, and references. WPV-ONTO will include an explicit hierarchical structure and relations supporting inheritance and compositional incident encoding. WPV-ONTO seeks to integrate the needs and conceptualizations of frontline workers, safety and aftercare professionals, sector policymakers, WPV researchers, and health/information systems experts. Recruitment of the expert panel is planned to begin in March 2027, and WPV-ONTO version 1.0 is expected at the end of the 24-month development period. CONCLUSIONS WPV-ONTO is expected to enable reasoning, inference, and consistent representation of relationships among WPV concepts for application in multiple contexts, including cross-sector reporting harmonization, software and API development, research data integration, and evaluation of prevention and aftercare initiatives. By providing a shared vocabulary and explicit structure, WPV-ONTO may also facilitate linkage with other relevant information systems, such as electronic medical records, and justice and police information systems, reducing methodological fragmentation and supporting coordinated cross-sector learning while retaining the contextual specificity necessary for public service environments. CLINICALTRIAL Not applicable.

I. Steenhout, Ronald Buyl · 0 citations
#software testing Open access Aug 2026

GIS-based multi-criteria landslide susceptibility assessment in the Northwestern Himalaya, Jammu and Kashmir, India

This study presents a landslide susceptibility assessment using multiple parameters by integrating remote sensing, GIS, and field observations in the Jammu and Kashmir region. Eight contributing variables were selected for landslide susceptibility analysis: Land Use and Land Cover (LULC), proximity to roads, streams, slope gradient, slope orientation (aspect), geology, geomorphology, and elevation. In addition, an extensive landslide inventory consisting of 669 landslide events was developed using the Field Landslide Inventory Mapping (FLIM) application, LISS-IV satellite data, and field observations, covering an area of 42,950.43 km². Landslide susceptibility mapping (LSM) was carried out using the Analytical Hierarchy Process (AHP) approach and validated with MaxEnt software and field-generated landslide data. The resulting landslide susceptibility map was classified into five categories: very high, high, medium, low, and very low susceptibility zones. Based on the AHP approach, these zones cover 3.65% (1,569.0483 km²), 24.43% (10,492.8912 km²), 51.56% (22,147.2369 km²), 18.81% (8,079.2199 km²), and 1.54% (662.0301 km²) of the study area, respectively. The weighted overlay and MaxEnt models proved effective for landslide vulnerability mapping, with MaxEnt achieving AUC values of 0.82 for training data and 0.807 for testing data, indicating good predictive performance. Field validation further showed that 87% of landslides occurred within high and very high susceptibility zones. The Jackknife test identified road proximity, slope, and stream proximity as the most significant independent variables influencing landslide occurrence. The generated landslide susceptibility map provides important insights for reducing landslide risk and serves as a valuable tool for infrastructure planning, community development, and disaster management in the region.

A. S. Jasrotia, Amit Sharma, I. C. Das et al. · 0 citations

Are we achieving what the algorithm tells us? Analysis of lumbar pedicle subtraction osteotomies with pre-bent rods.

Computer-assisted planning with patient-specific rods accurately reproduced the intended PSO and segmental lumbar correction but did not reliably predict global sagittal parameters, suggesting current planning tools might require refinement to improve the accuracy of predicted postoperative alignment using pre-bent rods.

Renzo A Laynes, Rafael Garcia de Oliveira, Kenneth T. Nguyen et al. · 0 citations
#software testing Open access Sep 2026

Spinal meningiomas: histopathological grading using a benchmark radiomics model with notes on disease control.

A benchmark radiomics model to preoperatively identify the histological grade of spinal meningiomas is constructed, suggesting a need to characterize the interplay between tumor grade and extent of resection as drivers of local disease control in SMs.

Adhith Palla, Nicolas K. Goff, Blake Perdikis et al. · 0 citations
#software testing Preprint Aug 2026

DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research

DSA is presented, an evidence-aware orchestration framework for multi-market stock research with large language model (LLM) agents that establishes implementation conformance for the tested software contracts, not superior report quality, forecasting accuracy, or investment returns.

Lingfeng Zhu, Yirui Shi · 0 citations
#software testing Preprint Aug 2026

Mutation Testing for Reproducibility Safeguards in Machine Learning Research Software: An Empirical Study

It is shown that, in this sample, existing validation workflows often did not detect the particular controlled reproducibility-relevant changes introduced by the study, motivating reproducibility-oriented mutation testing as a complementary way to assess whether research-software safeguards constrain experimentally important choices.

Ilya V. Shulepov · 0 citations
#software testing Preprint Aug 2026

Surf_2_Volume: a workflow for converting CIFTI parcellations to NIfTI volume space

Surf_2_Volume is presented, a workflow that combines Connectome Workbench, FreeSurfer, AFNI, AFNI, neuromaps, and Python image processing to convert cortical and subcortical CIFTI parcellations into Neuroimaging Informatics Technology Initiative (NIfTI) volumes.

Shu-Guang Yang, Zi-Yi Wang, Yue-Ru Shen et al. · 0 citations
#software testing Preprint Aug 2026

Algorithms for optimizing model-based incomplete block designs

This work proposes a model-based approach that optimizes model parameters, and evaluates first- and best-improvement algorithms, simulated annealing (SA), threshold accepting (TA), and two novel algorithms utilizing directional derivatives (dd) to guide exchanges.

Jonas Bjermo, Frank Miller · 0 citations
#software testing Preprint Aug 2026

An Empirical Evaluation of Using Large Language Models for Automated Model-Based Test Generation

This paper presents an empirical evaluation of Large Language Models (LLMs) for automated model-based test generation, compared with a state-of-the-art model-based testing tool (GraphWalker) and its built-in algorithms (random and quick random for edge and vertex coverage settings).

Hafize Sanli, Onur Kilinççeker, Cihat Çetinkaya · 0 citations
#software testing Review Aug 2026

"A Second Set of Eyes": The Process and Challenges of Software Documentation Review

The work identifies five distinct stages of the documentation review process: self review, technical review, editorial review, play testing, and post-publication feedback, and draws on practitioners with distinct expertise to address quality across content, presentation, and user experience.

Avinash Bhat, Ian Arawjo, Disha Shrivastava et al. · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.