Jul 2026· Revista de Estudos Interdisciplinares· Vol 25, pp. e3374· 0 citations
TL;DR
It is concluded that SIRSL is functionally viable and has the potential to make systematic reviews more integrated, organized, and traceable, however, further validation across different research fields and documentary datasets is still required.
Abstract
The expansion of scientific production and the fragmentation of tools used in Systematic Literature Reviews hinder the organization, traceability, and integration of methodological stages. This study aimed to develop and functionally validate the Intelligent System for Systematic Literature Reviews (SIRSL), based on Generative Artificial Intelligence. This applied technological research adopted a descriptive approach and was conducted through iterative stages involving requirements assessment, system modeling, implementation, testing, and refinement. The system was developed using React, TypeScript, Node.js, and Firebase, integrating the Gemini model. Functional validation was carried out by applying SIRSL to a systematic review associated with a doctoral research project within the PPGADT. The initial corpus comprised 470 records, of which 58 were excluded after duplicate identification and application of the selection criteria, leaving 412 documents. The results demonstrated that SIRSL integrated functionalities for developing search strategies, importing files, detecting duplicates, screening studies, extracting data, producing indicators, and generating reports. Artificial Intelligence was employed as an assistive resource, while methodological decisions remained under the researchers’ responsibility. It is concluded that SIRSL is functionally viable and has the potential to make systematic reviews more integrated, organized, and traceable. However, further validation across different research fields and documentary datasets is still required.
ABSTRACT Introduction Artificial intelligence (AI) is a branch of technology enabling machines to emulate complex human skills; it can also entail problem‐solving using bioinspired methods. It is used for automating systematic literature reviews (SLR), that is, defining a clinical question, locating relevant literature, preliminary screening, study evaluation, data extraction and analysis. Title and abstract screening is one of the most time‐consuming and error‐prone phases involved in developing a systematic review. While AI promises to expedite this process, adopting it faces challenges due to concerns about compatibility and transparency. This review aims to identify current evidence concerning AI use during preliminary SLR reference screening; it describes characteristics such as the different metrics used for reporting performance and how the different algorithms, pipelines, workflows or web applications are validated. AI resource users' reflections regarding SLR screening automation have also been summarized. Methods A scoping review was conducted following Joanna Briggs Institute's (JBI) methodology. Its objective was to identify existing evidence regarding the use of AI resources for title and abstract screening automation. Searches were limited to articles published between 2019 and 2026. The review included primary studies reporting the development, assessment, validation, or real‐world use of AI resources for screening automation, as well as systematic reviews and articles reporting experiences or recommendations for their use. Two types of data were extracted: (1) from primary studies—characteristics of AI resources and, where applicable, recommendations for their use; (2) from systematic reviews and experience‐based articles, recommendations for the use of AI resources. Results included frequency descriptions, tables, figures, and a decision flowchart reflecting the number of references and articles retrieved, excluded, or included in the final analysis. Results A total of 174 unique studies published between 2019 and 2026 were included in this scoping review. These were grouped into web applications (43%), model comparisons (32%), generative models (6%), pre‐trained models (3%) or pipelines/workflows (15%) used for title and abstract screening in systematic literature review (SLR). Most studies came from North America. Evaluating these tools often relied on retrospective comparisons with human reviewers' work (63%), sensitivity (n = 60), and specificity (n = 62) being the most reported metric for criterion assessment and Work Saved over Sampling (n = 28) being the most reported metric for assessing their utility. Considerations concerning AI resource use focused on the need for standardized evaluation metrics, stopping criteria, study design and the data sets used, resource characteristics facilitating usability, best practice and future research areas, with the persistence of the human component in the process (n = 26) being the most pressing recommendation. Conclusion The findings indicated substantial heterogeneity regarding the types of AI resources used, considerable variation concerning the metrics used for reporting performance, differences in how such metrics are defined and a clear need for standardizing reporting methods, study designs and related procedures. Although AI technologies will continue to evolve, maintaining a clear and consistent framework for interpreting research on AI resources for automating title and abstract screening can support understanding their level of maturity and facilitate informed decision‐making by users.
A. M. Barragán, Sara Elena Ortiz Bonett, Eliana-Isabel Rodríguez-Grande et al.· Cochrane evidence synthesis...· 0 citations
In this systematic literature review, we aimed to identify and thoroughly analyze the existing scientific knowledge related to a notably emerging field of generative AI. In strict adherence to the SPAR-4-SLR protocol, we focus on 1,104 peer-reviewed articles from the Scopus database, published between 2015 and 2025. Using bibliometric and thematic mapping methods, we address two main research questions: the first concerns the major publication trends in GenAI research, and the second deals with the intellectual structures and thematic domains shaping the field. Our results indicate an abrupt increase in scientific production since 2022, a consequence of the launch of models such as GPT, DALL·E, and Stable Diffusion, which are becoming increasingly powerful. We further identify five main research clusters: the technological foundations of generative AI, where researchers focus on building and utilizing LLMs and deep learning; professional and educational applications; ethical and governance issues; AI-assisted creativity; and user perceptions. Additionally, we find that higher education plays a significant role in the area, both in the application of ideas and the exploration of relevant questions. This review highlights the field's strong interdisciplinary character and, at the same time, reveals the current challenges that the sector faces. We outline a systematic research program to guide further studies of the implementation, impact, and problems of GenAI in enterprises and communities.
Majdouline Attaoui, Wissal Attaoui, Anas Moukrim et al.· International journal of mul...· 0 citations
In Ensuring fast and precise code review process has become one of the major problems of the modern DevOps world, where software development happens quickly for code reviews to be done manually. In this paper, we introduce a novel automated code review framework named AICR-DevOps, which utilizes a combination of rule-based static analysis, fine-tuned CodeBERT semantic classifier, and large language model using confidence-based aggregation. The framework architecture allows successfully combining the power of traditional program analysis approaches and artificial intelligence reasoning capabilities to achieve high review accuracy with minimized false positives. Together with GitHub Actions, the framework is capable of conducting pull request analysis instantly and constantly adapting to specific coding practices of projects by fine-tuning based on developer feedback using LoRA method. Experimentation with the proposed framework on 12,400 pull requests gathered from 40 Java and Python repositories reached 84.7% precision, 81.3% recall, and an F1-score of 82.9%, with review comments provided in 38 seconds on average. Moreover, a controlled study conducted with 48 developers showed that the acceptance rate of the developers to the framework reached 73.6%.
Nawnit Kumar, M. K. Shukla, Raushendra Kumar et al.· 2026 6th International Confe...· 0 citations
Generative artificial intelligence (GAI) is expanding from model-centered research into engineering and manufacturing activities, but its scope and maturity remain uneven. This PRISMA-guided bibliometric and abstract-level thematic review maps peer-reviewed industrial GAI research published from 2022 to 4 June 2026. Searches of Scopus, Web of Science, and the ACM Digital Library identified 492 records; 119 duplicates and 121 ineligible records were removed, leaving 252 studies. Keyword normalization, co-occurrence analysis, dominant and secondary thematic coding, and an abstract-reported evidence characterization were applied. The corpus shows two connected trajectories: engineering generation based on generative models for design, topology, materials, and electronics, and knowledge-intensive industrial intelligence based on large language models, retrieval-augmented generation, knowledge graphs, agents, and human–AI collaboration. Most studies report empirical or computational evaluation (72.2%), but 84.5% remain research-stage; only 0.8% indicate operational industrial evidence in their abstracts. The findings, therefore, distinguish publication activity from deployment maturity. Priority requirements for adoption include domain-grounded data, verification, manufacturability checks, traceability, cybersecurity, intellectual property protection, system integration, workforce preparation, and human accountability. This review contributes a reproducible cross-domain map, an overlap-aware synthesis, and stakeholder-specific guidance for trustworthy industrial GAI.
Requirement engineering is a foundational stage of the software development life cycle, and the accuracy with which
requirements are classified directly influences downstream design, testing, and cost estimation. Software requirements are
commonly separated into Functional Requirements (FR), which describe what a system must do, and Non-Functional
Requirements (NFR), which describe how well the system must do it, covering attributes such as performance, security, and
usability. When this separation is carried out manually, the process is slow, subjective, and prone to disagreement between
analysts, particularly as project size grows. This paper presents a Requirement Classification and Prioritization Tool that
combines Natural Language Processing (NLP) with machine learning to automate the FR/NFR decision. Requirement
statements are cleaned and normalised, then represented numerically through two complementary techniques: Term FrequencyInverse Document Frequency (TF-IDF) and contextual BERT embeddings. Four classifiers-Logistic Regression, Support Vector
Machine, Random Forest, and a
BERT-based model-are trained and benchmarked on a dataset of 6,086 labelled requirement statements using accuracy,
precision, recall, and
F1-score. A Gradio-based interface allows a requirement to be submitted and its predicted category, confidence score, and a short
explanation to be viewed immediately. The results indicate that transformer-based representations offer a modest but consistent
improvement in contextual understanding over TF-IDF, while classical classifiers remain competitive and considerably cheaper
to train.
Sai Sindhuja Bhukya, Naveen Kumar Nuthanapati· International Journal for Re...· 0 citations
Science, Technology and Innovation (STI) indicators are used to measure scientific capacity, technological development, and innovation performance. Collecting these indicators from the scientific literature is slow and difficult to scale. This study evaluates two large language models (LLMs)—Gemini and Claude—on automatic extraction and classification of STI indicators from five peer-reviewed articles. A Seed Prompt strategy—in which a model generates its own operational prompt from a set of predefined criteria—inspired by automatic prompt engineering (Zhou et al., 2022), was used. Model outputs were compared against a Gold Standard—a reference list of indicators manually identified by human experts—using a strict Exact Match criterion, in which an extracted indicator is accepted only if its normalized string is identical to one in the Gold Standard. Extraction F1 averaged 0.265 for Gemini and 0.356 for Claude, with high variation across documents (0.0–0.977). Mean Adjusted Rand Index (ARI) was 0.050 for Claude and 0.119 for Gemini, indicating near-random classification. Approximately 69% of indicators were left unclassified. Performance was higher for lexically standardized indicator categories and lower for categories with unstable definitions. These results show that current LLMs are not reliable for unsupervised STI indicator workflows. Hybrid human-AI approaches are recommended as a near-term alternative.