Skip to content

Category

large language models

451 papers

#large language models Open access Sep 2026

Integrating large language models for automated structural analysis

Automated analysis for engineering structures offers considerable potential for boosting efficiency by minimizing repetitive tasks. Although AI-driven methods are increasingly common, no systematic framework yet leverages Large Language Models (LLMs) for automatic structural analysis. This paper proposes a framework that employs domain-specific prompt design and in-context learning strategies to enhance LLM problem-solving capabilities and generative stability, enabling fully automated structural analysis from descriptive text to model outputs. A small-scale benchmark dataset consisting of 20 structural analysis word problems (SAWPs) is also introduced to evaluate the performance of different LLMs within the proposed framework. The results demonstrate that the proposed approach can increase the level of automation in solving SAWPs compared with traditional methods. Quantitatively, the framework built on GPT-5.4 and GPT-4o both achieved 100% accuracy, outperforming GPT-4 (85%), Gemini 1.5 Pro (80%), and Llama-3.3 (30%) on the test examples. Furthermore, integrating domain-specific instructions enhanced performance by 30% on problems with asymmetrical structural configurations.

Haoran Liang, Mohammad Talebi Kalaleh, Qipei Mei · 1 citation
#large language models Open access Sep 2026

Toward scalable generative AI: efficient language model distillation via zero-shot rationales

Abstract This paper investigates an efficient approach for distilling Large Language Models (LLMs) into smaller, application-specific models using zero-shot Chain of Thought (CoT) rationale generation and Optimization by Prompting (OPRO). To address the challenges of deploying computationally intensive generative AI for narrow tasks or resource-constrained environments, the approach leverages LLM reasoning capabilities to generate both labels and natural language explanations for unlabeled data. By reducing reliance on human-generated annotations, the approach substantially lowers annotation requirements and prompting costs while maintaining comparable performance in the evaluated settings. We formulate distillation as a multi-task learning problem in which student models are trained to jointly predict labels and learn from teacher-generated rationales, with the goal of improving data efficiency and generalization. Building on established zero-shot Chain of Thought (CoT) prompting and the OPRO prompt optimization technique, we use teacher-generated rationales to reduce annotation token requirements and examine the associated performance and efficiency gains. Additionally, we systematically investigate how explanation properties affect distillation efficiency. Across natural language inference and question answering benchmarks, results indicate that near-optimal performance can be achieved even when rationales are provided for only a subset of the training data, and that shorter explanations are often sufficient. These findings provide practical insights into the trade-offs between rationale generation cost and student model performance. Overall, this work contributes empirical evidence on the effectiveness and cost characteristics of rationale-based distillation for training compact, task-specific language models with minimal human intervention.

Lukas Vöge, Vincent Gurgul, Stefan Lessmann · 0 citations
#large language models Open access Sep 2026

Don’t Look up: Evaluating the Tradeoff Between Performance and Sustainability of Text Classification Using Open LLMs

The increasing adoption of Large Language Models (LLMs) as a text analysis method in social science presents a critical yet under-examined trade-off between model performance and environmental sustainability. This research provides a systematic evaluation comparing the performance, energy consumption, processing time, and CO 2 emissions of various computational text analysis methods (CTAM), including dictionaries, trained classifiers, and self-hosted open LLMs when performing sentiment analysis of parliamentary speeches, classification of open-ended survey responses, and named entity recognition of newspapers. The analysis is limited to self-hosted deployment in local and server environments where per-task energy consumption is directly measurable. Although self-hosted LLMs demonstrate strong performance in sentiment analysis, closely aligning with human judgment, they require significantly more energy and time than non-LLM approaches. For classification and named entity recognition, pretrained task-specific models achieve better F1 scores with a lower carbon footprint, challenging the primacy of larger models. To navigate this trade-off, we propose a CO 2 -Adjusted F1 Score that penalizes emissions while rewarding performance. Applying this metric, we show that smaller, task-specific models may be preferred over larger general-purpose LLMs for efficient text analysis. We highlight the necessity for thoughtful and responsible model selection, promoting a “right-fit” approach for CTAM.

Sean-Kelly Palicki, Isaac Bravo, Clint Claessen · 0 citations
#large language models Editorial Open access Sep 2026

Editorial: Advances in neurocritical care

Neurocritical care focuses on patients with actual or impending organ dysfunction that arises from, or accompanies, primary or secondary neurological injury and necessitates intensive care (1). As a subspecialty, neurocritical care continues to face unresolved and controversial questions in both clinical practice and research, many of which require further investigation. Chen et al. provide a broad overview of recent developments in severe nervous systematic diseases, neuromonitoring, hemodynamic and respiratory support, and post-cardiac-arrest car (1). Within this broader landscape, the present Research Topic, Advances in Neurocritical Care, examines more focused clinical, physiological, diagnostic, and methodological questions across these domains. It comprises 29 contributions: 15 research articles (14 Original Research articles and one Clinical Trial), three Systematic Reviews, three Reviews, two Study Protocols, and six Case Reports. The contributions are discussed here under four overlapping themes: dynamic physiological monitoring; biomarkers, measurement, and prediction; perioperative and neurocritical care management; and systemic, metabolic, or treatment-related neurological injury.Continuous monitoring derives clinical value from its integration with bedside assessment. Serial clinical examination remains indispensable, with trajectories serving as one component of repeated neurological assessment and contextual interpretation. Automated pupillometry has gained wider use because it provides an objective, repeatable measure of pupillary reactivity. Chen et al. highlighted the ORANGE cohort, in which abnormal Neurological Pupil Index (NPi) measurements were associated with unfavorable neurological outcomes and higher mortality after severe non-anoxic acute brain injur (1). The 2025 B-ICONIC Brussels consensus proposed an NPi of 3 or lower as suggestive of intracranial hypertension in traumatic brain injury (TBI) when invasive intracranial pressure (ICP) monitoring is unavailable and recommended integrating NPi with the clinical examination and at least one additional non-invasive modality (2). Kim et al. followed serial NPi values during the first 72 h after non-traumatic subarachnoid hemorrhage and found lower early NPi values in patients with unfavorable outcomes. These findings support serial assessment, although an intervention threshold remains to be defined. A low or falling NPi warrants repeat examination and review of imaging, physiology, and potential confounders before a protocolized response is initiated.Cerebral perfusion pressure (CPP) trajectories raise a related question. Wang et al. identified four phenotypes among 1,466 patients and found the highest mortality in the rapidly declining group. The gain in discrimination was modest. Earlier CENTER-TBI work linked time below an individualized lower limit of reactivity with mortality, while COGiTATE showed that autoregulation-guided CPP targeting was feasible and safe (3,4). Wang et al.'s study broadens the population beyond TBI. Its contribution lies in identifying how often and in whom CPP deteriorates. The appropriate bedside response to that pattern remains to be established prospectively.Positioning after craniotomy and hypotension during non-cardiac surgery may appear to be separate issues, yet both studies show why a threshold stripped of time and context is incomplete. In 21 postoperative patients, Li et al. found that a flat head-of-bed position increased ICP and reduced CPP, while tissue oxygenation remained stable. Baseline pressure and autoregulatory status altered the response. The brief physiological observation leaves the effect of position on recovery unresolved and cautions against assuming a uniform response. Ren et al. analyzed minute-by-minute arterial pressure in 789 operations. Complications increased as hypotension became sustained, prolonged, or fluctuating. Here, the clinically relevant exposure was the temporal pattern of hypotension.Signal quality is the central problem in two other monitoring studies. Hinsberger et al. traced most difficulties in electroencephalographic assessment during brain death determination to technical artifacts, especially electrode-related artifacts. The contribution of electroencephalography depends on compliance with technical standards and the applicable jurisdictional protocol. Jiang et al. evaluated diaphragm ultrasound in 188 neurosurgical intensive care unit (ICU) patients. First spontaneous breathing trial (SBT) success and first extubation success were similar across study phases, although ultrasound-guided phases had fewer reintubations and shorter ventilation. The sequential design makes the size of that benefit uncertain. The larger issue, from our perspective, is conceptual: diaphragm thickening fraction captures respiratory muscle performance, whereas airway protection depends on a broader set of functions. Current consensus treats extubation after acute brain injury as more than a respiratory test (5). Xu et al. had already operationalized that distinction in 226 neurosurgical patients: their STAGE score combined swallowing, tongue protrusion, spontaneous and suctioning cough, and the Glasgow Coma Scale motor response, and showed moderate discrimination for extubation success, with an area under the curve (AUC) of 0.72 (6). Badenes et al. later framed the same shift from respiratory load to airway protection, noting that, once an SBT is passed, vigorous cough may matter more than the SBT modality (7). We would therefore integrate diaphragm ultrasound with assessment of cough, secretion burden, consciousness, and bulbar function. Passing an SBT demonstrates short-term unsupported breathing capacity; airway safety requires separate evaluation.The Chinese neonatal extracorporeal membrane oxygenation (ECMO) consensus has a different purpose. Lu et al. describe how a Chinese expert consensus on neurological monitoring and long-term follow-up will be developed using a systematic review and the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach. We welcome its inclusion of neurodevelopment after discharge; survival alone is an incomplete outcome in this population. The protocol describes the methods, while the clinical recommendations and their feasibility await completion of the consensus process.Wang et al. studied the cerebrospinal fluid glucose-to-lactate ratio in 121 postoperative patients with acute brain injury and suspected intracranial infection. The area under the receiver operating characteristic curve was 0.866, with similar performance across glycemic strata. Because both measurements are routinely available, the ratio has practical appeal. We are less interested in the reported decimal than in whether the cut point survives changes in prior antibiotics, sampling time, case mix, and the reference diagnosis. Those details decide whether the test can travel.Three systemic biomarker studies remain further from a treatment decision. In 5,267 critically ill patients with stroke from the Medical Information Mart for Intensive Care IV (MIMIC-IV), Wang et al. found a graded relationship between the leuko-glycemic index and mortality and reproduced it in an institutional cohort of 424 patients. Another team serially measured soluble triggering receptor expressed on myeloid cells 1 and 2 (sTREM-1 and sTREM-2) in 120 patients after cardiac arrest and incorporated the measurements into machine-learning models. Kashatnikova et al. examined T-cell receptor excision circles and B-cell K-deleting recombination excision circles during rehabilitation after TBI; the recruitment setting leaves the acute phase and patients who never reached rehabilitation outside the frame. We read all three as candidate biological phenotypes.The two clinical prediction tools use familiar variables to estimate postoperative delirium after traumatic cervical spinal cord surgery and pulmonary infection after cerebral hemorrhage. Their simplicity is attractive, but delirium screening, extubation practice, antibiotic use, and definitions of pneumonia differ between hospitals. Zhang et al.'s Consensus-based Standards for the Selection of Health Measurement Instruments (COSMIN) review of the revised Richards-Campbell Sleep Questionnaire points to an even earlier problem: several language versions lack adequate evidence for measurement properties beyond internal consistency. A cleaner algorithm cannot rescue an unstable label. Before adding predictors, investigators need to show that the same outcome is being measured at the same time in the same way.Fluid choice in neurocritical care requires condition-specific interpretation of evidence derived from general ICU populations. In a secondary analysis of the BaSICS randomized trial, patients with TBI assigned to Plasma-Lyte 148 had a high probability of greater 90-day mortality than those assigned to saline; the subgroup design limits any class-wide inference about balanced solutions (8). The 2018 European Society of Intensive Care Medicine (ESICM) consensus recommended isotonic crystalloids for maintenance and resuscitation in acute brain injury and advised against hypotonic solutions; albumin was not recommended in TBI (9). The 2024 ESICM clinical practice guideline similarly emphasized tonicity and condition-specific selection, reflecting clinically important differences among balanced crystalloids (10). In this context, Duan et al.'s analysis of balanced crystalloid use in subarachnoid hemorrhage is clinically relevant. Its observational design leaves the optimal formulation, dose, and patient phenotype unresolved. For now, fluid should be prescribed as a drug: by indication, composition, dose, cumulative balance, and neurological context.Analgesia and sedation are particularly contentious after neurosurgery. Sedatives can obscure serial consciousness assessment and may delay recognition of hemorrhage, ischemia, or impending herniation; insufficient treatment of pain and agitation can also worsen physiological stress and expose patients to device removal or secondary injury. The Chinese expert consensus therefore recommends individualized targets according to brain injury, intracranial dynamics, ventilation, procedures, and the need for neurological assessment (11). Wang et al. describe a single-center, single-arm feasibility protocol without a randomized comparator. In 65 selected adults after craniotomy who are restless or agitated but do not require deep sedation, non-pharmacological measures and remifentanil-based analgesia are titrated to Richmond Agitation-Sedation Scale (RASS) scores of -2 to +1 and Critical-Care Pain Observation Tool (CPOT) scores of 0-1; midazolam or propofol remains available as rescue therapy. The primary endpoint is successful protocol management during the first 24 h. The protocol is designed to assess feasibility and safety. Its full report should clarify the frequency of rescue sedation and whether neurological assessment remains feasible without avoidable agitation or physiological harm.The two delirium studies warrant cautious interpretation. Sun et al. reported less postoperative delirium with preoperative warming plus dexmedetomidine in 153 analyzed older adults undergoing hip-fracture surgery, but some randomized participants were excluded and the setting was not neurocritical care. Dong et al. tested a virtual-reality package after cardiac surgery; only eight delirium events occurred among 40 participants. The package appeared feasible to deliver. However, orientation, sleep support, mobilization, and the additional staff attention may account for part of the observed effect.Temperature management illustrates the distinction between evidence synthesis and comparative treatment evidence. Using a 6S evidence framework, Zhang et al. appraised 20 sources (seven guidelines, six expert consensuses, four systematic reviews, two evidence summaries, and one clinical decision resource) and summarized 27 best-practice items across preparation, initiation, maintenance, complication management, rewarming, and prognostic management. Its design supports an implementation-oriented synthesis across heterogeneous neurological conditions; comparative treatment effects lie outside its scope. For comatose adults after out-of-hospital cardiac arrest, the Targeted Hypothermia Versus Targeted Normothermia After Out-of-Hospital Cardiac Arrest (TTM2) trial found no mortality or functional benefit from induced hypothermia at 33°C compared with targeted normothermia and early fever treatment (12). European Resuscitation Council-European Society of Intensive Care Medicine (ERC-ESICM) guidance emphasizes continuous core-temperature monitoring and active fever prevention for at least 72 h, while finding insufficient evidence for or against a 32-36°C target (13). Kortli and Nasa similarly emphasize post-resuscitation care as a bundle rather than a temperature target alone. The practical focus is the quality of temperature control: safe delivery, fever prevention, and delayed multimodal neurological prognostication (14,15). The review of decompressive hemicraniectomy pools 14 randomized trials and 1,003 patients with malignant middle cerebral artery infarction. Survival improved, while the functional interpretation varied with age, follow-up, and the chosen threshold; some analyses counted modified Rankin Scale scores of 0-4 as favorable. That definition matters to families. The decision and timing of decompression remain diagnosis-specific. Chen et al. highlighted the Randomized Evaluation of Surgery with Craniectomy for Patients Undergoing Evacuation of Acute Subdural Haematoma (RESCUE-ASDH) trial, in which patients undergoing evacuation of a traumatic acute subdural hematoma were randomized intraoperatively to replacement or non-replacement of the bone flap. Disability and quality of life at 12 months were similar; decompressive craniectomy reduced early additional cranial operations but increased wound complications (1,16). These findings apply specifically to the bone-flap decision during acute subdural hematoma evacuation; prophylactic decompression follows a different evidence base. For malignant middle cerebral artery infarction, European Stroke Organisation (ESO) guidance supports surgery within 48 h in adults aged 60 years or younger, with greater uncertainty in older patients and after 48 h (17). Hu et al. update these estimates within the populations represented by the available evidence. Karam et al., writing about blood stewardship and quantitative futility assessment in bleeding neurotrauma, expose the ethical counterpart of the same problem. When used to guide treatment limitation, prediction tools can create self-fulfilling bias. The appropriate safeguard is repeated clinical assessment combined with multidisciplinary discussion and an explicit account of the patient's values.Hyponatremia after neurological injury requires a mechanistic diagnosis

Linlin Zhang, Jian-Xin Zhou · 0 citations

Leakage-Aware Cross-Dataset Evaluation of Prompt Injection Detection Using Classical Machine Learning and Transformer Models

The widespread adoption of systems based on Large Language Models has made the reliable detection of prompt injection attacks a critical requirement. However, high performance achieved on training and test splits generated from the same data source does not guarantee that models can generalize to prompts from different sources. In this study, a leak-aware cross-dataset evaluation framework is presented to examine the robustness of classical machine learning and Transformer-based prompt injection detection models in the face of data source changes. During the data preparation process, empty records, duplicate prompts, conflicting labels, and text overlaps between datasets were checked. In this context, 38,184 duplicate records and 14 instances with conflicting labels were removed, and the 198 common prompts identified between the training and external test sets were removed only from the training set. Using WordHash and CharHash representations, SGD Logistic and Linear SVM models, as well as DistilBERT and DeBERTa-v3-small, were evaluated on an internal dataset consisting of 426,073 cleaned requests; the models were also tested on an independent dataset of 5,000 examples. While the models achieved performance in the range of approximately 0.997–1.000 in the internal evaluation, significant performance losses were observed in the external evaluation. DeBERTa-v3-small delivered the most balanced results, with an accuracy of 0.7360, a balanced accuracy of 0.7320, a Macro-F1 of 0.7284, and an attack sensitivity of 0.7119. The domain classifier achieved an ROC-AUC of 0.9770 hence indicating a remarkable shift in the distribution across data sources. The results show that internal validation results are not enough for prompt injection detection. Independent external validation, data leakage verification and domain shift analysis should be fundamental parts of reliable model evaluation.

Oğuzhan KİLİM · 0 citations
#artificial intelligence Open access Sep 2026

Implementing a Cognitively Grounded Artificial Moral Advisor: A Multi-LLM Multi-Agent Approach Based on the Cognitive–Reflective Equilibration Model

Large language model (LLM)-based artificial intelligence is increasingly used in ethically consequential human decision-making, yet fully autonomous machine ethics remains unrealistic, motivating architectures that support rather than replace human ethical judgment. This study introduces the Cognitive–Reflective Equilibration Architecture (CREA), a cognitively grounded artificial moral advisor that operationalizes the Cognitive–Reflective Equilibration Model (CREM), in which reflective reasoning guides ethical judgment from intuitive cognition toward a more advanced equilibrium among competing values, drawing on Piaget and Rawls. CREA implements CREM’s 20-step process through four stage-aligned reasoning agents—Cognitive, Reflective, Equilibration, and Evaluation—coordinated via multi-LLM orchestration, in which auxiliary models independently explore principles, generate counterarguments, and score supporting and opposing considerations to externalize reflective deliberation. The architecture was empirically evaluated by comparing four configurations—single-agent, multi-agent, multi-LLM, and multi-LLM with knowledge- and reasoning-bank augmentation—across four indicators of advice quality using 500 matched execution units per configuration. All comparisons are system-internal: advice quality was scored by CREA’s own multi-LLM measurement pipeline rather than by human ethicists, so the findings reflect relative differences among architectures under LLM-based self-evaluation, not normative validity. Within that scope, distributing reflective reasoning across multiple models was associated with higher reason-giving (justifiability) and normative-alignment scores relative to simpler configurations. CREA therefore offers an empirically characterized, auditable advisor architecture whose potential to scaffold human ethical judgment remains a hypothesis for user-centered validation rather than a demonstrated outcome.

Chulmin Kim, Seongjin Ahn · 0 citations
#artificial intelligence Open access Sep 2026

Using AI to support safety management – Analysing occupational safety data using machine learning within the framework of human factors

Organisations collect large amounts of safety data concerning safety events. However, many organisations are still in their infancy in this regard, especially when it comes to utilising large datasets containing textual safety data to support their safety management. Fast-developing Artificial Intelligence (AI) and machine learning methods offer new possibilities for efficient analysis of textual data. In this paper, we discuss aspects related to using AI to support safety management. The main contribution is to examine the possibilities and challenges of using machine learning methods to analyse unstructured textual safety data from organisations’ incident reports and accident investigation texts in order to identify and classify the contributing factors. To this end, we fine-tuned a pretrained Large Language Model (LLM), using Human Factors (HF) and the HF Tool as a framework, to classify the contributing factors mentioned in report texts. The HF Tool supports the systemic analysis of safety incidents and guides the identification of contributing factors at individual, work, group/team, and organisational levels. Furthermore, the LLM developed in this study, AN-HF-classifier-V1, was evaluated on two distinct datasets comprised of excerpts, which were classified according to HF Tool by human factors experts. Our findings suggested that this classifier may support the identification of HF from large textual datasets. However, we found that the quality of textual safety data needs to be improved to include more comprehensive and accurate factors contributing to incidents. The need to improve the quality of data applies regardless of whether AI is used in the analysis or not.

Maria Tiikkaja, Henriikka Kannisto, A. Nurmi et al. · 0 citations

A Knowledge Graph–Augmented Large Language Model for Automated Quality Compliance Checking in Construction

This research proposes a large language model–based approach combined with structured algorithms that process domain-specific competency questions and unstructured construction documents to build ontologies and extract structured knowledge graphs to support the automated generation of quality inspection checklists and compliance checking in construction.

Yi-Heng Wang, Juan Wang, W. Fang · 0 citations

A Decision Support Tool for Determining the Suitable Multifunctional Digital Twin Maturity Levels in Construction Projects

The research concludes that the DST provides stakeholders with high-level insights regarding the suitable DT maturity levels—essential for stakeholder buy-in—assists them in making data-driven decisions especially during the front-end planning stage, and helps determine realistic technology objectives that contribute to project success.

Amir Mahdiyar, Soheil Sabri, Sigrid Adriaenssens · 0 citations

Precision Design of Fluorogenic Probes via Orthogonal Tuning of Binding and Photophysics for Isoform-Selective ALDH2 Imaging.

Fluorogenic probes that report enzyme activity are essential for studying biological functions. However, designing them for targets with low catalytic turnover and narrow substrate specificity remains a significant challenge. Here, we present a precision design framework that separates the requirements for sensitivity and selectivity by integrating molecular docking, quantum chemical modeling of fluorogenic mechanisms, and targeted fine-tuning of the probe structures. As a proof of concept, we developed A5, a fluorogenic substrate for aldehyde dehydrogenase 2 (ALDH2) that exhibits high isoform selectivity and a >240-fold signal enhancement over the standard NADH assay. A5 enables quantitative imaging of ALDH2 activity across multiple biological scales─in blood samples, live cells, and intact mouse brains─and supports the identification of small-molecule activators with therapeutic potential in an Alzheimer's disease model. This work establishes a modular strategy for creating activity-based probes tailored to challenging enzymatic targets, with broad applications in precision imaging, drug discovery, and mechanistic biochemistry.

Rongrong Tao, Yu Chen, Taorui Yang et al. · 4 citations

From tech blogs

See all →
Microsoft Research Blog Aug 31, 2026

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.