Jul 2026· Signal Processing and Communications Applications Conference· pp. 1-4· 0 citations· 17 references
Abstract
X-ray computed tomography (CT) is one of today’s most critical imaging modalities, with a wide range of applications spanning from medical diagnosis to industrial inspection. CT images can be severely affected by physically induced degradations such as low-dose noise and beam hardening, which compromise image quality and diagnostic accuracy. Existing AI-based approaches largely treat this problem as a pure classification task, and a system that explains the physical mechanisms of artefacts and provides actionable recommendations to the user has not been systematically addressed. In this study, we propose a vision-language model (VLM) based pipeline that detects CT artefacts, explains their physical mechanisms in natural language, and generates structured, actionable recommendations. LLaVA-1.5-7B and Qwen2-VL-7B models were fine-tuned using QLoRA on the 2DeteCT dataset; following fine-tuning, LLaVA-1.5-7B achieved 99.8% accuracy while Qwen2-VL-7B reached 86.9%. The results demonstrate the effectiveness of domain adaptation for structured artefact assessment.
Introduction and aims Interpreting condylar osseous changes on CBCT is challenging for general practitioners. This study evaluated the ‘zero-shot’ diagnostic performance and utility of Vision-Language Models (VLMs) as AI assistants for detecting condylar abnormalities. Methods We analysed 72 CBCT images from the EHPN study for internal validation and constructed a balanced sample of 70 images from the MMDental dataset for external validation, with dual-radiologist consensus as the ground truth. Three frontier VLMs (Gemini-3, GPT-5.2, and Qwen3-VL) were used to evaluate representative sagittal slices without prior fine-tuning. Performance was measured by diagnostic accuracy, sensitivity, specificity, and balanced accuracy, while the Quality Assessment of Medical AI-generated Information (QAMAI) framework assessed the quality of AI-generated structured reports. This study was reported following STARD-AI and CLAIM guidelines as primary checklists, with TRIPOD-LLM as a supplementary framework. Results Gemini-3 achieved the highest diagnostic accuracy (90.3%; 95% CI: 81.0%-95.5%), outperforming GPT-5.2 (75.0%) and Qwen3-VL (55.6%). External validation on a constructed balanced sample from the MMDental dataset confirmed these findings, with Gemini-3 achieving 90.00% accuracy. Regarding report generation, Gemini-3 surpassed the other models across five QAMAI dimensions, providing more accurate, clear, and clinically useful justifications aligned with diagnostic standards. Conclusion VLMs, particularly Gemini-3, exhibit zero-shot capabilities in identifying condylar changes and generating high-quality diagnostic reports. While these models maintain a conservative diagnostic tendency, they demonstrate potential as preliminary, training-free screening aids in oral maxillofacial radiology. Further multicentre validation is warranted to establish clinical utility. Clinical Relevance Cloud-based VLMs could offer accessible screening assistance by allowing general practitioners to query condylar findings via simple web interfaces without specialized hardware. These tools may contribute to structured reporting workflows, though their diagnostic consistency in routine practice remains to be established, and they should complement rather than replace expert clinical judgment.
Ke Chen, Andrew Zhang, Xianju Xie et al.· International Dental Journal· 0 citations
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment. After pulmonary nodule detection, radiologists manually assess anatomical location, diameter, margin characteristics, and attenuation type to support risk assessment and clinical decision-making. However, this post-detection workflow is time-consuming and can be affected by inter-observer variability. Existing Artificial Intelligence methods often focus on isolated tasks, limiting their use as a unified, clinically grounded interpretation framework. This study presents FZ-VLM, a two-stage Florence-Zephyr Vision Language Model framework for unified structured pulmonary nodule characterization in lung CT. The framework uses a fine-tuned Florence-2 model to extract radiological attributes from expert-annotated 2D axial CT slices, while a Zephyr-7B model uses these attributes to generate nodule descriptions, follow-up recommendations, and longitudinal analyses. Results showed that the Stage 1 model achieved 77.18\% accuracy for anatomical location, 67.96\% accuracy for margin characteristics, and 79.13\% accuracy for attenuation type, with a Mean Absolute Error of 2.58 mm for diameter estimation, outperforming evaluated GPT-4-based baselines as well as the human baseline. Expert radiologist evaluation of Stage 2 showed 93.9\% accuracy, 98.6\% completeness score, 76.1\% clinical relevance, and an overall score of 89.5\%. Safety analysis showed that most outputs were clinically safe, although some follow-up recommendations still required expert review. To the best of our knowledge, this study presents the first two-stage Vision-Language Model framework for structured nodule characterization and clinical decision-making.
Pramita Dutta, Jenita Manokaran, Richa Mittal et al.· 0 citations
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in 2D medical image understanding, their extension to 3D volumetric imaging remains hindered by prohibitive annotation costs and dataset opacity. Current data formats, predominantly consisting of rigid Visual Question Answering (VQA) pairs or unstructured final clinical reports, typically fail to capture explicit clinical reasoning. To address this limitation, we introduce a large-scale structured reasoning dataset constructed via a novel slice-wise data synthesis paradigm. Inspired by the genuine diagnostic workflow of radiologists, this paradigm models visual cognition by decomposing the complex 3D reading process, translating global clinical priors into fine-grained, per-slice observations that are subsequently synthesized into an interpretable Chain-of-Thought (CoT). Crucially, this synthesized reasoning framework enforces essential clinical principles: sequential spatial tracking, multi-slice spatial awareness for artifact mitigation, and differential exclusion. To validate this approach, we instruction-tune a standard 2D-pretrained MLLM baseline using the synthesized data to enhance its volumetric comprehension. Comprehensive evaluations across multiple 3D medical benchmarks demonstrate that our method yields significant performance improvements over the 2D baseline. Furthermore, the resulting model exhibits robust spatial reasoning capabilities and rivals resource-intensive native 3D architectures, effectively bridging the performance gap. Ultimately, this data-centric strategy unlocks deep volumetric understanding and highly interpretable clinical logic without requiring computationally expensive 3D-specific pre-training. The complete repository, including datasets and training workflows, is publicly available at https://github.com/2020420145009/hounsfield.
Zhuoyuan Fu, Zeshang Li, Yiqiong Zhang et al.· 0 citations
The development of generalist artificial intelligence (AI) models for radiology is hindered by a lack of large-scale, three-dimensional (3D) imaging datasets with precise spatial annotations. While numerous datasets provide image-level labels for chest computed tomography (CT), these are insufficient for training models that can accurately localize findings. To address this gap, we present PatchChestCT, a large-scale, publicly available dataset for multi-abnormality localization in non-contrast chest CT. Using CT-RATE as the source cohort, PatchChestCT provides 3D patch-level annotations for nine clinically significant abnormalities across 2,201 physician-reviewed CT studies, with one reconstructed volume labeled per study. We introduce a token-aligned annotation scheme that is efficient, scalable, and naturally integrates with modern deep learning architectures like Transformers. To validate the dataset’s utility, we show that models trained with our patch-level labels achieved higher localization performance than weakly supervised baselines trained on image-level labels alone, across multiple architectures. PatchChestCT provides a public resource for training localization-aware models in chest CT.
Computed tomography (CT) images obtained in clinical settings are often acquired with diverse scanner types and acquisition parameters. They may exhibit significant variations in fields of view (FOVs) and levels of contrast enhancement. An automated method for navigating the content in these images is therefore essential for effective dataset curation and downstream analyses. This work introduces a framework called Body- Part-Phase Regression (BPPR) to automatically identify regions of interest and determine the contrast enhancement phase of body CT images. The framework consists of two key components: (1) A two-phase body part regression method for predicting the anatomical location of 2D slices within 3D volumes. (2) A circular regression model for predicting the contrast timing of CT images (i.e., the timing of the scan relative to contrast agent injection) from a continuous perspective, providing a fine-grained understanding of contrast differences, particularly in relation to patient-specific vascular effects. These two components are linked via a positional weighting mechanism which enhances volumelevel phase prediction by leveraging slice-level predictions. By unifying the "part" and "phase" regression models, our framework establishes a cohesive approach to continuous content navigation in CT images. We train and evaluate our models on large-scale datasets consisting of multi-contrast images and compare their performance with alternative approaches pursuing similar goals. The experiments demonstrate improvements in both slice localization and contrast phase prediction. In particular, the two-phase training scheme reduces the slice localization error of previous body part regression methods from 9.2 mm to 6.1 mm. We also discuss the distinctive advantages of BPR over segmentationbased approaches and highlight potential clinical applications that may benefit from the proposed BPPR framework.
Dingjie Su, Qingyun Yang, K. D. V. Schaik et al.· IEEE transactions on bio-med...· 0 citations