Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Comparison of the predictive roles of CT- and MRI-based endplate regional osteoporosis status measurements for cage subsidence after posterior lumbar interbody fusion.

OBJECTIVE This study aimed to compare the efficacy of endplate Hounsfield unit (HU) values and endplate bone quality (EBQ) scores in predicting cage subsidence (CS) after posterior lumbar interbody fusion (PLIF) in older patients and identify the most discriminative bone mineral density (BMD) assessment indicator. METHODS This retrospective analysis included consecutive patients who underwent PLIF at the authors' institution between January 2016 and February 2024. Clinical data were collected for all patients. Propensity scores were used to match patients with and without CS, and the matched cohort was subjected to conditional logistic regression to investigate the association between radiographic factors and CS. L1 and endplate HU values were derived from CT scans, whereas vertebral bone quality (VBQ) and EBQ scores were derived from MR images. Receiver operating characteristic curve analysis was conducted to assess the predictive value of endplate HU values and EBQ scores for CS and further compare their predictive value with that of L1 HU values and VBQ scores. RESULTS This study included 130 matched patients. The CS group demonstrated lower L1 (p < 0.001) and endplate HU (p < 0.001) values and higher VBQ (p = 0.002) and EBQ (p < 0.001) scores compared with the non-CS group. The conditional logistic regression analysis identified L1 HU value (OR 0.99, 95% CI 0.97-0.99; p = 0.036), endplate HU value (OR 0.99, 95% CI 0.98-0.99; p = 0.044), VBQ score (OR 2.40, 95% CI 1.34-4.32; p = 0.038), and EBQ score (OR 4.46, 95% CI 2.16-9.18; p = 0.003) as independent predictors of CS, demonstrating areas under the curve of 0.722, 0.815, 0.648, and 0.782, respectively. The optimal cutoff for the endplate HU value in predicting CS was 262.11 (sensitivity 83.08%, specificity 73.85%). CONCLUSIONS Endplate HU values indicated a relatively higher predictive performance for CS compared with EBQ scores and served as the most discriminative BMD indicator in patients who underwent PLIF. Measuring the endplate HU value preoperatively helps surgeons select a more appropriate surgical plan and is expected to improve patient outcomes.

Peng Du, Minghui Liang, Ruiyuan Chen et al. · 0 citations
Review Open access Jul 2026

Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians’ Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog

Background Orthopedic-related rare diseases are difficult to diagnose because of their low prevalence, heterogeneous phenotypes, and fragmented knowledge. Large language models (LLMs) can serve as dynamic knowledge-support tools, but their diagnostic performance and effect on physicians’ decision-making remain unclear. Objective This study aims to compare the diagnostic performance of advanced LLMs for orthopedic-related rare diseases and to evaluate the effect of a 2-stage LLM-assisted diagnostic workflow on physicians’ diagnostic accuracy and subjective acceptance. Methods We selected 40 orthopedic-related rare diseases from the Chinese Rare Disease Catalog. A total of 4 general-purpose LLMs each generated 1 primary diagnosis and 5 differential diagnoses per case. Diagnostic accuracy, defined as a correct primary diagnosis, was compared using the Cochran Q test and pairwise McNemar tests with Bonferroni correction. A representative LLM was integrated into a 2-stage workflow involving 27 intermediate and 15 senior orthopedic physicians. Physicians first diagnosed all cases independently and then rediagnosed the same cases after reviewing nonauthoritative LLM suggestions. Physician diagnostic data were primarily analyzed using mixed-effects logistic regression at the individual-diagnosis level. Case-level group accuracy was additionally assessed using >50% and ≥2/3 accurate-physician thresholds. After both rounds, physicians completed an 8-item Likert-scale questionnaire assessing subjective acceptance and workflow perceptions. Results Claude Sonnet 4.5, ChatGPT-5.0, and Gemini 2.5 Pro each achieved 90% (36/40) primary-diagnosis accuracy, whereas DeepSeek-V3.2 achieved 67.5% (27/40; Cochran Q P<.001). Before LLM assistance, mean physician-level accuracy was 42.22% for intermediate physicians and 58.67% for senior physicians; after assistance, it increased to 68.80% and 83.33%, respectively. In the primary mixed-effects logistic regression analysis, physician seniority group and LLM assistance stage were significantly associated with diagnostic correctness (both P<.001), whereas the group-by-stage interaction was not significant (P=.10). Secondary case-level analyses using the >50% threshold showed improvement from 40% (16/40) to 67.5% (27/40) for intermediate physicians and from 57.5% (23/40) to 82.5% (33/40) for senior physicians, with similar findings using the ≥2/3 threshold. Cases accurately diagnosed by all 3 agents increased from 16 to 27. The questionnaire showed high internal consistency (Cronbach α=0.902) and generally positive attitudes, with no significant differences between physician groups (P=.11 to P=.78). Conclusions LLMs achieved high diagnostic accuracy for orthopedic-related rare diseases. In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies. Senior physicians retained higher diagnostic correctness than intermediate physicians. Secondary case-level analyses suggested attenuation of group-level gaps in case-recognition patterns. Physicians reported broadly positive workflow perceptions. Given the same-day repeated-case design and potential short-term recall bias, these exploratory findings should be interpreted cautiously and warrant prospective randomized, crossover, washout-period, independent-case, or real-world evaluations of LLM-assisted diagnostic workflows in orthopedics.

Tusheng Li, Ziqian Ma, Baodong Wang et al. · 0 citations