Abstract Motivation Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation. Results We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson’s Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise’s value in complex loci like 17q21.31. Availability and implementation CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.
Nicole Kuznetsov, Kensuke Daida, M. Makarious et al.· Bioinformatics Advances· 0 citations
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce CLExEval, a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking. CLExEval combines 5,600 expert-physician annotations with 200 clinical reasoning traces derived from 40 rare diagnostic cases. Our analysis identifies three recurring failure patterns: (i) verbosity bias, where GPT-4o-mini's diagnostic accuracy drops from 95.0% to 32.5% under information scarcity; (ii) a hidden knowledge paradox, where a specialist model reaches 92.5% maximum diagnostic potential but fails to retrieve that knowledge reliably in verbose contexts; and (iii) a 68.6% reasoning-to-output mismatch, where correct diagnoses appear in reasoning traces but are not reflected in final answers. We further evaluate the LLM-as-a-Judge paradigm on a human-verified failure set (n = 142). GPT-4o-mini approved 47.9% of clinically incorrect outputs, while HuatuoGPT-o1 approved all validly scored failures and showed a positive self-preference bias. These results suggest that standalone automated clinical evaluations can substantially overestimate clinical reliability without expert-grounded validation.
Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al.· arXiv.org· 0 citations