Skip to content
Open access

Pangenomes aid accurate detection of large insertions and deletions from targeted sequencing: the case of cardiomyopathies

Aug 2026 · Genome Medicine · 0 citations

TL;DR

The results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Abstract

Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools’ performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1–19 bp; F1 0.975 vs. 0.968). Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.

Read PDF

Similar papers

Open access Aug 2026

Rapid minigene workflow for functional reclassification of splicing variants in hereditary cancer diagnostics.

BACKGROUND Next-generation sequencing of cancer predisposition genes is routinely used in hereditary cancer diagnostics. However, a substantial fraction of detected variants remains clinically unresolved. Using a customised 77-gene panel, we analysed 2142 individuals and identified 384 pathogenic or likely pathogenic variants across 54 genes, corresponding to a diagnostic yield of approximately 18%. Despite this, 17% of cases carried variants of uncertain significance, many of which were suspected to affect pre-mRNA splicing and are particularly challenging to interpret due to the limited reliability of in silico predictions and lack of experimental evidence. METHODS To address this diagnostic gap, we developed a streamlined minigene-based workflow for rapid functional evaluation of splicing variants and applied it retrospectively. The approach relies on synthetic DNA and recombination-based cloning, eliminating the need for patient-derived RNA and enabling efficient construct generation within a clinically compatible timeframe. Computational prioritisation using AlphaGenome was integrated to support variant selection, while experimental assays provided direct evidence of splicing outcomes. RESULTS Application of this strategy allowed the reclassification of previously unresolved variants and clarified cases with discordant computational evidence. Importantly, the workflow is designed for implementation in routine diagnostic settings, with a turnaround time aligned with clinical reporting requirements. CONCLUSION This approach provides a robust and scalable framework for functional interpretation of splicing variants, improving diagnostic resolution and supporting more informed clinical decision-making in hereditary cancer genetics.

Noemi Calandra, Elisabetta Mereu, P. Ogliara et al. · 0 citations
Open access Aug 2026

A high-resolution human pangenome structural variant resource for improved disease association

Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.

J. Lin, J. Gustafson, J. Wertz et al. · 0 citations
Open access Aug 2026

Detecting CYP2C19 deletions from genotyping array signals using neural networks

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

Burak Yelmen, R. Hofmeister, Viido Kaur Lutsar et al. · 0 citations
Open access Jul 2026

Whole-genome discovery of pathogenic snRNA variants and efficient extended-exome screening

Summary Pathogenic variants in small nuclear RNA (snRNA) genes have recently emerged as a major cause of Mendelian disorders, particularly neurodevelopmental disorders, yet they remain difficult to detect in routine diagnostics because conventional whole-exome sequencing (WES) does not capture snRNA loci. Here, we reanalyzed whole-genome sequencing (WGS) data from 1,578 unsolved probands and identified pathogenic variants in multiple snRNA genes, including RNU4-2, RNU2-2, RNU5B-1, and RNU4ATAC, accounting for 1.2% (19 patients) of previously unsolved cases. We then developed an snRNA-extended WES approach by incorporating capture probes targeting 50 snRNA genes into a standard exome design. Benchmarking demonstrated robust, uniform coverage across all targeted snRNA loci without increasing sequencing depth. Applying this approach to patient samples reliably detected disease-causing snRNA variants previously identified by WGS. Our results establish snRNA-extended WES as a cost-effective and scalable strategy to improve diagnostic yield and bridge the gap between recent gene discoveries and clinical genomic practice.

Yuka Nakano, Hisato Suzuki, Yukiko Kuroda et al. · 0 citations
Open access Aug 2026

Towards routine genetic testing of repeat expansions in neurogenetic diseases using multiplex CRISPR-Cas9-targeted long read sequencing

Abnormal expansion of nucleotide repeats was first identified 34 years ago as a unique mutational mechanism. It is now linked to numerous neurogenetic disorders, several of which discovered only recently. The identification of these expansions has led to various classifications based on clinical presentation, repeat nature and genomic location (coding or non-coding regions). Precise diagnosis of these conditions relies on molecular testing, currently performed on a gene-by-gene basis. Their analysis remains challenging, especially for long expansions. We evaluated CRISPR-Cas9-mediated target enrichment coupled to Oxford Nanopore Technologies (ONT) long read sequencing, to accelerate and improve the time-consuming molecular diagnosis of repeat expansion disorders. We simultaneously targeted nine loci involved in 10 repeat expansion disorders in a single capture panel, including FMR1 , HTT , DMPK , CNBP/ZNF9 , ATXN2 , JPH3 , FXN , C9ORF72 and RFC1 , covering a broad range of repeat types, sizes and diagnostic needs. Results were compared with standard routine testing methods. ONT sequencing using Flongle flow cells yielded results consistent with standard techniques for most loci, particularly for non-complex repeats. However, limitations were observed for structurally complex regions such as RFC1 , and inter-run variability required the aggregation of multiple Flongle runs per sample to achieve robust genotyping. These findings highlight both the potential and current limitations of CRISPR-Cas9-enriched ONT sequencing for multiplex diagnosis of repeat expansion disorders in a clinical setting. The approach deserves further development, particularly optimisation of protocols, inclusion of larger sample sizes, and comparison with alternative technologies.

P. Fergelot, C. Boury, B. Penaud et al. · 0 citations
Open access Jul 2026

Unraveling missing variants through target capture-based long-read sequencing in autosomal recessive disorders.

Short-read sequencing (SRS)-based disease-targeted NGS gene panels have revolutionized rare disease diagnostics but often leave autosomal recessive cases unsolved when only one pathogenic allele is detected. Missing variants may reside in deep intronic regions or involve structural variants (SVs) undetectable by SRS. To improve diagnostic yield, we implemented a cost-effective target capture-based long-read sequencing (LRS) assay covering 56 genes and retrospectively analyzed 78 patients suspected of autosomal recessive disorders who remained undiagnosed after SRS. Functional validation using reverse transcription PCR (RT-PCR) and minigene assays was performed to further determine pathogenicity. Target capture-based LRS solved 25.6% (20/78) of cases by identifying 10 SVs, 3 deep intronic variants experimentally confirmed to cause aberrant splicing, and 7 cases in which haplotype phasing confirmed that variants were in trans with the known pathogenic variant, leading to reclassification of the VUS as likely pathogenic. This study demonstrates that target capture-based LRS effectively detects diverse types of variants missed by SRS. Integrating this assay into stepwise diagnostic workflows offers a practical and cost-effective strategy to enhance diagnostic yield in autosomal recessive diseases. However, because this cohort was retrospectively defined based on a prior single-allele detection by SRS, this 25.6% (20/78) diagnostic yield reflects performance within a highly enriched population and should not be directly extrapolated to unselected rare disease cohorts.

Jee-Soo Lee, K. Ryu, Hyesu Lee et al. · 0 citations