ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.
Abstract
Abstract Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning–based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.
This work describes the full spectrum of genetic variation and shows that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions.
J. Lin, J. Gustafson, J. Wertz et al.· medRxiv· 0 citations
Abstract Motivation Accurate detection of genetic variants, including single nucleotide polymorphisms (SNPs), small insertions and deletions (INDELs), and structural variants (SVs), is essential for comprehensive genomic analysis. While short-read sequencing performs well for SNP and INDEL detection, it remains limited...
Can Luo, Y. Liu, Han Liu et al.· Bioinformatics Advances· 0 citations
Abstract Detecting structural variations (SVs) via long-read sequencing remains difficult due to algorithmic variations and genomic complexity, alongside a shortage of comprehensive benchmarks for somatic variants. We present a unified benchmarking framework covering both germline and somatic SV detection, which evalua...
Hua Shi, Yi-Hang Lin, Da-Chen Liu et al.· Briefings in Bioinformatics· 0 citations
Motivation Copy-number variants (CNVs) contribute to human disease and population trait variation. CNV detection from large whole-genome sequencing cohorts remains computationally demanding, as most methods require BAM or CRAM files. Genomic VCF (gVCF) files are smaller, routinely generated by standard variant-calling...
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) ar...