Skip to content

Author

K. Hoekzema

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Pangenome discovery and characterization of human protein-coding duplicated genes

Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.

Luyao Ren, DongAhn Yoo, Katarina Vlajic et al. · 0 citations
Open access Aug 2026

A high-resolution human pangenome structural variant resource for improved disease association

Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.

J. Lin, J. Gustafson, J. Wertz et al. · 0 citations