Skip to content
Dataset Open access

Whole-Genome Variants Resource of 144 Oryza rufipogon Accessions

Aug 2026 · Scientific Data · Vol 13 · 0 citations · 32 references
Medicine

TL;DR

The findings offer a comprehensive genomic variation resource for O. rufipogon, supporting future genetic research and rice breeding efforts, and confirming the data’s reliability.

Abstract

The global food security crisis highlights the urgent need to improve rice yield and quality. Oryza rufipogon, the wild progenitor of cultivated rice, offers valuable genetic diversity, yet its genome remains underexplored. We conducted whole-genome resequencing (WGS) of 144 O. rufipogon accessions from 13 countries, generating ~400 GB of data using the Illumina NovaSeq 6000 platform. After quality control, 97.58% of reads aligned to the reference genome, achieving 87.97% coverage with 12.47× average depth. We identified a total of 6,183,832 SNPs and 748,493 InDels across 12 chromosomes. Chromosome 1 exhibited the most variants (715,636 SNPs and 91,276 InDels), while chromosome 9 had the fewest (384,000 SNPs and 44,817 InDels). Mutation rates were 67 SNPs and 559 InDels per million base pairs. Notably, 40.75% of SNPs and 46.09% of InDels were found upstream of genes, with 19.13% and 22.26% located downstream, respectively. A transition-to-transversion ratio (Ti/Tv) of 2.72 confirmed the data’s reliability. Our findings offer a comprehensive genomic variation resource for O. rufipogon, supporting future genetic research and rice breeding efforts.

Read PDF

Similar papers

Review Open access Aug 2026

Whole-Genome survey and microsatellite analysis of spotbanded scat Selenotoca multifasciata

In this study, we conducted a comprehensive genome survey of Selenotoca multifasciata using Illumina short-read sequencing technology. A total of 51.63 Gb of high-quality clean sequencing data were generated, with Q20 and Q30 values reaching 98.77% and 96.64%, respectively. The 42.59 Gb of clean reads were assembled into 551,813 contigs (587.82 Mb) and 431,116 scaffolds (593.38 Mb). 17-mer frequency analysis estimated a genome size of 575.38 Mb, with 42.20% GC content, 0.43% heterozygosity, and 25.97% repeat ratio. 214,419 SSR loci were detected genome-wide, with dinucleotide repeats being the most prevalent type (80.49%) and AC/AG as the dominant motifs. Among 53 tested markers, 30 produced clear and stable bands, and 9 polymorphic loci were applied to assess the genetic diversity of a wild S. multifasciata population from Zhanjiang Bay. These results provide a valuable genomic basis for whole-genome sequencing and molecular marker development in S. multifasciata and related Scatophagidae species.

Chao Peng, Gaojie Chen, Peirong Ye et al. · 0 citations
Open access Aug 2026

Chromosome-scale genome assembly and annotation of the Vietnamese indica rice cultivar Khang Dan 18

Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.

T. Q. Nguyen, K. Do, T. M. Vu et al. · 0 citations
Open access Jul 2026

Identification of genome size and heterozygosity in 510 Jujube (Ziziphus jujuba Mill.) germplasms based on deep resequencing

Introduction Genome size and heterozygosity represent the fundamental genetic attributes of a species, but their intraspecies diversity and evolution remain largely elusive. Methods To address this, we developed a deep-resequencing-based approach for genome estimation by optimizing resequencing depth (30×) and kmer value (21) in Chinese jujube (Ziziphus jujuba Mill.) to achieve the best balance among accuracy, stability, and cost. Using this optimized pipeline, we estimated genome size and heterozygosity across a large panel of jujube accessions. Results We estimated the genome size of 296 cultivated jujube genotypes from 293 to 409 Mb (mean 347 Mb, CV 5.94%) with heterozygosity of 1.30% to 2.01% (mean 1.72%, CV 8.21%); for 214 wild jujube, genome size varies between 311 and 391 Mb (mean 339 Mb, CV 3.40%), and heterozygosity 1.22% to 2.13% (mean 1.74%, CV 6.79%). Additionally, analyzing 798 jujube accessions with resequencing data in NCBI (using k-mer 21) revealed that low sequencing depths (10∼20×) of most accessions (98.61%) results in underestimated genome sizes (4.30–309 Mb) and overestimated heterozygosity (2.17-11.8%). Shannon-Wiener Index analysis for genome heterozygosity and size in jujube indicate a medium diversity level with cultivated jujube higher that wild one. Both genome size and heterozygosity show normal distribution, and there is a significant negative correlation between them. Discussion During domestication from wild to cultivated jujube, average genome size increased by 8 Mb, while heterozygosity slightly decreased. A reference standard for genome size and heterozygosity is proposed and jujube is characterized as a small genome with mediumtohigh heterozygosity. This study establishes a powerful approach for estimating genome characteristics, enriched genomic resources for jujube, and provides insights for plant intraspecific genome diversity and evolution.

Wenshu Zhao, Yihan Yang, Hao Wu et al. · 0 citations
Open access Aug 2026

Whole Genome Sequencing of the Moruga Hill Rice (Oryza glaberrima) Reveals Its African Ancestry and the Presence of Candidate Stress-Tolerance Genes

Moruga Hill Rice (MHR) is an African rice (Oryza glaberrima Steud.) brought to Trinidad by formerly enslaved African Americans and has been grown for many generations in Trinidad at subsistence and commercial scale. Despite its historical and agricultural significance, genomic resources specific to MHR remain unexplored, and its genetic composition, evolutionary history, and potential agronomic traits have not been characterized. This current study presents the first draft genome assembly of the MHR genome using a hybrid sequencing approach. The MHR genome size was found to be ~372.9 Mb with 56,073 predicted genes. Variant analysis revealed a total of 3,318,242 variants, of which 2,440,476 were SNPs, and 877,766 were InDels. Several candidate genes encoding proteins with orthology to previously characterized biotic resistance and abiotic stress-responsive genes in rice were identified. Potential gene families identified prompt further investigation of their roles in MHR drought and salt stress responses. Phylogenomic analysis of O. glaberrima landraces suggests that MHR shares close genetic affinity with the IRGC−104595 Malian landrace, consistent with historical records. This assembly thus expands the African rice genomic repository, providing a foundation to understand the genetic architecture underlying key phenotypic traits and identifying potential novel gene sources in MHR for rice improvement in the Caribbean region.

Uddesh M. Sahadeo, Omar Ali, A. Ramsubhag et al. · 0 citations
Open access Aug 2026

First Chromosome-Scale Genome of Cornus kousa K2.

Kousa dogwood (Cornus kousa Hance) is a popular flowering ornamental tree in the United States (U.S.), largely due to its pest and disease tolerance. Here, we present the first chromosome-scale, diploid genome assembly of C. kousa K2, a foundational breeding parent. The final chromosome-scale genome assemblies are 1602.739 Mb for Hap 1 and 1598.929 Mb for Hap 2. The complete Benchmarking Universal Single-Copy Ortholog (BUSCO) for Hap 1 and 2 were 98.8% and 98.4%, respectively. Between the two haplotypes, 98.97% of the genome is placed into chromosomes. 30,799 and 31,044 genes were annotated in Hap 1 and Hap 2, respectively. This annotated genome assembly for C. kousa provides insight into the genomic composition of the species and will enhance our understanding of the genetic control of traits of interest in breeding programs and the evolutionary history of the Cornus genus.

Trinity P H Mehler, S. Boggess, Marcin Nowicki et al. · 0 citations