The genome will support comparative genomics, marker development, and breeding strategies aimed at improving fiber quality, disease resistance, and climate resilience in abaca and provides insights into genome organization, repeat landscape, and fiber-related gene content.
Abstract
Introduction Abaca (Musa textilis Née) is an important fiber crop cultivated primarily in the Philippines and valued for its exceptional fiber strength and industrial applications. Despite its economic importance, genomic resources for abaca remain limited, constraining efforts in molecular breeding and trait improvement. Here, we present a high-quality de novo genome assembly and functional annotation of M. textilis cv. Inosa, a commercially important cultivar known for superior fiber quality. Methods The genome was sequenced using PacBio HiFi technology and assembled de novo, followed by repeat annotation, gene prediction, functional characterization, and comparative genomic analyses with other Musa genomes. Orthology, synteny, and fiber-related gene analyses were performed to investigate genome evolution and identify genes associated with fiber development. Results The assembled genome spans 612.5 Mb across 388 contigs, with a contig N50 of 9.02 Mb and a BUSCO completeness score of 98.9%, indicating high assembly quality and completeness. Functional annotation identified 37,403 high-confidence protein-coding genes. Repetitive elements account for 59.18% of the genome, representing one of the highest repeat contents reported among Musa genomes. Notably, Polinton transposons, a rarely reported transposable element class in Musa, were identified. Comparative genomic analyses revealed strong macrosyntenic conservation with other M. textilis assemblies and identified 226 Inosa-specific orthogroups. In addition, 348 proteins associated with fiber biosynthesis were annotated, including key enzymes and regulatory proteins involved in cellulose and lignin biosynthesis pathways. Discussion This high-quality genome assembly expands the genomic resources available for abaca and provides insights into genome organization, repeat landscape, and fiber-related gene content. The genome will support comparative genomics, marker development, and breeding strategies aimed at improving fiber quality, disease resistance, and climate resilience in abaca.
Tea (
Camellia sinensis
) is a globally important economic crop. Among elite cultivars, ‘Fuding Dahao’ is particularly prized for its superior agronomic traits and its central role in premium white tea production. However, the lack of a high-quality chromosome-level genome for this regionally adapted cultivar has hindered molecular breeding efforts. Here, we present the first chromosome-level reference genome of ‘Fuding Dahao’ assembled using PacBio HiFi long-read sequencing and Hi-C chromatin interaction mapping. The final 3.30 Gb assembly is highly contiguous, with 90% of the sequences anchored to 15 pseudochromosomes. A total of 54,345 protein-coding genes were predicted, representing a substantial improvement in both assembly contiguity and annotation completeness compared with previously published tea genomes. This high quality genome provides a critical resource for dissecting the genetic basis of white tea quality traits and accelerating molecular breeding programs. Our results fill a major gap in tea genomics and lay a solid foundation for the development of superior, locally adapted tea cultivars.
Yang Chen, Lizhong Wang, Dengfeng Shen et al.· Scientific Data· 0 citations
Cyclamen is an economically important ornamental plant widely cultivated for its diverse floral characteristics and adaptation to cool climates. Despite its horticultural significance, genomic resources for this species remain limited, hindering molecular studies and genomics-assisted breeding. Here, we report the first highly contiguous nuclear genome assembly of C. persicum generated using high-fidelity long-read sequencing. The assembled genome spans 1.48 Gb, consisting of 126 contigs with an N50 length of 52.3 Mb. Telomeric repeat analysis identified eight contigs containing telomeric sequences at both ends, suggesting the presence of near-complete chromosome assemblies. Genome completeness assessment using BUSCO indicated 98.1% completeness. Repetitive sequences occupied 82.9% of the assembly, with long terminal repeat retrotransposons accounting for 42.1% of the genome. A total of 40,223 protein-coding genes were predicted, with a complete BUSCO score of 95.7%. Comparative orthogroup analysis with five representative eudicot species identified 430 orthogroups specific to C. persicum and 363 orthogroups shared exclusively between C. persicum and Primula kwangtungensis, indicating the presence of both lineage-specific and Primulaceae-conserved gene families. These findings provide critical insights into gene family evolution within Primulaceae and establish an essential comparative framework for future genomic studies. The genome resource presented here provides an invaluable foundation for investigating genome evolution, gene function, and trait-associated loci in cyclamen, effectively facilitating molecular breeding and genetic improvement in this ornamental species.
K. Shirasawa, Y. Akita, Y. Mizunoe et al.· bioRxiv· 0 citations
Khang Dan 18 (KD18) is an Oryza sativa L. subsp. indica rice cultivar widely cultivated in northern Vietnam and used as an experimental and breeding background in Vietnamese rice research. Although KD18 has previously been represented in low-depth population resequencing datasets, a contiguous and annotated cultivar-specific genome has not been available. Here, we report a chromosome-scale genome assembly of KD18 generated using Oxford Nanopore long-read and Illumina short-read sequencing. The 395.3-Mb assembly comprises 12 chromosome-scale pseudomolecules containing approximately 95% of the assembled sequence and 99.6% of the predicted protein-coding genes. The assembly showed 97.2% BUSCO completeness, an average Merqury quality value of 46 and a long terminal repeat assembly index of 13.21. A total of 56,546 protein-coding genes representing 71,237 transcripts were predicted, with 99% BUSCO and 98.68% OMArk completeness. These statistics are similar to those of other high-quality genome assemblies that were recently published for different Asian rice cultivars, therefore providing a cultivar-specific genomic resource for research involving KD18 and KD18-derived materials.
T. Q. Nguyen, K. Do, T. M. Vu et al.· bioRxiv· 0 citations
This high-quality, chromosome-level reference genome provides a foundational resource for understanding the population genetic structure, adaptive evolution and speciation mechanisms of C. appendiculata, thereby offering valuable insights into its evolutionary history and conservation.
Yongchao Tang, B. Xiao, Ruimin Yu et al.· Scientific Data· 0 citations
We present a chromosome-scale genome assembly and annotation of anise hyssop (Agastache foeniculum), an aromatic perennial herb widely used for medicinal, horticultural, and ornamental purposes. The genome was assembled using PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Omni-C proximity ligation data, with gene annotation supported by RNA-seq data from leaf tissue. The final assembly spans 482.39 Mb, of which 434.47 Mb (90.06%) were anchored into nine chromosome-scale pseudomolecules. Structural annotation identified 28,193 protein-coding genes. Genome completeness was assessed using BUSCO, yielding scores of 98.0% (embryophyta_odb10, genome mode) and 95.5% (protein mode). This chromosome-scale genome assembly provides a foundational genomic resource for comparative and functional genomics within the genus Agastache and the Lamiaceae family.
ABSTRACT Advancements in plant genome sequencing and assembly have enabled the production of increasingly accurate and contiguous genome sequences. Here, we present the chromosome‐level assembly of the durum wheat ( Triticum turgidum L. ssp. durum, cv. Svevo) reference genome produced using accurate long‐reads, optical mapping and Hi‐C. The new assembly (Svevo Rel.2.0) comprises 263 hybrid scaffolds with an N50 value of 112.3 Mb, arranged into 14 contiguous pseudomolecules spanning 10.4 Gb. The Svevo Rel.2.0 genome assembly was annotated using extensive short‐ and long‐read RNA sequencing data obtained from 60 tissue/treatment combinations. The resulting annotation comprises 68 154 high‐confidence protein‐coding genes, which have been integrated into a comprehensive transcriptome atlas accessible through an eFP browser. Annotation was manually curated for storage protein gene families and for Leucine‐Rich Repeat‐Containing Receptor genes yielding 3763 LRR‐CR loci. The genome assembly's accuracy and completeness were demonstrated by the correct reconstruction of the physical map of Tg1‐B (Tenacious glumes 1), a locus controlling the free threshing trait located on chromosome 2B that was not assembled in the previous genome release (Svevo Rel.1.0). A wealth of 6621 QTLs/MTAs from the literature were mapped onto Svevo Rel.2.0 to identify QTL hotspots and trait‐specific candidate genes. The ancestry of the durum genome to representative wild emmer populations from North‐Eastern and Southern‐Levant Fertile Crescent assessed by tracing haplotype transmission patterns revealed a clear mosaic pattern. This new durum reference genome, enhanced with advanced annotation and an expression atlas linked to QTLome data, is the most comprehensive tool available for durum wheat genomics.
E. Mazzucotelli, C. Forestan, Gina Zastrow-Hayes et al.· Plant Biotechnology Journal· 0 citations