This review synthesizes the transformative evolution of rapeseed genomics, traversing from initial fragmented references to the modern era of gap-free Telomere-to-Telomere (T2T) assemblies and graph-based pan-genomes, and underscores the pivotal shift from descriptive genomics to the precision engineering of climate-resilient, high-yielding polyploid crops.
Abstract
Oilseed rape (Brassica napus) serves as a cornerstone of global vegetable oil production, yet its genetic improvement has historically been impeded by a complex allopolyploid genome. This review synthesizes the transformative evolution of rapeseed genomics, traversing from initial fragmented references to the modern era of gap-free Telomere-to-Telomere (T2T) assemblies and graph-based pan-genomes. We highlight how these advanced resources resolve previously inaccessible repetitive regions and centromeres, revealing how structural variations (SVs) and homoeologous exchanges (HEs) drive key adaptive traits and morphotype diversification. Furthermore, we examine the integration of large-scale resequencing with sophisticated multi-omics pipelines to bridge the gap between statistical associations and biological causality. Through case studies, such as the characterization of BnRRF for seed weight and BnA09MYB47a for seed coloration, we illustrate the power of combining transcriptomics with Clustered Regularly Interspaced Short Palindromic Repeats-associated Protein 9 (CRISPR-Cas9) for functional validation. Finally, we explore the frontier of “Genomic Design,” where Artificial Intelligence (AI) algorithms like Target-Oriented Prioritization (TOP), combined with Speed Breeding 2.0 protocols, promise to accelerate the development of next-generation cultivars. This synthesis underscores the pivotal shift from descriptive genomics to the precision engineering of climate-resilient, high-yielding polyploid crops.
Background Advances in next-generation sequencing have accelerated genome-wide exploration of genetic diversity in underutilized oilseed crops. Salvia hispanica L. (chia), a high-nutrient pseudocereal rich in omega-3 fatty acids, is increasingly valued for its health benefits and commercial potential, yet it remains poorly characterized at the genomic level. Understanding the scale and nature of genomic variation is essential for improving complex traits such as oil yield, stress tolerance, and seed quality. Methods Two contrasting chia genotypes, Black-chia (CACH-B) and White- chia (CACH-W), were resequenced using the Bio-Resequencing Toolkit (BRT) pipeline. High-coverage sequencing, with a mapping rate exceeding 99% and an average depth of approximately 28×, facilitated the detection and annotation of single-nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy-number variations (CNVs), and structural variants (SVs). The functional classification of variant impacts enabled the identification of genes potentially linked to metabolic and adaptive traits. Results A total of 1.97 million SNPs, 401,493 InDels, 836 CNVs, and 15,288 SVs were identified across the chia genome. Notably, approximately 53% of exonic SNPs were non-synonymous (dN/dS ≈ 1.28), predominantly affecting lipid metabolism, transcriptional regulation, and stress response pathways, potentially altering key agronomic traits. In addition, CNV hotspots were concentrated in chromosomes 3 and 6, overlapping MYB, WRKY, and bZIP transcription factor loci, may potentially be involved in stress tolerance and yield. Furthermore, structural rearrangements, including inversions and duplications within the FAD2, FAD3, and CYP450 gene clusters, were potentially associated with seed pigmentation and omega-3 biosynthesis, pointing to their potential breeding relevance. Observed heterozygosity (Hₒ ≈ 0.71) and nucleotide diversity (π ≈ 7 × 10−3) indicated moderate to high allelic richness. In addition, the low FST value (0.038) indicates substantial genomic similarity between the two genotypes. Conclusion This study presents the first comprehensive map integrating SNPs, CNVs, and SVs in S. hispanica L. The results reveal a structurally dynamic genome characterized by substantial sequence and structural variation, providing valuable insights into genomic diversity and potential adaptive mechanisms in chia. The coexistence of high SNP diversity and abundant structural variation underpins chia's nutritional specialization and environmental resilience. These results deliver a foundational genomic resource for marker-assisted breeding, genome-wide association studies, and the development of climate-resilient chia cultivars.
C. Azzam, M. Rizk, R. Arafa et al.· Journal of Genetic Engineeri...· 0 citations
Wild perennial plants can be domesticated to make agriculture more diverse and resilient, but many have large genomes that have been recalcitrant to analysis. Here, we report phased genome assemblies for Silphium integrifolium Michx. and S. perfoliatum L., two species native to North America under domestication, and demonstrate the utility of trio-binning for genome assembly using an interspecific hybrid. These genomes have chromosomes reaching 1.8 Gb and a helical structure preserved during interphase with a loop circumference of 43 Mb. A genome-informed low coverage and target sequencing strategy enables the refinement of the genus phylogeny, reveals the spatial distribution and structure of natural populations, and identifies 81 loci associated with environmental and domestication traits. Variants in a MATE transporter, α/β hydrolase, and ortholog of Arabidopsis ACT Domain Repeat (ACR4) protein explain significant variance in floral architecture. These advances in genome assembly and genotyping could expand the range of candidates for de novo crop domestication. Silphium species native to North American prairies show strong drought tolerance. This study presents a haplotype-phased genome of a hybrid between S. integrifolium (oilseed crop) and S. perfoliatum (biomass/fiber crop), identifying loci linked to environmental adaptation and domestication.
Renan Souza, J. Clevenger, Jerry W. Jenkins et al.· Nature Communications· 0 citations
Rye (Secale cereale L.) is an important cereal crop known for its high yield potential and tolerance to biotic and abiotic stresses. However, its large, repeat-rich, and heterozygous genome has posed challenges for assembly compared to related species such as wheat and barley. Here, we present a high-quality, chromosome-scale genome assembly of the inbred line Lo7, generated using PacBio HiFi, Oxford Nanopore, Hi-C, and BioNano technologies with the TRITEX pipeline. The resulting Lo7_V3 assembly spans 6.76 Gb with a contig N50 of 128 Mb, correcting previous misorientations and fully assembling all seven centromeres. Repetitive clusters containing rye-specific satellite sequences (pSc200 and pSc250) are contiguously assembled. Their chromosomal positions are validated using FISH. Centromeric retrotransposon analysis reveals RLG_Abia and RLG_Abigail as abundant, recently active elements, unlike in wheat. Collectively, the Lo7_V3 genome assembly provides an improved genomic resource for future genomic research in rye and related cereal species. The large, repeat-rich, and highly heterozygous rye (Secale cereale L.) genome has posed significant challenges for genome assembly. Here, the authors present an improved rye genome assembly and uncover unique retrotransposon organizations within its centromeres.
Erwang Chen, Carlotta Marie Wehrkamp, Srijan Jhingan et al.· Nature Communications· 0 citations
ABSTRACT Sesame ( Sesamum indicum L., 2n = 26) is one of the oldest oilseed crops and is often called the ‘queen of oilseeds’ due to its high content of unsaturated fatty acids and natural antioxidants. Despite its long history, the origin and global spread of cultivated sesame remain unresolved. We assembled a telomere‐to‐telomere (T2T), high‐quality reference genome of sesame (cv. Yuzhi11) to investigate sequence differences between genomes and its origin and the local adaptation evolution of flowering time (DF). We generated a 305 Mb T2T sesame reference genome (cv. Yuzhi11) with > 99.99% base‐level accuracy, identifying 31 063 protein‐coding genes. Repetitive elements accounted for 52.03% of the genome. Population genomic analysis of 927 accessions from 14 regions identified four major groups. Integrative analyses of linkage disequilibrium decay (LD), nucleotide diversity (π), and fixation index (F ST) support East Africa as the center of origin, with subsequent migration through the Middle East, to South Asia, South‐East Asia, East Asia and ultimately to other parts of the world. Genome‐wide association studies (GWAS) and selection scans identified 30 genes associated with flowering time. SiUBP16 is a candidate associated with 7.6% of DF variation. Early‐flowering accessions carried up to 225 favourable alleles. A flowering time prediction model for high‐latitude regions achieved 96% accuracy. We present a high‐quality T2T reference genome for cultivated sesame, shedding light on its origin, evolutionary history, and regional flowering time adaptation. This genome insights valuable tools for breeding programs aimed at improving yield and environmental adaptation in sesame and related crops.