These findings reconstruct the evolutionary emergence of the DGF-1 architecture from pre-existing structural modules and provide a framework for one of the largest and most enigmatic gene families in trypanosomatids.
Abstract
Large multigene families are a hallmark of trypanosomatid genomes and play major roles in genome plasticity and parasite adaptation. Among them, the Dispersed Gene Family 1 (DGF-1) of Trypanosoma cruzi remains one of the least understood because of its large size, repetitiveness, and fragmented annotations. Here, using a high-quality long-read assembly of T. cruzi, we performed comprehensive re-annotation and structural characterization of DGF-1. We identified seven tandemly repeated structural modules, here termed ribs, that form the extracellular region. This is followed by a distinct C-terminal membrane-associated region containing a conserved hat domain and multiple transmembrane helices. These domains are highly conserved across paralogous copies and the 7 ribs cluster according to positional identity rather than gene origin, indicating that the canonical seven-rib architecture predates the expansion of the family. DGF-1 proteins are present in the early-branching trypanosomatid species Paratrypanosoma confusum and several other trypanosomatids but absent in Leishmania and African trypanosomes, consistent with multiple independent secondary losses. Comparative analyses across Euglenozoa indicate that the DGF-1 architecture did not originate in free-living bodonids. Instead, Bodo saltans contains proteins with rib-like domains and others containing the membrane-associated region, whereas the first complete DGF-1 architecture appears in P. confusum. Phylogenetic and comparative genomic analyses indicate that the family subsequently underwent lineage-specific losses and independent expansions, the largest expansion occurring in T. cruzi. Together, these findings reconstruct the evolutionary emergence of the DGF-1 architecture from pre-existing structural modules and provide a framework for one of the largest and most enigmatic gene families in trypanosomatids.
Transcriptional regulation of protein-coding genes is a hallmark of eukaryotic gene expression. Yet, a group of parasitic protists, trypanosomatids, appear to lack this capability. Here, we analyzed genomic, nascent transcriptomic, RNA polymerase occupancy and gene organization data to reconstruct the evolutionary origin and biological consequences of their unusual regulatory strategy. Across 59 Discoba protists, we show stepwise evolutionary erosion of conventional transcription regulation components in trypanosomatida lineage, including gene consolidation into polycistronic transcription units (PTUs), shortening of intra-PTU non-coding regions, and depletion of transcription factors and their enriched DNA-binding motifs. This transition was associated with near-constitutive expression of most genes, indicating broad loss of conditional gene expression. However, trypanosomatids retain some differential regulation at the PTU level, with >70% PTUs featuring significantly different nascent transcription than their neighbors or resident chromosomes. Moreover, gene expression is not uniform within PTUs: nascent transcription, translation efficiency, and protein abundance progressively decline with distance from the transcription start site. Consistent with this architecture-encoded regulatory logic, co-complex subunits and co-pathway enzymes preferentially occupy adjacent positions within PTUs despite each PTU’s overall functional heterogeneity. These findings reveal an evolutionary shift from gene-specific transcriptional regulation toward a regime where genome architecture becomes a regulator of gene expression.
Saurav Mallik, Meir Sylman, Moshe Kafri et al.· bioRxiv· 0 citations
Genes encoding intracellular nucleotide-binding site leucine-rich repeat (NBS-LRR) receptors represent the largest class of resistance (R) genes in plants, yet their evolutionary trajectory in trees remains poorly understood. Using high-quality long-read genome assemblies from eight Myrtaceae species, we identified 15,792 NBS-encoding genes with threefold variation in gene content across species. To investigate potential decoy domains for effector proteins of pathogens, we parsed the annotated NBS-encoding genes and determined 181 unique non-canonical domains. Notably, we observed frequent integration of Jacalin domains into TIR-NBS proteins. This novel gene family, named TNJ, is hypothesized to represent a new R gene class. Further exploration of TNJ sequence analysis shows up to seven Jacalin domains per protein and forming a monophyletic clade; however, AF3 modelling confirmed only six domains. Conserved residues in the TIR domain and functional motifs within the NB-ARC domain supports TNJ’s potential role in immune signalling. Most hypervariable sites and positively selected sites detected to be surface-exposed were clustered in the Jacalin region of TNJ, suggesting that surfaces of Jacalin domain may harbour residues determining pathogen recognition specificity. These findings support TNJ as a potentially new class of R genes in Myrtaceae with Jacalin as a replacement for LRR.
T. Tolessa, Scott Ferguson, Ashley W. Jones et al.· bioRxiv· 0 citations
The mitochondrial control region (CR) is the largest non-coding region in the vertebrate mitogenome and contains essential elements for replication and transcription. Despite its functional relevance, its evolutionary dynamics remain poorly understood. Here, we analyzed 5,235 complete vertebrate CRs spanning 11 classes to investigate how conserved sequences blocks (CSBs) and Extended Termination-Associated sequences (ETAS) shaped CR evolution. We hypothesized that CR length is positively associated with repeat accumulation, with tetrapods exhibiting longer and more complex CRs than fishes, while core elements remain conserved. Our analyses revealed marked inter- and intra-class variability, with longer CRs in tetrapods (1,283.27 ± 489.6 bp) than in fishes (969.25 ± 239.5 bp). Duplication events were restricted to tetrapods, especially birds and reptiles. Nucleotide composition was heterogeneous among orders, and structural divergence of CSBs was inferred across lineages. Repetitive elements were present in ~43% of CRs, with their abundance strongly correlated with CR length. Importantly, longer CRs were associated with higher GC content and greater variation in copy number of ETAS and CSBs. These results demonstrate that mitogenome CR expansion in vertebrates is largely driven by repeat proliferation, whereas key motifs required for replication and transcription are retained. We further identify lineage-specific trends, including pronounced CR elongation in amphibians and reptiles, contrasted with progressive reduction and simplification in birds and mammals. Our study provides the first comprehensive comparative framework of vertebrate CR evolution, highlighting how repetitive elements, conserved motifs, and nucleotide composition jointly contribute to both functional regulation and lineage-specific diversification.
Mauricio Ochoa Capera, Natalia S Medina, Paula Montaña-Lozano et al.· PLoS ONE· 0 citations
S-domain receptor-like kinases (S-RLKs) represent a typical RLK subfamily, which plays key roles in various biological processes in plants. However, the genome-wide evolutionary and functional differentiation of this family remains unclear in rice. In the present study, a comprehensive computational analysis was employed and identified 109 S-RLKs in 9311 genome and 103 S-RLKs in Nipponbare (Nip) genome. The S-RLKs were unevenly distributed across 12 chromosomes. Bioinformatics analysis indicate large-scale gene duplication and family expansion may be the main driving forces for the expansion of S-RLK family members. Although the S-RLKs are highly conserved between the two subspecies, it remains unknown whether these members have undergone functional differentiation during long-term evolution. Given the differences of roots and nitrogen (N) utilization in 9311 and Nip, we focused on the S-RLKs which exhibit different expressions in roots. OsNRS1 (Nitrogen-responsive Root S-domain kinase 1) was selected due to its expression in the roots of 9311 notably more than that in Nip. Phenotypic analysis showed that OsNRS1 CRISPR/Cas9 mutants in 9311 significantly inhibited the length of primary root and total root, and N accumulation, whereas OsNRS1 CRISPR/Cas9 mutants in Nip showed significant functional divergence. In-depth research revealed that these functional divergence of OsNRS1 may be caused by the sequence variation of an auxin response element AuxRR in its promoter. Collectively, this study not only provides novel insights into the evolution of the S-RLK gene family, but also identifies OsNRS1 as a potential key target for the genetic improvement of root and N utilization in rice.
Cong Chen, Jiajia Jin, Shujie Shi et al.· Plant physiology and biochem...· 0 citations
ABSTRACT Porcine teschoviruses (PTVs) are swine-specific picornaviruses associated with neurological and systemic disease, ranging from severe paralytic Teschen disease to milder forms of porcine encephalomyelitis, known as Talfan disease. Despite their global distribution and veterinary relevance, complete genomic information and molecular tools for PTVs are not available. Here, we report the first complete genome sequences of PTV-A11 (strain Dresden) and a recent PTV-A field isolate (strain Gi2020_1). Both genomes, each approximately 7.2 kb in size, include a previously uncharacterized 5′-terminal region upstream of the poly(C) tract (S-segment), thereby elucidating the complete genomic architecture of PTVs for the first time. The conserved 117–118 nucleotide S-segment closes a long-standing gap in teschovirus genomics and is essential for viral replication. Using reverse genetics, we established infectious molecular clones of both strains that recapitulate the properties of their parental strains. These systems enabled functional analyses of viral gene products, including the demonstration that the leader protein is dispensable for genome replication and morphogenesis but contributes to efficient virus growth. In addition, we developed a replication-competent subgenomic replicon and engineered fluorescent reporter viruses, including a stable mCherry-expressing virus that supports robust infection analysis and allows quantitative protein expression measurements. Together, these findings define the complete PTV genome organization and provide a versatile molecular toolbox for studying teschovirus replication, pathogenesis, and control. IMPORTANCE Teschen disease was once a devastating neurological disease of swine caused by highly virulent porcine teschovirus strains (PTVs) but has become rare following their eradication. Current control strategies base on hygiene measures and herd-specific vaccination providing limited protection and leaving swine populations vulnerable to the (re-)emergence of neuroinvasive PTVs. Less virulent strains remain endemic worldwide and continue to impair animal health, welfare, and production efficiency. Progress toward broadly effective vaccines and antivirals has been constrained by incomplete knowledge of the PTV genome. Here, we identify and functionally define the previously unrecognized S-segment that completes the 5′ UTR of PTVs. We further established infectious cDNA clones and developed subgenomic replicons, leaderless viruses, and fluorescent reporter viruses. Our tools enable direct genetic manipulation of PTVs. They provide a platform for mechanistic studies of viral replication, attenuation, and antigen design supporting rational development of vaccines and antiviral strategies against emerging PTVs. Teschen disease was once a devastating neurological disease of swine caused by highly virulent porcine teschovirus strains (PTVs) but has become rare following their eradication. Current control strategies base on hygiene measures and herd-specific vaccination providing limited protection and leaving swine populations vulnerable to the (re-)emergence of neuroinvasive PTVs. Less virulent strains remain endemic worldwide and continue to impair animal health, welfare, and production efficiency. Progress toward broadly effective vaccines and antivirals has been constrained by incomplete knowledge of the PTV genome. Here, we identify and functionally define the previously unrecognized S-segment that completes the 5′ UTR of PTVs. We further established infectious cDNA clones and developed subgenomic replicons, leaderless viruses, and fluorescent reporter viruses. Our tools enable direct genetic manipulation of PTVs. They provide a platform for mechanistic studies of viral replication, attenuation, and antigen design supporting rational development of vaccines and antiviral strategies against emerging PTVs.
Sebastian Affeldt, Sandra Barth, Stephanie Schlimbach et al.· Journal of Virology· 0 citations
Protein phosphatase 2C (PP2C) proteins are central regulators of plant signaling and stress responses, yet their macroevolutionary origin and diversification across the plant kingdom remain incompletely resolved. In this study, we performed a large-scale genome-wide analysis of the PP2C gene family using 402 representative plant genomes and integrated phylogenetic, duplication type, motif, expression, pangenome and selection-pressure analyses. A total of 36 960 PP2C genes were identified, showing substantial lineage-specific copy-number variation and marked expansion in angiosperms. Phylogenetic reconstruction classified plant PP2Cs into 12 subfamilies within three major clades and indicated that most subfamilies originated before the establishment of land plants, whereas angiosperm diversification mainly involved quantitative expansion rather than the emergence of new subfamilies. Duplication analysis revealed that dispersed and WGD/segmental duplication were the principal forces driving PP2C expansion, while phylogenetic tree reconciliation suggested extensive lineage-specific retention and loss after ancestral duplication events. Conserved motif analysis showed strong preservation of the catalytic scaffold, especially Motif1–Motif3 and Motif5, together with flexible remodeling of peripheral motifs. Cross-species transcriptome profiling in seven angiosperms revealed phylogenetically structured expression divergence: Clade II retained broad hormone, tissue and stress responsiveness, whereas Clades I and III displayed more condition- and organ-specific specialization. In 18 Brassica rapa accessions, 2478 PP2C genes were identified, most of which belonged to core orthogroups and syntenic regions. Ka/Ks analysis further indicated predominant purifying selection, with relaxed constraints in non-core genes. Collectively, these results provide a comprehensive evolutionary framework for PP2C functional diversification and candidate resources for stress-resistance improvement in horticultural crops.
Quanlong Liu, Jianbin Quan, Yuhua Cui et al.· Horticulture Research· 0 citations