This work introduces orthoSynAssign, a user-friendly, high-performance rewrite of the orthogroup refinement tool OrthoRefine, combining an intuitive Python interface with a core computing engine written in Rust, which provides a reliable and scalable framework for high-throughput phylogenomic workflows.
Abstract
Accurately identifying orthogroups is crucial for precise phylogenetic reconstruction, but clustering-based methods often generate complex, many-to-many orthogroups that include confounding paralogs. Incorporating synteny offers a robust strategy to refine these clusters into high-granularity, single-copy orthologs. We introduce orthoSynAssign, a user-friendly, high-performance rewrite of the orthogroup refinement tool OrthoRefine, combining an intuitive Python interface with a core computing engine written in Rust. This hybrid architecture ensures straightforward installation, seamless data parsing, and exceptional computational efficiency. Evaluated against the Yeast Gene Order Browser (YGOB) dataset, orthoSynAssign demonstrated outstanding performance, substantially elevating the Area Under the Precision-Recall Curve. Furthermore, multi-threading benchmarks across 193 Eurotiomycetes genomes confirmed strong scalability, drastically reducing execution runtime while maintaining a strictly bounded, thread-independent memory footprint. Ultimately, orthoSynAssign provides a reliable and scalable framework for high-throughput phylogenomic workflows.
PyiTOL validates inputs, generates 31 iTOL template schemas, performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay.
The results establish Sma3s v3 as a scalable and interpretable tool for functional annotation and re-annotation of proteomes, pangenomes, and metagenomic protein catalogues.
Alejandro Rubio, Jesús L. García-Junco Alcalá, Elisa Luque-Jiménez et al.· bioRxiv· 0 citations
Metagenomic analyses can be performed at multiple analytical levels, including read-based, assembly-based, and genome-resolved approaches, each capturing complementary biological information while introducing distinct analytical biases and trade-offs. However, existing workflows are commonly optimized for a single anal...
D. Corso, Edoardo Taccaliti, B. Barosa et al.· bioRxiv· 0 citations
ManiFasta is presented, a tool enabling users to generate standardized, reproducible, and robustly documented protein reference sets from diverse input sources and datatypes, and the value of maniFasta is highlighted in the context of salivary metaproteomics, addressing the need for a taxonomically comprehensive refere...
Christopher Handelmann, Ashley K. Miles, Yin-Yin Ye et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.