Theoretical estimates on the expected number of mutations needed to reconstruct clonal lineage trees
Abstract
Abstract Motivation Phylogenetics faces a growing challenge from increasingly large and complicated data sets enabled by ever-improving sequencing technologies. The issue is particularly acute for somatic evolution studies, such as cancer cell lineages, where single-cell data sets may include tens of thousands of mutations in hundreds of thousands of genetically distinct cells. Simultaneously, the biological complexity of somatic evolution has led to complex phylogeny methods that struggle to scale to even modest data sizes. Results We explore the theoretical and empirical basis for one strategy for managing these large data sets: subsampling mutations for the computationally challenging phylogeny problem followed by faster placement of mutations on a putatively known guide tree. We specifically focus on determining the number of mutations sufficient to recover the true phylogeny at some level of resolution with high probability. We theoretically analyze variants of several common models that underlie popular tools for building clonal lineage trees. We further test these bounds through simulations of these models, extensions of them, and real biological datasets. The results suggest that modest numbers of mutations suffice to reconstruct clonal tree topologies for typical numbers of clones, supporting subsampling as a general strategy for managing the challenges of ever-growing data. Availability and implementation All analysis code and scripts used for data simulation are implemented in Python 3 and available at https://github.com/CMUSchwartzLab/mutation-subsampling.git.