Compositionally Biased Regions Within Structured Protein Domains
Abstract
In proteins, tracts compositionally biased by a subset of amino acids are often linked to intrinsic disorder. Such compositionally biased regions (CBRs) are sometimes analyzed for ‘sequence complexity’ (‘information entropy’) despite being an unlikely substrate for natural selection per se, and therefore not functionally implicated. Here, an algorithmic strategy applying compositional bias detection was designed to capture the wide diversity of CBRs in structured protein domains (termed ‘sCBRs’), ranging from trihomopeptides to >200 residues, with conservation and partner binding as functional lenses. sCBRs are common, with about 1/4th of domains harbouring them, but domains dominated by sCBRs over >50% of their lengths are rare (~1 in 200). Despite general assumptions, very short sCBRs are highly significantly sequence-conserved, and associated with ligand binding, even when common nucleotide/phosphate-binding or glycine-rich cases are disregarded. However, regardless of length, ~50% of cases are evolutionarily dynamic, undergoing clade-specific expansion/contraction. Protein-binding associations include aversions for short (≤16 residues) valine-rich regions in protein interfaces, and enrichments of longer alanine-rich cases (>16 residues). Only ~5% of cases are (at least partly) in intrinsically disordered loops, and are significantly shorter than sCBRs generally. Functional implications of sCBRs are discussed with many examples. The sCBR data might help with hypothesis generation and protein design/engineering.