bioRxiv Science⌕ Search

Biology subjects

Kotha, S. R.

Publications and source records attributed to Kotha, S. R..

2 recordsLinked to original sources

Conservation of function without conservation of amino acid sequence in intrinsically disordered transcriptional activation domains

Protein function is canonically believed to be more conserved than amino acid sequence, but this idea is only well supported in folded domains, where highly diverged sequences can fold into equivalent 3D structures with identical function. Intrinsically disordered protein regions (IDRs) often experience rapid amino acid sequence divergence, but because they do not fold into stable 3D structures, it remains unknown when and how function is conserved. As a model system for studying the evolution of IDRs, we examined transcriptional activation domains, the regions of transcription factors that bind to coactivator complexes. We systematically identified activation domains on 502 homologs of the transcriptional activator Gcn4 spanning 600 MY of fungal evolution in the Ascomycota. We find that the central activation domain shows strong conservation of function without conservation of sequence. We identify the molecular mechanism for this conservation of function without conservation of sequence: evolutionary turnover (gain and loss) of acidic and aromatic residues that are important for function. We further see turnover of complete N-terminal activation domains. This turnover at two length scales confounds multiple sequence alignments, explaining why traditional comparative genomics cannot detect functional conservation of activation domains. Evolutionary turnover of key residues is likely a general mechanism for conservation of function without conservation of sequence in IDRs.

systems biology↗

The balance of acidic and hydrophobic residues predicts acidic transcriptional activation domains from protein sequence

Transcription factors activate gene expression in development, homeostasis, and stress with DNA binding domains and activation domains. Although there exist excellent computational models for predicting DNA binding domains from protein sequence (Stormo, 2013), models for predicting activation domains from protein sequence have lagged behind (Erijman et al., 2020; Ravarani et al., 2018; Sanborn et al., 2021), particularly in metazoans. We recently developed a simple and accurate predictor of acidic activation domains on human transcription factors (Staller et al., 2022). Here, we show how the accuracy of this human predictor arises from the balance between hydrophobic and acidic residues, which together are necessary for acidic activation domain function. When we combine our predictor with the predictions of neural network models trained in yeast, the intersection is more predictive than individual models, emphasizing that each approach carries orthogonal information. We synthesize these findings into a new set of activation domain predictions on human transcription factors.

systems biology↗