bioRxiv ScienceSearch

Biology subjects

Shah, N. H.

Publications and source records attributed to Shah, N. H..

4 recordsLinked to original sources

Deep mutational analysis reveals functional trade-offs in the sequences of EGFR autophosphorylation sites

Upon activation, the epidermal growth factor receptor (EGFR) phosphorylates tyrosine residues in its cytoplasmic tail, which triggers the binding of Src Homology 2 (SH2) and Phosphotyrosine Binding (PTB) domains and initiates downstream signaling. The sequences flanking the tyrosine residues (referred to as phosphosites) must be compatible with phosphorylation by the EGFR kinase domain and the recruitment of adapter proteins, while minimizing phosphorylation that would reduce the fidelity of signal transmission. In order to understand how phosphosite sequences encode these functions within a small set of residues, we carried out high-throughput mutational analysis of three phosphosite sequences in the EGFR tail. We used bacterial surface-display of peptides, coupled with deep sequencing, to monitor phosphorylation efficiency and the binding of the SH2 and PTB domains of the adapter proteins Grb2 and Shc1, respectively. We found that the sequences of phosphosites in the EGFR tail are restricted to a subset of the range of sequences that can be phosphorylated efficiently by EGFR. Although efficient phosphorylation by EGFR can occur with either acidic or large hydrophobic residues at the -1 position with respect to the tyrosine, hydrophobic residues are generally excluded from this position in tail sequences. The mutational data suggest that this restriction results in weaker binding to adapter proteins, but also disfavors phosphorylation by the cytoplasmic tyrosine kinases c-Src and c-Abl. Our results show how EGFR-family phosphosites achieve a trade-off between minimizing off-pathway phosphorylation while maintaining the ability to recruit the diverse complement of effectors required for downstream pathway activation.

biochemistry

Fine-tuning of substrate preferences of the Src-family kinase Lck revealed through a high-throughput specificity screen

To obtain a comprehensive map of the intrinsic specificities of tyrosine kinase domains, we developed a high-throughput method that uses bacterial surface-display and next-generation sequencing to analyze the specificity of any tyrosine kinase against a library of thousands of peptides derived from human tyrosine phosphorylation sites. Using this approach, we identified a difference in the electrostatic recognition of substrates between the cytoplasmic Src-family tyrosine kinases Lck and c-Src. This divergence likely reflects the specialization of Lck to act in concert with the tyrosine kinase ZAP-70 in T cell receptor signaling. The current understanding of substrate recognition by tyrosine kinases emphasizes the role of localization by non-catalytic domains, but our results point to the importance of direct recognition at the kinase active site in fine-tuning specificity. Our method provides a simple approach that leverages next-generation sequencing to readily map the specificity of any tyrosine kinase at the proteome level.

biochemistry

Interpretation of biological experiments changes with evolution of Gene Ontology and its annotations

Gene Ontology (GO) enrichment analysis is ubiquitously used for interpreting high throughput molecular data and generating hypotheses about underlying biological phenomena of experiments. However, the two building blocks of this analysis -- the ontology and the annotations -- evolve rapidly. We used gene signatures derived from 104 disease analyses to systematically evaluate how enrichment analysis results were affected by evolution of the GO over a decade. We found low consistency between enrichment analyses results obtained with early and more recent GO versions. Furthermore, there continues to be strong annotation bias in the GO annotations where 58% of the annotations are for 16% of the human genes. Our analysis suggests that GO evolution may have affected the interpretation and possibly reproducibility of experiments over time. Hence, researchers must exercise caution when interpreting GO enrichment analyses and should reexamine previous analyses with the most recent GO version.

bioinformatics

Integrated molecular and clinical analysis for understanding human disease relationships

Existing knowledge of human disease relationships is incomplete. To establish a comprehensive understanding of disease, we integrated transcriptome profiles of 41,000 human samples with clinical profiles of 2 million patients, across 89 diseases. Based on transcriptome data, autoimmune diseases clustered with their specific infectious triggers, and brain disorders clustered by disease class. Clinical profiles clustered diseases according to the similarity of their initial manifestation and later complications, identifying disease relationships absent in prior co-occurrence analyses. Our integrated analysis of transcriptome and clinical profiles identified overlooked, therapeutically actionable disease relationships, such as between myositis and interstitial cystitis. Our improved understanding of disease relationships will identify disease mechanisms, offer novel therapeutic targets, and create synergistic research opportunities.

bioinformatics