bioRxiv Science⌕ Search

Biology subjects

Eicholt, L. A.

Publications and source records attributed to Eicholt, L. A..

3 recordsLinked to original sources

Sequence, Structure and Functional space of Drosophila de novo proteins

During de novo emergence, new protein coding genes emerge from previously non-genic sequences. The de novo proteins they encode are dissimilar in composition and predicted biochemical properties to conserved proteins. However, many functional de novo proteins indeed exist. Both identification of functional de novo proteins and their structural characterisation are experimentally laborious. To identify functional and structured de novo proteins in silico, we applied recently developed machine learning based tools and refined the results for de novo proteins. We found that most de novo proteins are indeed different from conserved proteins both in their structure and sequence. However, some de novo proteins are predicted to adopt known protein folds, participate in cellular reactions, and to form biomolecular condensates. Apart from broadening our understanding of de novo protein evolution, our study also provides a large set of testable hypotheses for focused experimental studies on structure and function of de novo proteins in Drosophila.

bioinformatics↗

Random, de novo and conserved proteins: How structure and disorder predictors perform differently

Understanding the emergence and structural characteristics of de novo and random proteins is crucial for unraveling protein evolution and designing novel enzymes. However, experimental determination of their structures remains challenging. Recent advancements in protein structure prediction, particularly with AlphaFold2 (AF2), have expanded our knowledge of protein structures, but their applicability to de novo and random proteins is unclear. In this study, we investigate the structural predictions and confidence scores of AF2 and protein language model (pLM)-based predictor ESMFold for de novo, random, and conserved proteins. We find that the structural predictions for de novo and random proteins differ significantly from conserved proteins. Interestingly, a positive correlation between disorder and confidence scores (pLDDT) is observed for de novo and random proteins, in contrast to the negative correlation observed for conserved proteins. Furthermore, the performance of structure predictors for de novo and random proteins is hampered by the lack of sequence identity. We also observe varying predicted disorder among different sequence length quartiles for random proteins, suggesting an influence of sequence length on disorder predictions. In conclusion, while structure predictors provide initial insights into the structural composition of de novo and random proteins, their accuracy and applicability to such proteins remain limited. Experimental determination of their structures is necessary for a comprehensive understanding. The positive correlation between disorder and pLDDT could imply a potential for conditional folding and transient binding interactions of de novo and random proteins.

bioinformatics↗

Chaperones facilitate heterologous expression of naturally evolved putative de novo proteins

Over the past decade, evidence has accumulated that new protein coding genes can emerge de novo from previously non-coding DNA. Most studies have focused on large scale computational predictions of de novo protein coding genes across a wide range of organisms. In contrast, experimental data concerning the folding and function of de novo proteins is scarce. This might be due to difficulties in handling de novo proteins in vitro, as most are predicted to be short and disordered. Here we propose a guideline for the effective expression of eukaryotic de novo proteins in Escherichia coli. We used 11 sequences from Drosophila melanogaster and 10 from Homo sapiens, that are predicted de novo proteins from former studies, for heterologous expression. The candidate de novo proteins have varying secondary structure and disorder content. Using multiple combinations of purification tags, E. coli expression strains and chaperone systems, we were able to increase the number of solubly expressed putative de novo proteins from 30 % to 62 %. Our findings indicate that the best combination for expressing putative de novo proteins in E.coli is a GST-tag with T7 Express cells and co-expressed chaperones. We found that, overall, proteins with higher predicted disorder were easier to express.

molecular biology↗