bioRxiv Science⌕ Search

Biology subjects

Hadarovich, A.

Publications and source records attributed to Hadarovich, A..

3 recordsLinked to original sources

Characterization of Alternative Splicing During Mammalian Brain Development Reveals the Magnitude of Isoform Diversity and its Effects on Protein Conformational Changes

Regulation of gene expression is critical for fate commitment of stem and progenitor cells during tissue formation. In the context of mammalian brain development, a plethora of studies have described how changes in the expression of individual genes characterize cell types across ontogeny and phylogeny. However, little attention was paid to the fact that different transcripts can arise from any given gene through alternative splicing (AS). Considered a key mechanism expanding transcriptome diversity during evolution, assessing the full potential of AS on isoform diversity and protein function has been notoriously difficult. Here we capitalize on the use of a validated reporter mouse line to isolate neural stem cells, neurogenic progenitors and neurons during corticogenesis and combine the use of short- and long-read sequencing to reconstruct the full transcriptome diversity characterizing neurogenic commitment. Extending available transcriptional profiles of the mammalian brain by nearly 50,000 new isoforms, we found that neurogenic commitment is characterized by a progressive increase in exon inclusion resulting in the profound remodeling of the transcriptional profile of specific cortical cell types. Most importantly, we computationally infer the biological significance of AS on protein structure by using AlphaFold2 and revealing how radical protein conformational changes can arise from subtle changes in isoforms sequence. Together, our study reveals that AS has a greater potential to impact protein diversity and function than previously thought independently from changes in gene expression.

developmental biology↗

SHARK enables homology assessment in unalignable anddisordered sequences

Intrinsically disordered regions (IDRs) are structurally flexible protein segments with regulatory functions in multiple contexts, such as in the assembly of biomolecular condensates. Since IDRs undergo more rapid evolution than ordered regions, identifying homology of such poorly conserved regions remains challenging for state-of-the-art alignment-based methods that rely on position-specific conservation of residues. Thus, systematic functional annotation and evolutionary analysis of IDRs have been limited, despite comprising [~]21% of proteins. To accurately assess homology between unalignable sequences, we developed an alignment-free sequence comparison algorithm, SHARK (Similarity/Homology Assessment by Relating K-mers). We trained SHARK-dive, a machine learning homology classifier, which achieved superior performance to standard alignment in assessing homology in unalignable sequences, and correctly identified dissimilar IDRs capable of functional rescue in IDR-replacement experiments reported in the literature. SHARK-dive not only predicts functionally similar IDRs, but also identifies cryptic sequence properties and motifs that drive remote homology, thereby facilitating systematic analysis and functional annotation of the unalignable protein universe.

bioinformatics↗

PICNIC identifies condensate-forming proteins across organisms

Biomolecular condensates are membraneless organelles that can concentrate hundreds of different proteins to operate essential biological functions. However, accurate identification of their components remains challenging and biased towards proteins with high structural disorder content with focus on self-phase separating (driver) proteins. Here, we present a machine learning algorithm, PICNIC (Proteins Involved in CoNdensates In Cells) to classify proteins involved in biomolecular condensates regardless of their role in condensate formation. PICNIC successfully predicts condensate members by identifying amino acid patterns in the protein sequence and structure in addition to the intrinsic disorder and outperforms previous methods. We performed extensive experimental validation in cellulo and demonstrated that PICNIC accurately predicts 21 out of 24 condensate-forming proteins regardless of their structural disorder content. Even though increasing disorder content was associated with organismal complexity, we found no correlation between predicted condensate proteome content and disorder content across organisms. Overall, we applied a novel machine learning classifier to interrogate condensate components at single protein and whole-proteome levels across the tree of life (picnic.cd-code.org).

bioinformatics↗