bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,207 records · Page 67Linked to original sources

Growing Glycans in Rosetta: Accurate de novo glycan modeling, density fitting, and rational sequon design

Carbohydrates and glycoproteins modulate key biological functions. Computational approaches inform function to aid in carbohydrate structure prediction, structure determination, and design. However, experimental structure determination of sugar polymers is notoriously difficult as glycans can sample a wide range of low energy conformations, thus limiting the study of glycan-mediated molecular interactions. In this work, we expanded the RosettaCarbohydrate framework, developed and benchmarked effective tools for glycan modeling and design, and extended the Rosetta software suite to better aid in structural analysis and benchmarking tasks through the SimpleMetrics framework. We developed a glycan-modeling algorithm, GlycanTreeModeler, that computationally builds glycans layer-by-layer, using adaptive kernel density estimates (KDE) of common glycan conformations derived from data in the Protein Data Bank (PDB) and from quantum mechanics (QM) calculations. After a rigorous optimization of kinematic and energetic considerations to improve near-native sampling enrichment and decoy discrimination, GlycanTreeModeler was benchmarked on a test set of diverse glycan structures, or "trees". Structures predicted by GlycanTreeModeler agreed with native structures at high accuracy for both de novo modeling and experimental density-guided building. GlycanTreeModeler algorithms and associated tools were employed to design de novo glycan trees into a protein nanoparticle vaccine that are able to direct the immune response by shielding regions of the scaffold from antibody recognition. This work will inform glycoprotein model prediction, aid in both X-ray and electron microscopy density solutions and refinement, and help lead the way towards a new era of computational glycobiology.

molecular biology↗

Comprehensive structural and interactome analysis reveals 1 novel interactions and protein binding sites in miR-675: a non-coding RNA critically involved in multiple diseases

miR-675 is a microRNA expressed from exon 1 of H19 long non-coding RNA. H19 lncRNA is temporally expressed in humans and atypical expression of miR-675 has been linked with several diseases and disorders. To execute its function inside the cell, miR-675 is folded into a particular conformation which aids in its interaction with several other biological molecules. However, the exact folding dynamics of miR-675 and its complete interaction map are currently unknown. Moreover, how H19 lncRNA and miR-675 crosstalk and modulate each others activities is also unclear. Detailed structural analysis of miR-675 in this study determines its conformation and identifies novel protein binding sites on miR-675 which can make it an excellent therapeutic target against numerous diseases. Mapping of the interactome identified some of known and unknown interactors of miR-675 which aid in expanding our repertoire of miR-675 involved pathways in the cell. This analysis also identified some of the previously unknown and yet to be characterised proteins as probable interactors of miR-675. Structural and base pair conservation analysis between H19 lncRNA and miR-675 results in structural transformations in miR-675 thus describing the earlier unknown mechanism of interaction between these two molecules. Comprehensively, this study details the conformation of miR-675, its interacting biological partners and explains its relationship with H19 lncRNA which can be interpreted to understand the role of miR-675 in the development and progression of various diseases.

molecular biology↗

High-throughput, microscopy-based screening, and quantification of genetic elements

Synthetic biology relies on the screening and quantification of genetic components to assemble sophisticated gene circuits with specific functions. Microscopy is powerful tool for characterizing complex cellular phenotypes with increasing spatial and temporal resolution to library screening of genetic elements. Microscopy-based assays are powerful tools for characterizing cellular phenotypes with spatial and temporal resolution, and can be applied to large-scale samples for library screening of genetic elements. However, strategies for high-throughput microscopy experiments remain limited. Here, we present a high-throughput, microscopy-based platform that can simultaneously complete the preparation of an 8x12-well agarose pads plate, allowing for the screening of 96 independent strains or experimental conditions in a single experiment. Using this platform, we screened a library of natural intrinsic promoters from Pseudomonas aeruginosa and identified a small subset of robust promoters that drives stable levels of gene expression under varying growth conditions. Additionally, the platform allowed for single-cell measurement of genetic elements over time, enabling the identification of complex and dynamic phenotypes to map genotype in high-throughput. We expected that the platform could be employed to accelerate the identification and characterization of genetic elements in various biological systems, as well as to understand the relationship between cellular phenotypes and internal states, including genotypes and gene expression programs. Impact statementThe high-throughput microscopy-based platform, presented in this study, enables efficient screening of 96 independent strains or experimental conditions in a single experiment, facilitating the rapid identification of genetic elements with desirable features, thereby advancing synthetic biology. The robust promoters identified through this platform, which provide predictable and consistent control over gene expression under varying growth conditions, can be utilized as reliable tools to regulate gene expression in various biological applications, including synthetic biology, metabolic engineering, and gene therapy, where consistent system performance is required.

molecular biology↗

Structure-based probe reveals the presence of large transthyretin aggregates in plasma of ATTR amyloidosis patients

ATTR amyloidosis is a relentlessly progressive disease caused by the misfolding and systemic accumulation of amyloidogenic transthyretin into amyloid fibrils. These fibrils cause diverse clinical phenotypes, mainly cardiomyopathy and/or polyneuropathy. Little is known about the aggregation of transthyretin during disease development and whether this has implications for diagnosis and treatment. Using the cryogenic electron microscopy structures of mature ATTR fibrils, we developed a peptide probe for fibril detection. With this probe, we have identified previously unknown aggregated transthyretin species in plasma of patients with ATTR amyloidosis. These species are large, non-native, and distinct from monomeric and tetrameric transthyretin. Observations from our study open many questions about the biology of ATTR amyloidosis and reveals a potential diagnostic and therapeutic target.

molecular biology↗

A scalable CRISPR-Cas9 gene editing system facilitates CRISPR screens in the malaria parasite Plasmodium berghei

Many Plasmodium genes remain uncharacterised due to low genetic tractability. Previous large scale knockout screens have only been able to target about half of the genome in the more genetically tractable rodent malaria parasite Plasmodium berghei. To overcome this limitation, we have developed a scalable CRISPR system called PbHiT, which uses a single cloning step to generate targeting vectors with 100 bp homology arms physically linked to a guide RNA (gRNA) that effectively integrate into the target locus. We show that PbHiT coupled with gRNA sequencing robustly recapitulates known knockout mutant phenotypes in pooled transfections. Furthermore, we provide vector designs and sequences to target the entire P. berghei genome and scale-up vector production using a pooled ligation approach. This work presents for the first time a tool for high-throughput CRISPR screens in Plasmodium for studying the parasites biology at scale.

molecular biology↗

Profiling of yeast Saccharomyces cerevisiae mitochondrial AMPylome reveals a regulation of ATP synthase coupling trough subunit delta

The adenylation (AMPylation) of proteins as a posttranslational modification is used by bacteria during infection of host cells. These new virulence factors - AMPylases mainly belonging to the FIC domain containing proteins and constitute a potential drug target. Human FIC protein (HYPE) controls the activity of BiP chaperone under endoplasmic reticulum stress. No FIC family proteins have yet been identified in yeast Saccharomyces cerevisiae. The second family of AMPylases are SelO proteins which control the redox homeostasis in mitochondria and chloroplasts. We describe here the first global screening of AMPylated proteins in yeast S. cerevisiae mitochondrial proteome from wild type and SelO (Fmp40) lacking cells. Through quantitative mass-spectrometry-based proteomics, we identified a total of 169 AMPylated proteins in mitochondria while AMPylated peptides of 115 proteins were identified in fmp40{Delta} mitochondria, indicating on the presence of another, besides Fmp40, not yet identified AMPylase in yeast. We confirmed AMPylation of Atp1, Atp2, Atp3 and Atp16 subunits of mitochondrial ATP synthase by western blotting. Interestingly, we found AMPylation and phosphorylation of many residues, what indicates on the complex regulation of the ATP synthase activity. We confirmed the importance of one of such residues in Atp16, showing that its post-translational modification serves to regulate ATP synthase and OXPHOS coupling in both fermentative and respiratory growth conditions. This regulation serves to maintain the proper potential of the inner mitochondrial membrane, particularly under conditions of fermentative growth. This dataset represents the first library of AMPylated mitochondrial yeast proteins reported to date and supplements the AMPylome of human chronic lymphocytic leukemia cell line from human HYPE containing and HYPE lacking cells. The data represents a foundation for substrate specific investigations that can ultimately decipher the biological role of the AMPylation in the mitochondria.

molecular biology↗

SLAM-ITseq: Sequencing cell type-specific transcriptomes without cell sorting

Cell type-specific transcriptome analysis is an essential tool in understanding biological processes but can be challenging due to the limits of microdissection or fluorescence-activated cell sorting (FACS). Here, we report a novel in vivo sequencing method, which captures the transcriptome of a specific type of cells in a tissue without prior cellular or molecular sorting. SLAM-ITseq provides an accurate snapshot of the transcriptional state in vivo.

molecular biology↗

PIWI Proteins Act at Multiple Steps in the Production of Their Own Guides

In animals, piRNAs guide PIWI-proteins to silence transposons and regulate gene expression. The mechanisms for making piRNAs have been proposed to differ among cell types, tissues, and animals. Our data instead suggest a single model that explains piRNA production in most animals. piRNAs initiate piRNA production by guiding PIWI proteins to slice precursor transcripts. Next, PIWI proteins direct the stepwise fragmentation of the sliced precursor transcripts, yielding tail-to-head strings of phased pre-piRNAs. Our analyses detect evidence for this piRNA biogenesis strategy across an evolutionarily broad range of animals including humans. Thus, PIWI proteins initiate and sustain piRNA biogenesis by the same mechanism in species whose last common ancestor predates the branching of most animal lineages. The unified model places PIWI-clade Argonautes at the center of piRNA biology and suggests that the ancestral animal--the Urmetazoan--used PIWI proteins both to generate piRNA guides and to execute piRNA function.

molecular biology↗

DDX3 depletion selectively represses translation of structured mRNAs

DDX3 is an RNA chaperone of the DEAD-box family that regulates translation. Ded1, the yeast ortholog of DDX3, is a global regulator of translation, whereas DDX3 is thought to preferentially affect a subset of mRNAs. However, the set of mRNAs that are regulated by DDX3 are unknown, along with the relationship between DDX3 binding and activity. Here, we use ribosome profiling, RNA-seq, and PAR-CLIP to define the set of mRNAs that are regulated by DDX3 in human cells. We find that while DDX3 binds highly expressed mRNAs, depletion of DDX3 particularly affects the translation of a small subset of the transcriptome. We further find that DDX3 binds a site on helix 16 of the human ribosome, placing it immediately adjacent to the mRNA entry channel. Translation changes caused by depleting DDX3 levels or expressing an inactive point mutation are different, consistent with different association of these genetic variant types with disease. Taken together, this work defines the subset of the transcriptome that is responsive to DDX3 inhibition, with relevance for basic biology and disease states where DDX3 is altered.

molecular biology↗

Identification of a carbohydrate-recognition motif of purinergic receptors

As a major class of biomolecules, carbohydrates play indispensable roles in various biological processes. However, it remains largely unknown how carbohydrates directly modulate important drug targets, such as G-protein coupled receptors (GPCRs). Here, we employed P2Y purinoceptor 14 (P2Y14), a drug target for inflammation and immune responses, to uncover the sugar nucleotide activation of GPCRs. Integrating molecular dynamics simulation with functional study, we identified the uridine diphosphate (UDP)-sugar-binding site on P2Y14, and revealed that a UDP-glucose might activate the receptor by bridging the transmembrane helices (TM) 2 and 7. Between TM2 and TM7 of P2Y14, a conserved salt bridging chain (K2.60-D2.64-K7.35-E7.36, KDKE) was identified to distinguish different UDP-sugars, including UDP-glucose, UDP-galactose, UDP-glucuronic acid and UDP-N-acetylglucosamine. We identified the KDKE chain as a conserved functional motif of sugar binding for both P2Y14 and P2Y purinoceptor 12 (P2Y12), and then designed three sugar nucleotides as agonists of P2Y12. These results not only expand our understanding for activation of purinergic receptors but also provide insights for the carbohydrate drug development for GPCRs.

molecular biology↗

GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations

Understanding complex graph-structured data is a cornerstone of modern research in fields like cheminformatics and bioinformatics, where molecules and biological systems are naturally represented as graphs. However, traditional graph neural networks (GNNs) often fall short by focusing mainly on node features while overlooking the rich information encoded in edges. To bridge this gap, we present GraFusionNet, a framework designed to integrate node, edge, and molecular-level semantic features for enhanced graph classification. By employing a dual-graph autoencoder, GraFusionNet transforms edges into nodes via a line graph conversion, enabling it to capture intricate relationships within the graph structure. Additionally, the incorporation of Chem-BERT embeddings introduces semantic molecular insights, creating a comprehensive feature representation that combines structural and contextual information. Our experiments on benchmark datasets, such as Tox21 and HIV, highlight GraFusionNets superior performance in tasks like toxicity prediction, significantly surpassing traditional models. By providing a holistic approach to graph data analysis, GraFusion-Net sets a new standard in leveraging multi-dimensional features for complex predictive tasks. CCS CONCEPTSO_LIComputing methodologies [->] Neural networks. C_LI ACM Reference FormatMd Toki Tahmid, Tanjeem Azwad Zaman, and Mohammad Saifur Rahman. 2018. GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym XX). ACM, New York, NY, USA, 9 pages. https://doi.org/XXXXXXX.XXXXXXX

molecular biology↗

DNA2 and MSH2 activity collectively mediate chemically stabilized G4 for efficient telomere replication

G-quadruplexes (G4s) are widely existing stable DNA secondary structures in mammalian cells. A long-standing hypothesis is that timely resolution of G4s is needed for efficient and faithful DNA replication. In vitro, G4s may be unwound by helicases or alternatively resolved via DNA2 nuclease mediated G4 cleavage. However, little is known about the biological significance and regulatory mechanism of the DNA2-mediated G4 removal pathway. Here, we report that DNA2 deficiency or its chemical inhibition leads to a significant accumulation of G4s and stalled replication forks at telomeres, which is demonstrated by a high-resolution technology: Single molecular analysis of replicating DNA (SMARD). We further identify that the DNA repair complex MutS (MSH2-MSH6) binds G4s and stimulates G4 resolution via DNA2-mediated G4 excision. MSH2 deficiency, like DNA2 deficiency or inhibition, causes G4 accumulation and defective telomere replication. Meanwhile, G4-stabilizing environmental compounds block G4 unwinding by helicases but not G4 cleavage by DNA2. Consequently, G4 stabilizers impair telomere replication and cause telomere instabilities, especially in cells deficient in DNA2 or MSH2.

molecular biology↗

Nucleic Acid Capture from Human Blood Plasma Uncovers G-Quadruplex Structures

Ultrashort (US) cell-free DNA (cfDNA) is a population of approximately 50-nucleotide single-stranded DNA molecules in human plasma.1-3 Although it holds biological and diagnostic potential, US cfDNA escapes detection by conventional double-stranded library preparation methods.1-3 Independent studies have linked US cfDNA to regulatory genomic regions and to sequences predicted to form noncanonical structures, prompting the hypothesis that higher-order DNA structure contributes to its molecular properties. This interpretation, however, has so far rested on computational prediction rather than experimental evidence. Here we test this hypothesis using computational, biophysical and biochemical approaches. In silico size-selected US cfDNA from 20 healthy donors was selectively enriched at putative quadruplex sequences (PQS) that overlap both experimentally observed quadruplex sequences and accessible chromatin of blood cells. Synthetic oligonucleotides corresponding to the most enriched of these loci adopted predominantly parallel G-quadruplex (G4) structures, as revealed by circular dichroism. Endogenous nucleic acids captured directly from pooled plasma by poly(A)-tailing and immobilization, without extraction, denaturation or annealing at any step, displayed folded G4 structures. Two orthogonal probes detected these structures: the BG4 antibody and the fluorogenic ligand N-methyl mesoporphyrin IX. Reciprocal competition with a third, chemically unrelated G4 ligand, pyridostatin, confirmed the signal. Together, these experiments provide direct experimental evidence for the presence of folded G4 structures in human blood plasma.

molecular biology↗

Fe(III) heme sets an activation threshold for processing distinct groups of pri-miRNAs in mammalian cells

The essential biological cofactor heme is synthesized in cells in the Fe(II) form. Oxidized Fe(III) heme is specifically required for processing primary transcripts of microRNAs (pri-miRNAs) by the RNA-binding protein DGCR8, a core component of the Microprocessor complex. It is unknown how readily available Fe(III) heme is in the largely reducing environment in human cells and how changes in cellular Fe(III) heme availability alter microRNA (miRNA) expression. Here we address the first question by characterizing DGCR8 mutants with various degrees of deficiency in heme-binding. We observed a strikingly simple correlation between Fe(III) heme affinity in vitro and the Microprocessor activity in HeLa cells, with the heme affinity threshold for activation estimated to be between 0.6-5 pM under typical cell culture conditions. The threshold is strongly influenced by cellular heme synthesis and uptake. We suggest that the threshold reflects a labile Fe(III) heme pool in cells. Based on our understanding of DGCR8 mutants, we reanalyzed recently reported miRNA sequencing data and conclude that heme is generally required for processing canonical pri-miRNAs, that heme modulates the specificity of Microprocessor, and that cellular heme level and differential DGCR8 heme occupancy alter the expression of distinct groups of miRNAs in a hierarchical fashion. Overall, our study provides the first glimpse of a labile Fe(III) heme pool important for a fundamental physiological function and reveal principles governing how Fe(III) heme modulates miRNA maturation at a genomic scale. We also discuss potential states and biological significance of the labile Fe(III) heme pool.

molecular biology↗

iCodon: ideal codon design for customized gene expression

Messenger RNA (mRNA) stability substantially impacts steady-state gene expression levels in a cell. mRNA stability, in turn, is strongly affected by codon composition in a translation-dependent manner across species, through a mechanism termed codon optimality. We have developed iCodon (www.iCodon.org), an algorithm for customizing mRNA expression through the introduction of synonymous codon substitutions into the coding sequence. iCodon is optimized for four vertebrate transcriptomes: mouse, human, frog, and fish. Users can predict the mRNA stability of any coding sequence based on its codon composition and subsequently generate more stable (optimized) or unstable (deoptimized) variants encoding for the same protein. Further, we show that codon optimality predictions correlate with expression levels using fluorescent reporters and endogenous genes in human cells and zebrafish embryos. Therefore, iCodon will benefit basic biological research, as well as a wide range of applications for biotechnology and biomedicine.

molecular biology↗

MitoQuicLy: a high-throughput method for quantifying cell-free DNA from human plasma, serum, and saliva

Circulating cell-free mitochondrial DNA (cf-mtDNA) is an emerging biomarker of psychobiological stress and disease which predicts mortality and is associated with various disease states. To evaluate the contribution of cf-mtDNA to health and disease states, standardized high-throughput procedures are needed to quantify cf-mtDNA in relevant biofluids. Here, we describe MitoQuicLy: Mitochondrial DNA Quantification in cell-free samples by Lysis. We demonstrate high agreement between MitoQuicLy and the commonly used column-based method, although MitoQuicLy is faster, cheaper, and requires a smaller input sample volume. Using 10 {micro}L of input volume with MitoQuicLy, we quantify cf-mtDNA levels from three commonly used plasma tube types, two serum tube types, and saliva. We detect, as expected, significant inter-individual differences in cf-mtDNA across different biofluids. However, cf-mtDNA levels between concurrently collected plasma, serum, and saliva from the same individual differ on average by up to two orders of magnitude and are poorly correlated with one another, pointing to different cf-mtDNA biology or regulation between commonly used biofluids in clinical and research settings. Moreover, in a small sample of healthy women and men (n=34), we show that blood and saliva cf-mtDNAs correlate with clinical biomarkers differently depending on the sample used. The biological divergences revealed between biofluids, together with the lysis-based, cost-effective, and scalable MitoQuicLy protocol for biofluid cf-mtDNA quantification, provide a foundation to examine the biological origin and significance of cf-mtDNA to human health.

molecular biology↗

Concatemer Assisted Stoichiometry Analysis (CASA): a targeted mass spectrometry method for protein quantification

Large multi-protein machines are central to multiple biological processes. However, stoichiometric determination of protein complex subunits in their native states presents a significant challenge. This study addresses the limitations of current tools in accuracy and precision by introducing concatemer-assisted stoichiometry analysis (CASA). CASA leverages stable isotope-labeled concatemers and liquid chromatography parallel reaction monitoring mass spectrometry (LC-PRM-MS) to achieve robust quantification of proteins with sub-femtomole sensitivity. As a proof-of-concept, CASA was applied to study budding yeast kinetochores. Stoichiometries were determined for ex vivo reconstituted kinetochore components, including the canonical H3 nucleosomes, centromeric (Cse4CENP-A) nucleosomes, centromere proximal factors (Cbf1 and CBF3 complex), inner kinetochore proteins (Mif2CENP-C, Ctf19CCAN complex), and outer kinetochore proteins (KMN network). Absolute quantification by CASA revealed Cse4CENP-A as a cell-cycle controlled limiting factor for kinetochore assembly. These findings demonstrate that CASA is applicable for stoichiometry analysis of multi-protein assemblies. SummaryThis study presents Concatemer-Assisted Stoichiometry Analysis (CASA) to address a common challenge in cell biological research: quantifying the number of each protein subunit in a native protein complex.

molecular biology↗

TEMI: Tissue Expansion Mass Spectrometry Imaging

The spatial distribution of diverse biomolecules in multicellular organisms is essential for their physiological functions. High-throughput in situ mapping of biomolecules is crucial for both basic and medical research, and requires high scanning speed, spatial resolution, and chemical sensitivity. Here, we developed a Tissue Expansion method compatible with matrix-assisted laser desorption/ionization Mass spectrometry Imaging (TEMI). TEMI reaches single-cell spatial resolution without sacrificing voxel throughput and enables the profiling of hundreds of biomolecules, including lipids, metabolites, peptides (proteins), and N-glycans. Using TEMI, we mapped the spatial distribution of biomolecules across various mammalian tissues and uncovered metabolic heterogeneity in tumors. TEMI can be easily adapted and broadly applied in biological and medical research, to advance spatial multi-omics profiling.

molecular biology↗