bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,153 records · Page 64Linked to original sources

Expression of four mitochondrial tRNAs from only two loci

Transfer RNAs (tRNAs) are among the few genes retained in animal mitochondrial genomes after more than a billion years of gene loss. These ancient bacterial vestiges are often structurally aberrant and less stable than their bacterial or cytosolic tRNA counterparts. In some lineages, mitochondrial tRNAs (mt-tRNAs) have become so truncated that the loss of one or both arms has expanded our understanding of what constitutes a functional tRNA. Here, we report another radical departure from canonical tRNA gene architecture: two overlapping tRNAs produced from opposite strands of the same locus. These mirror tRNA pairs eliminate the need to retain separate loci for all tRNA genes, as a single locus can produce tRNAs to decode two different amino acids. We show that these mirror tRNAs are aminoacylated and demonstrate their presence in mitoribosomes. Furthermore, mirror tRNAs display strand-specific patterns of nucleotide modification and RNA editing, reflecting specific post-transcriptional maturation that depends on transcriptional orientation. To our knowledge, this demonstration of functional, bidirectional tRNA expression is a first for any genome or organism and reveals an unexpected strategy by which mitochondrial genomes maintain a complete set of tRNAs in the face of unrelenting gene loss. The discovery of mirror tRNAs has broad implications for the evolution of tRNA-interacting enzymes, mitochondrial biology, and even the origins of the protein synthesis machinery itself.

molecular biology↗

CreLite: An Optogenetically Controlled Cre/loxP System Using Red Light

Precise manipulation of gene expression with temporal and spatial control is essential for functional studies and the determination of cell lineage relationships in complex biological systems. The Cre-loxP system is commonly used for gene manipulation at desired times and places. However, specificity is dependent on the availability of tissue- or cell-specific regulatory elements used in combination with Cre or CreER (tamoxifen-inducible). Here we present CreLite, an optogenetically-controlled Cre system using red light in developing zebrafish embryos. Cre activity is disabled by splitting Cre and fusing the inactive halves with the Arabidopsis thaliana red light-inducible binding partners, PhyB and PIF6. In addition, PhyB-PIF6 binding requires phycocyanobilin (PCB), providing an additional layer of control. Upon exposure to red light (660 nm) illumination, the PhyB-CreC and PIF6-CreN fusion proteins come together in the presence of PCB to restore Cre activity. Red-light exposure of transgenic zebrafish embryos harboring a Cre-dependent multi-color fluorescent protein reporter (ubi:zebrabow) injected with CreLite mRNAs and PCB, resulted in Cre activity as measured by the generation of multi-spectral cell labeling in various tissues, including heart, skeletal muscle and epithelium. We show that CreLite can be used for gene manipulations in whole embryos or small groups of cells at different stages of development. CreLite provides a novel optogenetic tool for precise temporal and spatial control of gene expression in zebrafish embryos that may also be useful in cell culture, ex vivo organ culture, and other animal models for developmental biology studies.

molecular biology↗

Structural basis for differential p19 targeting by IL-23 biologics

BackgroundIL-23 is central to the pathogenesis of psoriasis, and is structurally comprised of p19 and p40 subunits. "Targeted" IL-23 inhibitors risankizumab, tildrakizumab, and guselkumab differ mechanistically from ustekinumab because they bind p19, whereas ustekinumab binds p40; however, a knowledge gap exists regarding the structural composition of their epitopes and how these molecular properties relate to their clinical efficacy. ObjectivesTo characterize and differentiate the structural epitopes of the IL-23 inhibitors risankizumab, guselkumab, tildrakinumab, and ustekinumab, and correlate their molecular characteristics with clinical response in plaque psoriasis therapy. MethodsWe utilized epitope data derived from hydrogen-deuterium exchange studies for risankizumab, tildrakizumab, and guselkumab, and crystallographic data for ustekinumab to map drug epitope locations, hydrophobicity, and surface charge onto the IL-23 molecular surface (Protein Data Bank ID Code 3D87) using UCSF Chimera. PDBePISA was used to calculate solvent accessible surface area (SASA). Epitope composition was determined by classifying residues as acidic, basic, polar, or hydrophobic and calculating their contribution to epitope SASA. Linear regression and analysis of variance was performed. ResultsAll the p19-specific inhibitor epitopes differ in location and size, with risankizumab and guselkumab having large epitope surface areas (SA), and tildrakizumab and ustekinumab having smaller SA. The tildrakizumab epitope was mostly hydrophobic (56%), while guselkumab, risankizumab, and ustekinumab epitopes displayed >50% non-hydrophobic residues. Risankizumab and ustekinumab exhibited acidic surface charges, while tildrakizumab and guselkumab were net neutral. Each inhibitor binds an epitope with a unique size and composition, and with mostly distinct locations except for a 10-residue overlap region that lies outside of the IL-23 receptor epitope. We observed a strong correlation between epitope SA and PASI-90 rates (R2 = 0.9969, p = 0.0016), as well as between epitope SA and KD (R2 = 0.9772, p = 0.0115). In contrast, we found that total epitope hydrophobicity, polarity, and charge content do not correlate with clinical efficacy. ConclusionsStructural analysis of IL-23 inhibitor epitopes reveals strong association between epitope SA and early drug efficacy in plaque psoriasis therapy, exemplifying how molecular data can explain clinical observations, inform future innovation, and help clinicians in specific drug selection for patients.

molecular biology↗

Identification of Pathways Associated with Chemosensitivity through Network Embedding

Basal gene expression levels have been shown to be predictive of cellular response to cytotoxic treatments. However, such analyses do not fully reveal complex genotype-phenotype relationships, which are partly encoded in highly interconnected molecular networks. Biological pathways provide a complementary way of understanding drug response variation among individuals. In this study, we integrate chemosensitivity data from a recent pharmacogenomics study with basal gene expression data from the CCLE project and prior knowledge of molecular networks to identify specific pathways mediating chemical response. We first develop a computational method called PACER, which ranks pathways for enrichment in a given set of genes using a novel network embedding method. It examines known relationships among genes as encoded in a molecular network along with gene memberships of all pathways to determine a vector representation of each gene and pathway in the same low-dimensional vector space. The relevance of a pathway to the given gene set is then captured by the similarity between the pathway vector and gene vectors. To apply this approach to chemosensitivity data, we identify genes with basal expression levels in a panel of cell lines that are correlated with cytotoxic response to a compound, and then rank pathways for relevance to these response-correlated genes using PACER. Extensive evaluation of this approach on benchmarks constructed from databases of compound target genes, compound chemical structure, as well as large collections of drug response signatures demonstrates its advantages in identifying compound-pathway associations, compared to existing statistical methods of pathway enrichment analysis. The associations identified by PACER can serve as testable hypotheses about chemosensitivity pathways and help further study the mechanism of action of specific cytotoxic drugs. More broadly, PACER represents a novel technique of identifying enriched properties of any gene set of interest while also taking into account networks of known gene-gene relationships and interactions.

pharmacology and toxicology↗

Stop Bickering! Reconciling Signaling Pathway Databases with Network Topologies

A major goal of molecular systems biology is to understand the coordinated function of genes or proteins in response to cellular signals and to understand these dynamics in the context of disease. Signaling pathway databases such as KEGG, NetPath, NCI-PID, and Panther describe the molecular interactions involved in different cellular responses. While the same pathway may be present in different databases, prior work has shown that the particular proteins and interactions differ across database annotations. However, to our knowledge no one has attempted to quantify their structural differences. It is important to characterize artifacts or other biases within pathway databases, which can provide a more informed interpretation for downstream analyses. In this work, we consider signaling pathways as graphs and we use topological measures to study their structure. We find that topological characterization using graphlets (small, connected subgraphs) distinguishes signaling pathways from appropriate null models of interaction networks. Next, we quantify topological similarity across pathway databases. Our analysis reveals that the pathways harbor database-specific characteristics implying that even though these databases describe the same pathways, they tend to be systematically different from one another. We show that pathway-specific topology can be uncovered after accounting for database-specific structure. This work present the first step towards elucidating common pathway structure beyond their specific database annotations.

systems biology↗

Simple biochemical features underlie transcriptional activation domain diversity and dynamic, fuzzy binding to Mediator

Gene activator proteins comprise distinct DNA-binding and transcriptional activation domains (ADs). Because few ADs have been described, we tested domains tiling all yeast transcription factors for activation in vivo and identified 150 ADs. By mRNA display, we showed that 73% of ADs bound the Med15 subunit of Mediator, and that binding strength was correlated with activation. AD-Mediator interaction in vitro was unaffected by a large excess of free activator protein, pointing to a dynamic mechanism of interaction. Structural modeling showed that ADs interact with Med15 without shape complementarity ("fuzzy" binding). ADs shared no sequence motifs, but mutagenesis revealed biochemical and structural constraints. Finally, a neural network trained on AD sequences accurately predicted ADs in human proteins and in other yeast proteins, including chromosomal proteins and chromatin remodeling complexes. These findings solve the longstanding enigma of AD structure and function and provide a rationale for their role in biology.

molecular biology↗

Sequence features of transcriptional activation domains are consistent with the surfactant mechanism of gene activation

Transcriptional activation domains (ADs) of gene activators remain enigmatic for decades as they are short, extremely variable in sequence, structurally disordered, and interact fuzzily to a spectrum of targets. We showed that the single required characteristic of the most common acidic ADs is an amphiphilic aromatic-acidic surfactant-like property which is the key for the local gene-promoter chromatin phase transition and the formation of "transcription factory" condensates. We demonstrate that the presence of tryptophan and aspartic acid residues in the AD sequence is sufficient for in vivo functionality, even when present only as a single pair of residues within a 20-amino-acid sequence containing only 18 additional glycine residues. We demonstrate that breaking the amphipathic -helix in AD by prolines increases AD functionality. The proposed mechanism is paradigm-shifting for gene activation area and generally for biochemistry as it relies on near-stochastic allosteric interactions critical for the key biological function.

molecular biology↗

Genetic activation of canonical RNA interference in mice

Canonical RNA interference (RNAi) is sequence-specific mRNA degradation guided by small interfering RNAs (siRNAs) made from double-stranded RNA (dsRNA) by RNase III Dicer. RNAi has different roles including gene regulation, antiviral immunity or defense against transposable elements. In mammals, RNAi is constrained by Dicer, which is adapted to produce microRNAs, another class of small RNAs. However, RNAi exists in mouse oocytes, which employs a truncated Dicer variant. A homozygous mutation to express only the truncated variant ({Delta}HEL1) causes dysregulation of microRNAs and perinatal lethality in mice. Here, we report the phenotype and RNAi activity in Dicer{Delta}HEL1/wtmice, which are viable, show minimal miRNome changes but their endogenous siRNA levels are increased by an order of magnitude. We show that siRNA abundance is limited by available dsRNA but not by PKR, a dsRNA sensor of innate immunity. Expressing dsRNA from a transgene, functional RNAi in vivo was induced in heart. Dicer{Delta}HEL1/wt mice thus represent a new model for researching mammalian canonical RNAi in vivo and offer an unprecedented platform for addressing claims about its biological roles.

molecular biology↗

High-Throughput Epigenetic Profiling Immunoassays for Accelerated Disease Research and Clinical Development

Epigenetics, which examines the regulation of genes without modification of the DNA sequence, plays a crucial role in various biological processes and disease mechanisms. Among the different forms of epigenetic modifications, histone post-translational modifications (PTMs) are important for modulating chromatin structure and gene expression. Aberrant levels of histone PTMs are implicated in a wide range of diseases, including cancer, making them promising targets for biomarker discovery and therapeutic intervention. In this context, blood, tissues, or cells serve as valuable resources for epigenetic research and analysis. Traditional methods such as mass spectrometry and western blotting are widely used to study histone PTMs, providing qualitative and (semi)quantitative information. However, these techniques often face limitations that could include throughput and scalability, particularly when applied to clinical samples. To overcome these challenges, we developed and validated 13 Nu.Q(R) immunoassays to detect and quantify specific histone PTM-nucleosomes from K2EDTA plasma samples. Then, we tested these assays on other types of samples, including chromatin extracts from frozen tissues, as well as cell lines and white blood cells Our findings demonstrate that the Nu.Q(R) assays offer high specificity, sensitivity, precision and linearity, making them effective tools for epigenetic profiling. A comparative analysis of HeLa cells using mass spectrometry, Western blot, and Nu.Q(R) immunoassays revealed a consistent histone PTMs signature, further validating the effectiveness of these assays. Additionally, we successfully applied Nu.Q(R) assays across various biological samples, including human tissues from different organs and specific white blood cell subtypes, highlighting their versatility and applicability in diverse biological contexts.

molecular biology↗

N-Terminal Deleted Isoforms of E3 Ligase RNF220 (Isoform 4) Are Ubiquitously Expressed and Required for Mouse Muscle Differentiation.

Four isoform peptides of the novel E3 ligase RNF220 have been identified in humans. However, all of previous studies have predominantly focused on isoform 1, which consists of 566 amino acids (aa). Here, we show that a shorter isoform, isoform 4 (308 aa), lacking most of the N-terminus, is the predominant and ubiquitously expressed variant that warrants functional investigation. Both isoform 1 and isoform 4 are expressed in the brain; however, isoform 4 is the major isoform expressed in all other tissues in mice. Consistently, H3K4me3 ChIP-seq data from ENCODE reveal that the transcription start site for isoform 4 demonstrates broader and stronger activity across human tissues than that of isoform 1. Isoform 4 produces two peptides (4a and 4b) through alternative translation initiation, with isoform 4b displaying distinct subcellular localization and subnuclear structures. Notably, during embryonic stem cell differentiation into neural stem cells, isoform 1 expression increases, whereas isoform 4 expression decreases. In murine myoblasts, isoform 4 is the sole expressed isoform and is required for MyoD and myogenin expression, as well as for muscle differentiation. Our findings highlight isoform 4 as the ubiquitously and highly expressed variant, likely playing a fundamental role across tissues while exhibiting functional differences from isoform 1. These results emphasize the critical importance of isoform 4 in future studies investigating the biological functions of RNF220.

molecular biology↗

The extra-terminal domain drives the role of BET proteins in transcription

BET proteins facilitate the transcription of most eukaryotic genes, yet the specific mechanisms underlying their function remain incompletely understood. As chromatin readers, BET proteins use tandem bromodomains to recognize acetylated lysine residues on histones and other protein partners. However, recent evidence indicates that bromodomain activity alone does not account for the full spectrum of BET protein functions, underscoring the importance of additional conserved domains. Here, we systematically evaluated all conserved domains of BET proteins and identified the extra-terminal (ET) domain as essential for cell viability, genome-wide transcription, and BET chromatin occupancy. Moreover, we demonstrate that the ET domain exerts these effects by acting as a central hub for interactions with multiple transcriptional regulators. Our findings advance the current understanding of BET protein biology and reveal potential mechanisms by which cells can evade bromodomain inhibition under pathological conditions.

molecular biology↗

Efficient CRISPR/Cas-mediated homologous recombination in the model diatom Thalassiosira pseudonana

CRISPR/Cas enables targeted genome editing in many different plant and algal species including the model diatom Thalassiosira pseudonana. However, efficient gene targeting by homologous recombination (HR) to date is only reported for photosynthetic organisms in their haploid life-cycle phase and there are no examples of efficient nuclease-meditated HR in any photosynthetic organism. Here, a CRISPR/Cas construct, assembled using Golden Gate cloning, enabled highly efficient HR for the first time in a diploid photosynthetic organism. HR was induced in T. pseudonana by means of sequence specific CRISPR/Cas, paired with a donor matrix, generating substitution of the silacidin gene by a resistance cassette (FCP:NAT). Approximately 85% of NAT resistant T. pseudonana colonies screened positive for HR using a nested PCR approach and confirmed by sequencing of the PCR products. The knockout of the silacidin gene in T. pseudonana caused a significant increase in cell size, confirming the role of this gene for cell-size regulation in centric diatoms. Highly efficient gene targeting by HR makes T. pseudonana as genetically tractable as Nannochloropsis and Physcomitrella, hence rapidly advancing functional diatom biology, bionanotechnology and any biotechnological application targeted on harnessing the metabolic potential of diatoms.

molecular biology↗

Chd1 regulates repair of promoter-proximal DNA breaks to sustain hypertranscription in embryonic stem cells

Stem and progenitor cells undergo a global elevation of nascent transcription, or hypertranscription, during key developmental transitions involving rapid cell proliferation. The chromatin remodeler Chd1 binds to genes transcribed by RNA Polymerase (Pol) I and II and is required for hypertranscription in embryonic stem (ES) cells in vitro and the early post-implantation epiblast in vivo. Biochemically, Chd1 has been shown to facilitate transcription at least in part by removing nucleosomal barriers to elongation, but its mechanism of action in stem cells remains poorly understood. Here we report a novel role for Chd1 in the repair of promoter-proximal endogenous double-stranded DNA breaks (DSBs) in ES cells. An unbiased proteomics approach revealed that Chd1 interacts with several DNA repair factors including Atm, Parp1, Kap1 and Topoisomerase 2{beta}. We show that wild-type ES cells display high levels of phosphorylated H2A.X and Kap1 at chromatin, notably at rDNA in the nucleolus, in a Chd1-dependent manner. Loss of Chd1 leads to an extensive accumulation of DSBs at Chd1-bound Pol II-transcribed genes and rDNA. Genes prone to DNA breaks in Chd1 KO ES cells tend to be longer genes with GC-rich promoters, a more labile nucleosomal structure and roles in chromatin regulation, transcription and signaling. These results reveal a vulnerability of hypertranscribing stem cells to endogenous DNA breaks, with important implications for developmental and cancer biology.

molecular biology↗

A comprehensive dataset of TLX1 positive ALL-SIL lymphoblasts and primary T-cell acute lymphoblastic leukemias

Most currently available transcriptome data of T-cell acute lymphoblastic leukemia (T-ALL) are based on polyA[+] RNA sequencing methods thus lacking non-polyadenylated transcripts. Here, we present the data of polyA[+] and total RNA sequencing in the context of in vitro TLX1 knockdown in ALL-SIL cells and a primary T-ALL cohort. We extended this dataset with ATAC sequencing and H3K4me1 and H3K4me3 ChIP sequencing data to map putative gene regulatory regions. In this data descriptor, we present a detailed report of how the data were generated and which bioinformatics analyses were performed. Through several technical validations, we showed that our sequencing data are of high quality and that our in vitro TLX1 knockdown was successful. We also validated the quality of the ATAC and ChIP sequencing data and showed that ATAC and H3K4me3 ChIP peaks are enriched at transcription start sites. We believe that this comprehensive set of sequencing data can be reused by others to further unravel the complex biology of T-ALL in general and TLX1 in particular.

molecular biology↗

Cell-type-resolved RNP topologies reveal dynamic structural mechanisms of splicing and therapeutic targets

Resolving RNA conformations in native ribonucleoprotein (RNP) complexes remains a fundamental challenge. Here, we introduce spatial hydroxyl acylation reversible crosslinking with immunoprecipitation (SHARCLIP) to simultaneously capture RNA-RNA, RNA-protein and protein-protein contacts in cells. SHARCLIP profiling of HNRNPC-associated RNA conformations established a global phased map of ribonucleosomes, resolving a decades-old debate on heterogenous nuclear (hn)RNP assembly. We built a dynamic structural atlas across seven cell lineages for >10,000 RNAs, generating ~200 million contacts, and identifying millions of dynamic loops, steric blockers and conformational switches that control splicing outcome. Deciphering the structural logic of mutually exclusive exons (MXEs) enabled rational design of structure-breaking and stabilizing antisense oligonucleotides (ASOs). We demonstrate effective isoform swapping in 12 genes linked to genetic disorders. SHARCLIP provides a comprehensive roadmap for cellular RNA structural biology and structure-guided RNA therapeutics.

molecular biology↗

MiRNA Atlas: A Literature-Derived Database of MicroRNAs Bridging Osteoarthritis and Appendage Regeneration

Many microRNAs (miRNAs) regulate tissue remodeling, cellular plasticity, and repair across evolutionarily distant vertebrate lineages that are capable of regenerating appendages such as amputated limbs, fins, and antlers, as well as in human articular cartilage responding to injury. These miRNAs often belong to the same families and exert conserved, though occasionally inverted, regulatory effects. This strong cross-species overlap motivates the present literature-derived analysis. Osteoarthritis (OA), the most prevalent joint disease, still lacks disease-modifying therapies, in part because of the longstanding assumption that adult mammalian cartilage demonstrates no intrinsic reparative capacity. Yet human cartilage retains a latent repair program activated by mechanical and inflammatory stress. Some injured or degenerating joints may never be clinically recognized as osteoarthritic because their intrinsic repair capacity is sufficient to restore tissue integrity; in others, where repair capacity is diminished or damage exceeds it, the repair program is insufficient and OA becomes clinically manifest. MiRNAs are established post-transcriptional regulators of cartilage homeostasis, degeneration, and appendage regeneration, yet because the OA and regeneration research fields have advanced largely independently, the insights available at their intersection have gone unrecognized. To close this gap, we systematically mined both literatures to construct an auto-updating, cross-referenced atlas of OA- and appendage regeneration-associated miRNAs. Integrating these datasets identified a core set of shared miRNA families, delineated miRNAs unique to each field, and mapped convergent families onto common pathways governing matrix remodeling, dedifferentiation, senescence, and inflammation. We propose that regeneration-competent species can inform the identification of therapeutic miRNAs, such as miR-133, miR-21, and let-7, capable of activating endogenous cartilage repair. Collectively, this synthesis and its accompanying web-based miRNA atlas (https://mirnaatlas.shinyapps.io/mirnaatlas/) establish a comparative framework for regenerative miRNA biology and provide a continually updated resource to accelerate discovery of disease-modifying, RNA-based therapies for OA.

molecular biology↗

Ribosome stalling caused by the Argonaute-miRNA-SGS3 complex regulates production of secondary siRNA biogenesis in plants

The path of ribosomes on mRNAs can be impeded by various obstacles. One such example is halting of ribosome movement by microRNAs, though the exact mechanism and physiological role remain unclear. Here, we find that ribosome stalling caused by the Argonaute-microRNA-SGS3 complex regulates the production of secondary small interfering RNAs (siRNAs) in plants. We show that the double-stranded RNA-binding protein SGS3 directly interacts with the 3' end of the microRNA in an Argonaute protein, resulting in ribosome stalling. Importantly, microRNA-mediated ribosome stalling positively correlates with efficient production of secondary siRNAs from target mRNAs. Our results illustrate a role for paused ribosomes in regulation of small RNA function that may have broad biological implications across the plant kingdom.

molecular biology↗

Chromosome specific telomere lengths and the minimal functional telomere revealed by nanopore sequencing

We developed a method to tag telomeres and measure telomere length by nanopore sequencing in the yeast S. cerevisiae. Nanopore allows long read sequencing through the telomere, subtelomere and into unique chromosomal sequence, enabling assignment of telomere length to a specific chromosome end. We observed chromosome end specific telomere lengths that were stable over 120 cell divisions. These stable chromosome specific telomere lengths may be explained by stochastic clonal variation or may represent a new biological mechanism that maintains equilibrium unique to each chromosomes end. We examined the role of RIF1 and TEL1 in telomere length regulation and found that TEL1 is epistatic to RIF1 at most telomeres, consistent with the literature. However, at telomeres that lack subtelomeric Y sequences, tel1{Delta} rif1{Delta} double mutants had a very small, but significant, increase in telomere length compared to the tel1{Delta} single mutant, suggesting an influence of Y elements on telomere length regulation. We sequenced telomeres in a telomerase-null mutant (est2{Delta}) and found the minimal telomere length to be around 75bp. In these est2{Delta} mutants there were many apparent telomere recombination events at individual telomeres before the generation of survivors, and these events were significantly reduced in est2{Delta} rad52{Delta} double mutants. The rate of telomere shortening in the absence of telomerase was similar across all chromosome ends at about 5 bp per generation. This new method gives quantitative, high resolution telomere length measurement at each individual chromosome end, suggests possible new biological mechanisms regulating telomere length, and provides capability to test new hypotheses.

molecular biology↗