bioRxiv ScienceSearch

Biology subjects

Chan, S.

Publications and source records attributed to Chan, S..

9 recordsLinked to original sources

Resource: Scalable whole genome sequencing of 40,000 single cells identifies stochastic aneuploidies, genome replication states and clonal repertoires

Essential features of cancer tissue cellular heterogeneity such as negatively selected genome topologies, sub-clonal mutation patterns and genome replication states can only effectively be studied by sequencing single-cell genomes at scale and high fidelity. Using an amplification-free single-cell genome sequencing approach implemented on commodity hardware (DLP+) coupled with a cloud-based computational platform, we define a resource of 40,000 single-cell genomes characterized by their genome states, across a wide range of tissue types and conditions. We show that shallow sequencing across thousands of genomes permits reconstruction of clonal genomes to single nucleotide resolution through aggregation analysis of cells sharing higher order genome structure. From large-scale population analysis over thousands of cells, we identify rare cells exhibiting mitotic mis-segregation of whole chromosomes. We observe that tissue derived scWGS libraries exhibit lower rates of whole chromosome anueploidy than cell lines, and loss of p53 results in a shift in event type, but not overall prevalence in breast epithelium. Finally, we demonstrate that the replication states of genomes can be identified, allowing the number and proportion of replicating cells, as well as the chromosomal pattern of replication to be unambiguously identified in single-cell genome sequencing experiments. The combined annotated resource and approach provide a re-implementable large scale platform for studying lineages and tissue heterogeneity.

genomics

Aspirin treatment does not increase microhemorrhage size in young or aged mice

Microhemorrhages are common in the aging brain and are thought to contribute to cognitive decline and the development of neurodegenerative diseases, such as Alzheimers disease. Chronic aspirin therapy is widespread in older individuals and decreases the risk of coronary artery occlusions and stroke. There remains a concern that such aspirin usage may prolong bleeding after a vessel rupture in the brain, leading to larger bleeds that cause more damage to the surrounding tissue. Here, we aimed to understand the influence of aspirin usage on the size of cortical microhemorrhages and explored the impact of age. We used femtosecond laser ablation to rupture arterioles in the cortex of both young (2-5 months old) and aged (18-29 months old) mice dosed on aspirin in their drinking water and measured the extent of penetration of both red blood cells and blood plasma into the surrounding tissue. We found no difference in microhemorrhage size for both young and aged mice dosed on aspirin, as compared to controls (hematoma diameter = 104 +/- 39 (97 +/- 38) m in controls and 109 +/- 25 (101 +/- 28) m in aspirin-treated young (aged) mice; mean +/- SD). In contrast, young mice treated with intravenous heparin had an increased hematoma diameter of 136 +/- 44 m. These data suggest that aspirin does not increase the size of microhemorrhages, supporting the safety of aspirin usage.

neuroscience

CRISPR-bind: a simple, custom CRISPR/dCas9-mediated labeling of genomic DNA for mapping in nanochannel arrays

Bionano genome mapping is a robust optical mapping technology used for de novo construction of whole genomes using ultra-long DNA molecules, able to efficiently interrogate genomic structural variation. It is also used for functional analysis such as epigenetic analysis and DNA replication mapping and kinetics. Genomic labeling for genome mapping is currently specified by a single strand nicking restriction enzyme followed by fluorophore incorporation by nick-translation (NLRS), or by a direct label and stain (DLS) chemistry which conjugates a fluorophore directly to an enzyme-defined recognition site. Although these methods are efficient and produce high quality whole genome mapping data, they are limited by the number of available enzymes--and thus the number of recognition sequences--to choose from. The ability to label other sequences can provide higher definition in the data and may be used for countless additional applications. Previously, custom labeling was accomplished via the nick-translation approach using CRISPR-Cas9, leveraging Cas9 mutant D10A which has one of its cleavage sites deactivated, thus effectively converting the CRISPR-Cas9 complex into a nickase with customizable target sequences. Here we have improved upon this approach by using dCas9, a nuclease-deficient double knockout Cas9 with no cutting activity, to directly label DNA with a fluorescent CRISPR-dCas9 complex (CRISPR-bind). Unlike labeling with CRISPR-Cas9 D10A nickase, in which nicking, labeling, and repair by ligation, all occur as separate steps, the new assay has the advantage of labeling DNA in one step, since the CRISPR-dCas9 complex itself is fluorescent and remains bound during imaging. CRISPR-bind can be added directly to a sample that has already been labeled using DLS or NLRS, thus overlaying additional information onto the same molecules. Using the dCas9 protein assembled with custom target crRNA and fluorescently labeled tracrRNA, we demonstrate rapid labeling of repetitive DUF1220 elements. We also combine NLRS-based whole genome mapping with CRISPR-bind labeling targeting Alu loci. This rapid, convenient, non-damaging, and cost-effective technology is a valuable tool for custom labeling of any CRISPR-Cas9 amenable target sequence.

genomics

Structural variability of EspG chaperones from mycobacterial ESX-1, ESX-3 and ESX-5 type VII secretion systems

Type VII secretion systems (ESX) are responsible for transport of multiple proteins in mycobacteria. How different ESX systems achieve specific secretion of cognate substrates remains elusive. In the ESX systems, the cytoplasmic chaperone EspG forms complexes with heterodimeric PE-PPE substrates that are secreted from the cells or remain associated with the cell surface. Here we report the crystal structure of the EspG1 chaperone from the ESX-1 system determined using a fusion strategy with T4 lysozyme. EspG1 adopts a quasi 2-fold symmetric structure that consists of a central {beta}-sheet and two -helical bundles. Additionally, we describe the structures of EspG3 chaperones from four different crystal forms. Alternate conformations of the putative PE-PPE binding site are revealed by comparison of the available EspG3 structures. Analysis of EspG1, EspG3 and EspG5 chaperones using small-angle X-ray scattering (SAXS) reveals that EspG1 and EspG3 chaperones form dimers in solution, which we observed in several of our crystal forms. Finally, we propose a model of the ESX-3 specific EspG3-PE5-PPE4 complex based on the SAXS analysis.\n\nHighlightsO_LIThe crystal structure of EspG1 reveals the common architecture of the type VII secretion system chaperones\nC_LIO_LIStructures of EspG3 chaperones display a number of conformations that could reflect alternative substrate binding modes\nC_LIO_LIEspG3 chaperones dimerize in solution\nC_LIO_LIA model of EspG3 in complex with its substrate PE-PPE dimer is proposed based on SAXS data\nC_LI

biochemistry

Improved Aedes aegypti mosquito reference genome assembly enables biological discovery and vector control

Female Aedes aegypti mosquitoes infect hundreds of millions of people each year with dangerous viral pathogens including dengue, yellow fever, Zika, and chikungunya. Progress in understanding the biology of this insect, and developing tools to fight it, has been slowed by the lack of a high-quality genome assembly. Here we combine diverse genome technologies to produce AaegL5, a dramatically improved and annotated assembly, and demonstrate how it accelerates mosquito science and control. We anchored the physical and cytogenetic maps, resolved the size and composition of the elusive sex-determining \"M locus\", significantly increased the known members of the glutathione-S-transferase genes important for insecticide resistance, and doubled the number of chemosensory ionotropic receptors that guide mosquitoes to human hosts and egg-laying sites. Using high-resolution QTL and population genomic analyses, we mapped new candidates for dengue vector competence and insecticide resistance. We predict that AaegL5 will catalyse new biological insights and intervention strategies to fight this deadly arboviral vector.

genomics

Genome-Wide Identification of Early-Firing Human Replication Origins by Optical Replication Mapping

The timing of DNA replication is largely regulated by the location and timing of replication origin firing. Therefore, much effort has been invested in identifying and analyzing human replication origins. However, the heterogeneous nature of eukaryotic replication kinetics and the low efficiency of individual origins in metazoans has made mapping the location and timing of replication initiation in human cells difficult. We have mapped early-firing origins in HeLa cells using Optical Replication Mapping, a high-throughput single-molecule approach based on Bionano Genomics genomic mapping technology. The single-molecule nature and 290-fold coverage of our dataset allowed us to identify origins that fire with as little as 1% efficiency. We find sites of human replication initiation in early S phase are not confined to well-defined efficient replication origins, but are instead distributed across broad initiation zones consisting of many inefficient origins. These early-firing initiation zones co-localize with initiation zones inferred from Okazaki-fragment-mapping analysis and are enriched in ORC1 binding sites. Although most early-firing origins fire in early-replication regions of the genome, a significant number fire in late-replicating regions, suggesting that the major difference between origins in early and late replicating regions is their probability of firing in early S-phase, as opposed to qualitative differences in their firing-time distributions. This observation is consistent with stochastic models of origin timing regulation, which explain the regulation of replication timing in yeast.

genomics

RES complex is associated with intron definition and required for zebrafish early embryogenesis

Pre-mRNA splicing is a critical step of gene expression in eukaryotes. Transcriptome-wide splicing patterns are complex and primarily regulated by a diverse set of recognition elements and associated RNA-binding proteins. The retention and splicing (RES) complex is formed by three different proteins (Bud13p, Pml1p and Snu17p) and is involved in splicing in yeast. However, the importance of the RES complex for vertebrate splicing, the intronic features associated with its activity, and its role in development are unknown. In this study, we have generated loss-of-function mutants for the three components of the RES complex in zebrafish and showed that they are required during early development. The mutants showed a marked neural phenotype with increased cell death in the brain and a decrease in differentiated neurons. Transcriptomic analysis of bud13, snip1 (pml1) and rbmx2 (snu17) mutants revealed a global defect in intron splicing, with strong mis-splicing of a subset of introns. We found these RES-dependent introns were short, rich in GC and flanked by GC depleted exons, all of which are features associated with intron definition. Using these features we developed a predictive model that classifies RES dependent introns. Altogether, our study uncovers the essential role of the RES complex during vertebrate development and provides new insights into its function during splicing.

genomics

Systematic mapping of the free energy landscapes of a growing immunoglobulin domain identifies a kinetic intermediate associated with co-translational proline isomerization

Co-translational folding is a fundamental molecular process that ensures efficient protein biosynthesis and minimizes the wasteful or hazardous formation of misfolded states. However, the complexity of this process makes it extremely challenging to obtain structural characterizations of co-translational folding pathways. Here we contrast observations in translationally-arrested nascent chains with those of a systematic C-terminal truncation strategy. We create a detailed description of chain length-dependent free energy landscapes associated with folding of the FLN5 filamin domain, in isolation and on the ribosome. By using this approach we identify and characterize two folding intermediates, including a partially folded intermediate associated with the isomerization of a conserved proline residue, which, together with measurements of folding kinetics, raises the prospect that neighboring unfolded domains might accumulate during biosynthesis. We develop a simple model to quantify the risk of misfolding in this situation, and show that catalysis of folding by peptidyl-prolyl isomerases is essential to eliminate this hazard.

biophysics

The DOE Systems Biology Knowledgebase (KBase)

The U.S. Department of Energy Systems Biology Knowledgebase (KBase) is an open-source software and data platform designed to meet the grand challenge of systems biology -- predicting and designing biological function from the biomolecular (small scale) to the ecological (large scale). KBase is available for anyone to use, and enables researchers to collaboratively generate, test, compare, and share hypotheses about biological functions; perform large-scale analyses on scalable computing infrastructure; and combine experimental evidence and conclusions that lead to accurate models of plant and microbial physiology and community dynamics. The KBase platform has (1) extensible analytical capabilities that currently include genome assembly, annotation, ontology assignment, comparative genomics, transcriptomics, and metabolic modeling; (2) a web-browser-based user interface that supports building, sharing, and publishing reproducible and well-annotated analyses with integrated data; (3) access to extensive computational resources; and (4) a software development kit allowing the community to add functionality to the system.

bioinformatics