bioRxiv ScienceSearch

Biology subjects

Fang, H.

Publications and source records attributed to Fang, H..

8 recordsLinked to original sources

Metabolic engineering of Escherichia coli for de novo biosynthesis of vitamin B12

The only known source of vitamin B12 (adenosylcobalamin) is from bacteria and archaea, and the only unknown step in its biosynthesis is the production of the intermediate adenosylcobinamide phosphate. Here, using genetic and metabolic engineering, we generated an Escherichia coli strain that produces vitamin B12 via an engineered de novo aerobic biosynthetic pathway. Excitingly, the BluE and CobC enzymes from Rhodobacter capsulatus transform L-threonine into (R)-1-Amino-2-propanol O-2-Phosphate, which is then condensed with adenosylcobyric acid to yield adenosylcobinamide phosphate by either CobD from the aeroic R. capsulatus or CbiB from the anerobic Salmonella typhimurium. These findings suggest that the biosynthetic steps from co(II)byrinic acid a,c-diamide to adocobalamin are the same in both the aerobic and anaerobic pathways. Finally, we increased the vitamin B12 yield of a recombinant E. coli strain by more than [~]250-fold to 307.00 {micro}g/g DCW via metabolic engineering and optimization of fermentation conditions. Beyond our scientific insights about the aerobic and anaerobic pathways and our demonstration of E. coli as a microbial biosynthetic platform for vitamin B12 production, our study offers an encouraging example of how the several dozen proteins of a complex biosynthetic pathway can be transferred between organisms to facilitate industrial production.

synthetic biology

WDR45 contributes to neurodegeneration through regulation of ER homeostasis and neuronal death

Mutations in the autophagy gene WDR45 cause {beta}-propeller protein-associated neurodegeneration (BPAN); however the molecular and cellular mechanism of the disease process is largely unknown. Here we generated constitutive Wdr45 knockout (KO) mice that displayed cognitive impairments, abnormal synaptic transmission and lesions in hippocampus and basal ganglia. Immunohistochemistry analysis shows loss of neurons in prefrontal cortex and basal ganglion in aged mice, and increased apoptosis in these regions, recapitulating a hallmark of neurodegeneration. Quantitative proteomic analysis shows accumulation of endoplasmic reticulum (ER) proteins in KO mouse. Furthermore, we show that a defect in autophagy results in impaired ER turnover and ER stress. The unfolded protein response (UPR) is elevated through IRE1 and possibly other kinase signaling pathways, and eventually leads to neuronal apoptosis. Suppression of ER stress, or activation of autophagy through inhibition of mTOR pathway rescues neuronal death. Thus, our study not only provides mechanistic insights for BPAN, but also suggests that a defect in macroautophagy machinery leads to impairment in selective organelle autophagy.

cell biology

Complex rearrangements and oncogene amplifications revealed by long-read DNA and RNA sequencing of a breast cancer cell line

The SK-BR-3 cell line is one of the most important models for HER2+ breast cancers, which affect one in five breast cancer patients. SK-BR-3 is known to be highly rearranged although much of the variation is in complex and repetitive regions that may be underreported. Addressing this, we sequenced SK-BR-3 using long-read single molecule sequencing from Pacific Biosciences, and develop one of the most detailed maps of structural variations (SVs) in a cancer genome available with nearly 20,000 variants present, most of which were missed by prior efforts. Surrounding the important HER2 locus, we discover a complex sequence of nested duplications and translocations, suggesting a punctuated progression. Full-length transcriptome sequencing further revealed several novel gene fusions within the nested genomic variants. Combining long-read genome and transcriptome sequencing enables an in-depth analysis of how SVs disrupt the transcriptome and sheds new light on the complexity of cancer progression.

genomics

Accurate detection of complex structural variations using single molecule sequencing

Structural variations (SVs) are the largest source of genetic variation, but remain poorly understood because of limited genomics technology. Single molecule long read sequencing from Pacific Biosciences and Oxford Nanopore has the potential to dramatically advance the field, although their high error rates challenge existing methods. Addressing this need, we introduce open-source methods for long read alignment (NGMLR, https://github.com/philres/ngmlr) and SV identification (Sniffles, https://github.com/fritzsedlazeck/Sniffles) that enable unprecedented SV sensitivity and precision, including within repeat-rich regions and of complex nested events that can have significant impact on human disorders. Examining several datasets, including healthy and cancerous human genomes, we discover thousands of novel variants using long reads and categorize systematic errors in short-read approaches. NGMLR and Sniffles are further able to automatically filter false events and operate on low amounts of coverage to address the cost factor that has hindered the application of long reads in clinical and research settings.

bioinformatics

Orientation-dependent Dxz4 contacts shape the 3D structure of the inactive X chromosome

The mammalian inactive X chromosome (Xi) condenses into a bipartite structure with two superdomains of frequent long-range contacts separated by a boundary or hinge region. Using in situ DNase Hi-C in mouse cells with deletions or inversions within the hinge we show that the conserved repeat locus Dxz4 alone is sufficient to maintain the bipartite structure and that Dxz4 orientation controls the distribution of long-range contacts on the Xi. Frequent long-range contacts between Dxz4 and the telomeric superdomain are either lost after its deletion or shifted to the centromeric superdomain after its inversion. This massive reversal in contact distribution is consistent with the reversal of CTCF motif orientation at Dxz4. De-condensation of the Xi after Dxz4 deletion is associated with partial restoration of TADs normally attenuated on the Xi. There is also an increase in chromatin accessibility and CTCF binding on the Xi after Dxz4 deletion or inversion, but few changes in gene expression, in accordance with multiple epigenetic mechanisms ensuring X silencing. We propose that Dxz4 represents a structural platform for frequent long-range contacts with multiple loci in a direction dictated by the orientation of a bank of CTCF motifs at Dxz4, which may work as a ratchet to form the distinctive bipartite structure of the condensed Xi.

molecular biology

Scikit-ribo: Accurate estimation and robust modeling of translation dynamics at codon resolution

Ribosome profiling (Riboseq) is a powerful technique for measuring protein translation, however, sampling errors and biological biases are prevalent and poorly understand. Addressing these issues, we present Scikit-ribo (https://github.com/hanfang/scikit-ribo), the first open-source software for accurate genome-wide A-site prediction and translation efficiency (TE) estimation from Riboseq and RNAseq data. Scikit-ribo accurately identifies A-site locations and reproduces codon elongation rates using several digestion protocols (r = 0.99). Next we show commonly used RPKM-derived TE estimation is prone to biases, especially for low-abundance genes. Scikit-ribo introduces a codon-level generalized linear model with ridge penalty that correctly estimates TE while accommodating variable codon elongation rates and mRNA secondary structure. This corrects the TE errors for over 2000 genes in S. cerevisiae, which we validate using mass spectrometry of protein abundances (r = 0.81) and allows us to determine the Kozak-like sequence directly from Riboseq. We conclude with an analysis of coverage requirements needed for robust codon-level analysis, and quantify the artifacts that can occur from cycloheximide treatment.

bioinformatics

Persistent homology demarcates a leaf morphospace

Current morphometric methods that comprehensively measure shape cannot compare the disparate leaf shapes found in seed plants and are sensitive to processing artifacts. We explore the use of persistent homology, a topological method applied across the scales of a function, to overcome these limitations. The described method isolates subsets of shape features and measures the spatial relationship of neighboring pixel densities in a shape. We apply the method to the analysis of 182,707 leaves, both published and unpublished, representing 141 plant families collected from 75 sites throughout the world. By measuring leaves from throughout the seed plants using persistent homology, a defined morphospace comparing all leaves is demarcated. Clear differences in shape between major phylogenetic groups are detected and estimates of leaf shape diversity within plant families are made. This approach does not only predict plant family, but also the collection site, confirming phylogenetically invariant morphological features that characterize leaves from specific locations. The application of a persistent homology method to measure leaf shape allows for a unified morphometric framework to measure plant form, including shape and branching architectures.

plant biology

HadoopCNV: A Dynamic Programming Imputation Algorithm To Detect Copy Number Variants From Sequencing Data

BACKGROUNDWhole-genome sequencing (WGS) data may be used to identify copy number variations (CNVs). Existing CNV detection methods mostly rely on read depth or alignment characteristics (paired-end distance and split reads) to infer gains/losses, while neglecting allelic intensity ratios and cannot quantify copy numbers. Additionally, most CNV callers are not scalable to handle a large number of WGS samples.\n\nMETHODSTo facilitate large-scale and rapid CNV detection from WGS data, we developed a Dynamic Programming Imputation (DPI) based algorithm called HadoopCNV, which infers copy number changes through both allelic frequency and read depth information. Our implementation is built on the Hadoop framework, enabling multiple compute nodes to work in parallel.\n\nRESULTSCompared to two widely used tools - CNVnator and LUMPY, HadoopCNV has similar or better performance on both simulated data sets and real data on the NA12878 individual. Additionally, analysis on a 10-member pedigree showed that HadoopCNV has a Mendelian precision that is similar or better than other tools. Furthermore, HadoopCNV can accurately infer loss of heterozygosity (LOH), while other tools cannot. HadoopCNV requires only 1.6 hours for a human genome with 30X coverage, on a 32-node cluster, with a linear relationship between speed improvement and the number of nodes. We further developed a method to combine HadoopCNV and LUMPY result, and demonstrated that the combination resulted in better performance than any individual tools.\n\nCONCLUSIONSThe combination of high-resolution, allele-specific read depth from WGS data and Hadoop framework can result in efficient and accurate detection of CNVs.

bioinformatics