bioRxiv ScienceSearch

Biology subjects

Huang, T.

Publications and source records attributed to Huang, T..

3 recordsLinked to original sources

Designer Sinorhizobium meliloti strains and multi-functional vectors for direct inter-kingdom transfer of high G+C content DNA

Storage and manipulation of large DNA fragments is crucial for synthetic biology applications, yet DNA with high G+C content can be unstable in many host organisms. Here, we report the development of Sinorhizobium meliloti as a new universal host that can store DNA, including high G+C content, and mobilize DNA to Escherichia coli, Saccharomyces cerevisiae, and the eukaryotic microalgae Phaeodactylum tricornutum. We deleted the S. meliloti hsdR restriction-system to enable DNA transformation with up to 1.4 x 105 efficiency. Multi-host and multi-functional shuttle vectors (MHS) were constructed and shown to stably replicate in S. meliloti, E. coli, S. cerevisiae, and P. tricornutum, with a copy-number inducible E. coli origin for isolating plasmid DNA. Crucially, we demonstrated that S. meliloti can act as a universal conjugative donor for MHS plasmids with a cargo of at least 62 kb of G+C rich DNA derived from Deinococcus radiodurans.

synthetic biology

Widespread and polymorphous noncoding amino acid residues in human sperm proteome

Proteins are usually deciphered by translation of the coding genome; however, their amino acid residues are seldom determined directly across the proteome. Herein, we describe a systematic workflow for identifying all possible protein residues that differ from the coding genome, termed noncoded amino acids (ncAAs). By measuring the mass differences between the coding amino acids and the actual protein residues in human spermatozoa, over a million nonzero delta masses were detected, fallen into 424 high-quality Gaussian clusters and 571 high-confidence ncAAs spanning 29,053 protein sites. Most ncAAs are novel with unresolved side-chains and discriminative between healthy individuals and patients with oligoasthenospermia. For validation, 40 out of 98 ncAAs that matched with amino acid substitutions were confirmed by exon sequencing. This workflow revealed the widespread existence of previously unreported ncAAs in the sperm proteome, which represents a new dimension on the understanding of amino acid polymorphisms at the proteomic level.\n\nHighlightsO_LI571 ncAAs spanning 108,000 protein sites were identified in human sperm proteome.\nC_LIO_LIMost ncAAs are novel with unresolved sidechains and found at unreported protein sites.\nC_LIO_LIExon sequencing confirmed 40 of 98 ncAAs that matched with amino acid substitutions.\nC_LIO_LIMany ncAAs are linked with disease and have potential for diagnosis and targeting.\nC_LI\n\neTOC BlurbWe describe a systematic identification of all possible protein residues that were not encoded by their genomic sequences. A total of 571 high-confidence most novel noncoded amino acids were identified in human sperm proteome, corresponding to over 108,000 ncAA-containing protein sites. For validation, 40 out of 98 ncAAs that matched to amino acid substitutions were confirmed by exon sequencing. These ncAAs are discriminative between individuals and expand our understanding of amino acid polymorphisms in human proteomes and diseases.

molecular biology

ActiveDriverDB: human disease mutations and genome variation in post-translational modification sites of proteins

Interpretation of genetic variation is required for understanding genotype-phenotype associations, mechanisms of inherited disease, and drivers of cancer. Millions of single nucleotide variants (SNVs) in human genomes are known and thousands are associated with disease. An estimated 20% of disease-associated missense SNVs are located in protein sites of post-translational modifications (PTMs), chemical modifications of amino acids that extend protein function. ActiveDriverDB is a comprehensive human proteo-genomics database that annotates disease mutations and population variants using PTMs. We integrated >385,000 published PTM sites with [~]3.8 million missense SNVs from The Cancer Genome Atlas (TCGA), the ClinVar database of disease genes, and inter-individual variation from human genome sequencing projects. The database includes interaction networks of proteins, upstream enzymes such as kinases, and drugs targeting these enzymes. We also predicted network-rewiring impact of mutations by analyzing gains and losses of kinase-bound sequence motifs. ActiveDriverDB provides detailed visualization, filtering, browsing and searching options for studying PTM-associated SNVs. Users can upload mutation datasets interactively and use our application programming interface for pipelines. Integrative analysis of SNVs and PTMs helps decipher molecular mechanisms of phenotypes and disease, as exemplified by case studies of disease genes TP53, BRCA2 and VHL. The open-source database is available at https://www.ActiveDriverDB.org.

bioinformatics