bioRxiv ScienceSearch

Biology subjects

Yu, F.

Publications and source records attributed to Yu, F..

10 recordsLinked to original sources

Multi-Center Study of Resectable Lung Lesions by Ultra-Deep Sequencing of Targeted Genes in Plasma Cell-Free DNA to Assess Nodule Malignancy and Detect Lung Cancers

BACKGROUNDEarly detection of lung cancer to allow curative treatment remains challenging. Cell-free circulating tumor DNA (ctDNA) analysis may aid in malignancy assessment and early cancer diagnosis of lung nodules found in screening imagery.\n\nMETHODSThe multi-center clinical study enrolled 192 patients with operable occupying lung diseases. Plasma ctDNA, white blood cell genomic DNA (gDNA) and tumor tissue gDNA of each patient were analyzed by ultra-deep sequencing to an average of 35,000X of the coding regions of 65 lung cancer-related genes.\n\nRESULTSThe cohort consists of a quarter of benign lung diseases and three quarters of cancer patients with all histopathology subtypes. 64% of the cancer patients is at Stage I. Gene mutations detection in tissue gDNA and plasma ctDNA results in a sensitivity of 91% and specificity of 88%. When ctDNA assay was used as the test, the sensitivity was 69% and specificity 96%. As for the lung cancer patients, the assay detected 63%, 83%, 94% and 100%, for Stage I, II, III and IV, respectively. In a linear discriminant analysis, combination of ctDNA, patient age and a panel of serum biomarkers boosted the overall sensitivity to 80% at a specificity of 99%. 29 out of the 65 genes harbored mutations in the lung cancer patients with the largest number found in TP53 (30% plasma and 62% tumor tissue samples) and EGFR (20% and 40%, respectively).\n\nCONCLUSIONPlasma ctDNA was analyzed in lung nodule assessment and early cancer detection while an algorithm combining clinical information enhanced the test performance.

clinical trials

Understanding the limit of open search in the identification of peptides with post-translational modifications -- A simulation-based study

MotivationAnalyzing tandem mass spectrometry data to recognize peptides in a sample is the fundamental task in computational proteomics. Traditional peptide identification algorithms perform well when identifying unmodified peptides. However, when peptides have post-translational modifications (PTMs), these methods cannot provide satisfactory results. Recently, Chick et al., 2015 and Yu et al., 2016 proposed the spectrum-based and tag-based open search methods, respectively, to identify peptides with PTMs. While the performance of these two methods is promising, the identification results vary greatly with respect to the quality of tandem mass spectra and the number of PTMs in peptides. This motivates us to systematically study the relationship between the performance of open search methods and quality parameters of tandem mass spectrum data, as well as the number of PTMs in peptides.\n\nResultsThrough large-scale simulations, we obtain the performance trend when simulated tandem mass spectra are of different quality. We propose an analytical model to describe the relationship between the probability of obtaining correct identifications and the spectrum quality as well as the number of PTMs. Based on the analytical model, we can quantitatively describe the necessary condition to effectively apply open search methods.\n\nAvailabilitySource codes of the simulation are available at http://bioinformatics.ust.hk/PST.html.\n\nContactboningli@ust.hk or eeyu@ust.hk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

The Early Diagnosis in Lung Cancer by the Detection of Circulating Tumor DNA

BackgroundRemarkable advances for clinical diagnosis and treatment in cancers including lung cancer involve cell-free circulating tumor DNA (ctDNA) detection through next generation sequencing. However, before the sensitivity and specificity of ctDNA detection can be widely recognized, the consistency of mutations in tumor tissue and ctDNA should be evaluated. The urgency of this consistency is extremely obvious in lung cancer to which great attention has been paid to in liquid biopsy field.\n\nMethodsWe have developed an approach named systematic error correction sequencing (Sec-Seq) to improve the evaluation of sequence alterations in circulating cell-free DNA. Averagely 10 ml preoperative blood samples were collected from 30 patients containing pulmonary space occupying pathological changes by traditional clinic diagnosis. cfDNA from plasma, genomic DNA from white blood cells, and genomic DNA from solid tumor of above patients were extracted and constructed as libraries for each sample before subjected to sequencing by a panel contains 50 cancer-associated genes encompassing 29 kb by custom probe hybridization capture with average depth >40000, 7000, or 6300 folds respectively.\n\nResultsDetection limit for mutant allele frequency in our study was 0.1%. The sequencing results were analyzed by bioinformatic expertise based on our previous studies on the baseline mutation profiling of circulating cell-free DNA and the clinicopathological data of these patients. Among all the lung cancer patients, 78% patients were predicted as positive by ctDNA sequencing when the shreshold was defined as at least one of the hotspot mutations detected in the blood (ctDNA) was also detected in tumor tissue. Pneumonia and pulmonary tuberculosis were detected as negative according to the above standard. When evaluating all hotspots in driver genes in the panel, 24% mutations detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). When evaluating all genetic variations in the panel, including all the driver genes and passenger genes, 28% detected in tumor tissue (tDNA) were also detected in patients blood (ctDNA). Positive detection rates of plasma ctDNA in stage I lung cancer patients is 85%, compared with 17% of tumor biomarkers.\n\nConclusionWe demonstrated the importance of sequencing both circulating cell-free DNA and genomic DNA in tumor tissue for ctDNA detection in lung cancer currently. We also determined and confirmed the consistency of ctDNA and tumor tissue through NGS according to the criteria explored in our studies. Our strategy can initially distinguish the lung cancer from benign lesions of lung. Our work shows that the consistency will be benefited from the optimization in sensitivity and specificity in ctDNA detection.

cancer biology

Structural basis of the Cope rearrangement and C-C bond-forming cascade in hapalindole/fischerindole biogenesis

STRUCTURESThe atomic coordinates and structure factors for:\n\nHpiC1 W73M/K132M SeMet (P212121) -1.7 [A]\n\nHpiC1 native (C2) -1.5 [A]\n\nHpiC1 native (P42) -2.1 [A]\n\nHpiC1 Y101F (C2) -1.4 [A]\n\nHpiC1 Y101S (C2) -1.4 [A]\n\nHpiC1 F138S (P21) -1.7 [A]\n\nHpiC1 Y101F/F138S (P21 -1.65 [A] have been deposited with the Research Collaboratory for Structural Bioinformatics as Protein Data Bank entries 5WPP, 5WPR, 6AL6, 5WPR, 5WPU, 6AL7, and 6AL8 (www.rcsb.org).\n\nGRANTSThis work was supported by: The authors thank the National Science Foundation under the CCI Center for Selective C-H Functionalization (CHE-1205646), the National Institutes of Health (CA70375 to RMW and DHS), R35 GM118101, R01 GM076477 and the Hans W. Vahlteich Professorship (to DHS) for financial support. M.G-B. thanks the Ramon Areces Foundation for a postdoctoral fellowship. J.N.S. acknowledges the support of the National Institute of General Medical Sciences of the National Institutes of Health under Award Number F32GM122218. Computational resources were provided by the UCLA Institute for Digital Research and Education (IDRE) and the Extreme Science and Engineering Discovery Environment (XSEDE), which is supported by the NSF (OCI-1053575). The content does not necessarily represent the official views of the National Institutes of Health.\n\nABSTRACTHapalindole alkaloids are a structurally diverse class of cyanobacterial natural products defined by their varied polycyclic ring systems and diverse biological activities. These polycyclic scaffolds are generated from a common biosynthetic intermediate by the Stig cyclases in three mechanistic steps, including a rare Cope-rearrangement, 6-exo-trig cyclization, and electrophilic aromatic substitution. Here we report the structure of HpiC1, a Stig cyclase that catalyzes the formation of 12-epi-hapalindole U in vitro. The 1.5 [A] structure reveals a dimeric assembly with two calcium ions per monomer and the active sites located at the distal ends of the protein dimer. Mutational analysis and computational methods uncovered key residues for an acid catalyzed [3,3]-sigmatropic rearrangement and specific determinants that control the position of terminal electrophilic aromatic substitution leading to a switch from hapalindole to fischerindole alkaloids.

biochemistry

PhaMers identifies novel bacteriophage sequences from thermophilic hot springs

Metagenomic sequencing approaches have become popular for the purpose of dissecting environmental microbial diversity, leading to the characterization of novel microbial lineages. In addition of bacterial and fungal genomes, metagenomic analysis can also reveal genomes of viruses that infect microbial cells. Because of their small genome size and limited knowledge of phage diversity, discovering novel phage sequences from metagenomic data is often challenging. Here we describe PhaMers (Phage k-Mers). a phage identification tool that uses supervised learning to classify metagenomic contigs as phage or non-phage on the basis of tetranucleotide frequencies. a technique that does not depend on existing gene annotations. PhaMers compares the tetranucleotide frequencies of metagenomic contigs to phage and bacteria references from online databases. resulting in assignments of lower level phage taxonomy based on sequence similarity. Using PhaMers. we identified 103 novel phage sequences from hot spring samples of Yellowstone National Park based on data generated from a microfluidic-based minimetagenomic approach. We analyzed assembled contigs over 5 kbp in length using PhaMers and compared the results with those generated by VirSorter, a publicly available phage identification and annotation package. We analyzed the performance of phage genome prediction and taxonomic classification using PhaMers. and presented putative hosts and taxa for some of the novel phage sequences. Finally. mini-metagenomic occurrence profiles of phage and prokaryotic genomes were used to verify putative hosts.

bioinformatics

Xolik: finding cross-linked peptides with maximum paired scores in linear time

MotivationCross-linking technique coupled with mass spectrometry (MS) is widely used in the analysis of protein structures and protein-protein interactions. In order to identify cross-linked peptides from MS data, we need to consider all pairwise combinations of peptides, which is computationally prohibitive when the sequence database is large. To alleviate this problem, some heuristic screening strategies are used to reduce the number of peptide pairs during the identification. However, heuristic screening criteria may ignore true findings.\n\nResultsWe directly tackle the combination challenge without using any screening strategies. With the additive scoring function and the data structure of double-ended queue, the proposed algorithm reduces the quadratic time complexity of exhaustive searching down to the linear time complexity. We implement the algorithm in a tool named Xolik, and the running time of Xolik is validated using databases with different number of proteins. Experiments using synthetic and empirical datasets show that Xolik outperforms existing tools in terms of running time and statistical power.\n\nAvailabilitySource code and binaries of Xolik are freely available at http://bioinformatics.ust.hk/Xolik.html.\n\nContacteeyu@ust.hk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Microfluidic-based mini-metagenomics enables discovery of novel microbial lineages from complex environmental samples

Metagenomics and single-cell genomics have enabled the discovery of many new genomes from previously unknown branches of life. However, extracting novel genomes from complex mixtures of metagenomic data can still be challenging and in many respects represents an ill-posed problem which is generally approached with ad hoc methods. Here we present a microfluidic-based mini-metagenomic method which offers a statistically rigorous approach to extract novel microbial genomes from complex samples. In addition, by generating 96 sub-samples from each environmental sample, this method maintains high throughput, reduces sample complexity, and preserves single-cell resolution. We used this approach to analyze two hot spring samples from Yellowstone National Park and extracted 29 new genomes larger than 0.5 Mbps. These genomes represent novel lineages at different taxonomic levels, including three deeply branching lineages. Functional analysis revealed that these organisms utilize diverse pathways for energy metabolism. The resolution of this mini-metagenomic method enabled accurate quantification of genome abundance, even for genomes less than 1% in relative abundance. Our analyses also revealed a wide range of genome level single nucleotide polymorphism (SNP) distributions with nonsynonymous to synonymous ratio indicative of low to moderate environmental selection. The scale, resolution, and statistical power of microfluidic-based mini-metagenomic make it a powerful tool to dissect the genomic structure microbial communities while effectively preserving the fundamental unit of biology, the single cell.

microbiology

New kids on the block: Intercontinental dissemination and transmission of newly emerging lineages of multi-drug resistant Escherichia coli with highly dynamic resistance gene acquisition.

The increase in infections as a result of multi-drug resistant strains of Escherichia coli is a global health crisis. The emergence of globally disseminated lineages of E. coli carrying ESBL genes has been well characterised. An increase in strains producing carbapenemase enzymes and mobile colistin resistance is now being reported, but to date there is little genomic characterisation of such strains. Routine screening of patients within an ICU of West China Hospital identified a number of E. coli carrying the blaNDM-5 carbapenemase gene, found to be two distinct clones, E. coli ST167 and ST617. Interrogation of publically available data shows isolation of ESBL and carbapenem resistant strains of both lineages from clinical cases across the world. Further analysis of a large collection of publically available genomes shows that ST167 and ST617 have emerged in distinct patterns from the ST10 clonal complex of E. coli, but share evolutionary events involving switches in LPS genetics, intergenic regions and anaerobic metabolism loci. These may be evolutionary events which underpin the emergence of carbapenem resistance plasmid carriage in E. coli.

microbiology

ECL 2.0: Exhaustively Identifying Cross-Linked Peptides with a Linear Computational Complexity

Chemical cross-linking coupled with mass spectrometry is a powerful tool to study protein-protein interactions and protein conformations. Two linked peptides are ionized and fragmented to produce a tandem mass spectrum. In such an experiment, a tandem mass spectrum contains ions from two peptides. The peptide identification problem becomes a peptide-peptide pair identification problem. Currently, most existing tools dont search all possible pairs due to the quadratic time complexity. Consequently, a significant percentage of linked peptides are missed. In our earlier work, we developed a tool named ECL to search all pairs of peptides exhaustively. While ECL does not miss any linked peptides, it is very slow due to the quadratic computational complexity, especially when the database is large. Furthermore, ECL uses a score function without statistical calibration, while researchers1,2 have demonstrated that using a statistical calibrated score function can achieve a higher sensitivity than using an uncalibrated one.\n\nHere, we propose an advanced version of ECL, named ECL 2.0. It achieves a linear time and space complexity by taking advantage of the additive property of a score function. It can analyze a typical data set containing tens of thousands of spectra using a large-scale database containing thousands of proteins in a few hours. Comparison with other five state-of-the-art tools shows that ECL 2.0 is much faster than pLink, StavroX, ProteinProspector, and ECL. Kojak is the only one tool that is faster than ECL 2.0. But Kojak does not exhaustively search all possible peptide pairs. We also adopt an e-value estimation method to calibrate the original score. Comparison shows that ECL 2.0 has the highest sensitivity among the state-of-the-art tools. The experiment using a large-scale in vivo cross-linking data set demonstrates that ECL 2.0 is the only tool that can find PSMs passing the false discovery rate threshold. The result illustrates that exhaustive search and well calibrated score function are useful to find PSMs from a huge search space.

bioinformatics

Hardy Weinberg Exact Test In Large Scale Variant Calling Quality Control

Hardy Weinberg Equilibrium (HWE) test is widely used as a quality control measure to detect sequencing artifacts like mismapping, allelic dropout and biases. However, in the high throughput sequencing era, where the sample size is beyond a thousand scale, the utility of HWE test in reducing the false positive rate remains unclear. In this paper, we demonstrate that HWE test has limited power in identifying sequencing artifacts when the variant allele frequency is lower than 1% in a variant call set produced from more than five thousand whole genome sequenced samples from two homogeneous populations. We develop a novel strategy of implementing HWE filtering in which we incorporate site frequency spectrum information and determine the p-value cutoff which optimizes the tradeoff between sensitivity and specificity. The novel strategy is shown to outperform the exact test of HWE with an empirical constant p-value cutoff regardless of the sequencing sample size. We also present best practice recommendations for identifying possible sources of false positives from large sequencing datasets based on an analysis of intrinsic biases in the variant calling process. Our novel strategy of determining the HWE test p-value cutoff and applying the test to the common variants provides a practical approach for the variant level quality controls in the upcoming sequencing projects with tens to hundreds of thousand of samples.

bioinformatics