bioRxiv ScienceSearch

Biology subjects

Clement, K.

Publications and source records attributed to Clement, K..

6 recordsLinked to original sources

Interpretable and accurate prediction models for metagenomics data

Biomarker discovery using metagenomic data is becoming more prevalent for patient diagnosis, prognosis and risk evaluation. Selected groups of microbial features provide signatures that characterize host disease states such as cancer or cardio-metabolic diseases. Yet, the current predictive models stemming from machine learning still behave as black boxes. Moreover, they seldom generalize well when learned on small datasets. Here, we introduce an original approach that focuses on three models inspired by microbial ecosystem interactions: the addition, subtraction, and ratio of microbial taxon abundances. While being extremely simple, their performance is surprisingly good and compares to or is better than Random Forest, SVM or Elastic Net. Such models besides being interpretable, allow distilling biological information of the predictive core-variables. Collectively, this approach builds up both reliable and trustworthy diagnostic decisions while agreeing with societal and legal pressure that require explainable AI models in the medical domain.

bioinformatics

Analysis and comparison of genome editing using CRISPResso2

Genome editing technologies are rapidly evolving, and analysis of deep sequencing data from target or off-target regions is necessary for measuring editing efficiency and evaluating safety. However, no software exists to analyze base editors, perform allele-specific quantification or that incorporates biologically-informed and scalable alignment approaches. Here, we present CRISPResso2 to fill this gap and illustrate its functionality by experimentally measuring and analyzing the editing properties of six genome editing agents.

bioinformatics

CRISPR-SURF: Discovering regulatory elements by deconvolution of CRISPR tiling screen data

Tiling screens using CRISPR-Cas technologies provide a powerful approach to map regulatory elements to phenotypes of interest, but computational methods that effectively model these experimental approaches for different CRISPR technologies are not readily available. Here we present CRISPR-SURF, a deconvolution framework to identify functional regulatory regions in the genome from data generated by CRISPR-Cas nuclease, CRISPR interference (CRISPRi), or CRISPR activation (CRISPRa) tiling screens. We validated CRISPR-SURF on previously published and new data, identifying both experimentally validated and new potential regulatory elements. With CRISPR tiling screens now being increasingly used to elucidate the regulatory architecture of the non-coding genome, CRISPRSURF provides a generalizable and accessible solution for the discovery of regulatory elements.

bioinformatics

AmpUMI: Design and analysis of unique molecular identifiers for deep amplicon sequencing

MotivationUnique molecular identifiers (UMIs) are added to DNA fragments before PCR amplification to discriminate between alleles arising from the same genomic locus and sequencing reads produced by PCR amplification. While computational methods have been developed to take into account UMI information in genome-wide and single-cell sequencing studies, they are not designed for modern amplicon based sequencing experiments, especially in cases of high allelic diversity. Importantly, no guidelines are provided for the design of optimal UMI length for amplicon-based sequencing experiments.\n\nResultsBased on the total number of DNA fragments and the distribution of allele frequencies, we present a model for the determination of the minimum UMI length required to prevent UMI collisions and reduce allelic distortion. We also introduce a user-friendly software tool called AmpUMI to assist in the design and the analysis of UMI-based amplicon sequencing studies. AmpUMI provides quality control metrics on frequency and quality of UMIs, and trims and deduplicates amplicon sequences with user specified parameters for use in downstream analysis. AmpUMI is open-source and freely available at http://github.com/pinellolab/AmpUMI.\n\nContactIpinello@mgh.harvard.edu

bioinformatics

"Unexpected mutations after CRISPR-Cas9 editing in vivo" are most likely pre-existing sequence variants and not nuclease-induced mutations

Schaefer et al. recently advanced the provocative conclusion that CRISPR-Cas9 nuclease can induce off-target alterations at genomic loci that do not resemble the intended on-target site.1 Using high-coverage whole genome sequencing (WGS), these authors reported finding SNPs and indels in two CRISPR-Cas9-treated mice that were not present in a single untreated control mouse. On the basis of this association, Schaefer et al. concluded that these sequence variants were caused by CRISPR-Cas9. This new proposed CRISPR-Cas9 off-target activity runs contrary to previously published work2-8 and, if the authors are correct, could have profound implications for research and therapeutic applications. Here, we demonstrate that the simplest interpretation of Schaefer et al.s data is that the two CRISPR-Cas9-treated mice are actually more closely related genetically to each other than to the control mouse. This strongly suggests that the so-called \"unexpected mutations\" simply represent SNPs and indels shared in common by these mice prior to nuclease treatment. In addition, given the genomic and sequence distribution profiles of these variants, we show that it is challenging to explain how CRISPR-Cas9 might be expected to induce such changes. Finally, we argue that the lack of appropriate controls in Schaefer et al.s experimental design precludes assignment of causality to CRISPR-Cas9. Given these substantial issues, we urge Schaefer et al. to revise or re-state the original conclusions of their published work so as to avoid leaving misleading and unsupported statements to persist in the literature.

molecular biology

Genetic architecture of early childhood growth phenotypes gives insights into their link with later obesity

Early childhood growth patterns are associated with adult metabolic health, but the underlying mechanisms are unclear. We performed genome-wide meta-analyses and follow-up in up to 22,769 European children for six early growth phenotypes derived from longitudinal data: peak height and weight velocities, age and body mass index (BMI) at adiposity peak (AP ~9 months) and rebound (AR ~5-6 years). We identified four associated loci (P< 5x10-8): LEPR/LEPROT with BMI at AP, FTO and TFAP2B with Age at AR and GNPDA2 with BMI at AR. The observed AR-associated SNPs at FTO, TFAP2B and GNPDA2 represent known adult BMI-associated variants. The common BMI at AP associated variant at LEPR/LEPROT was not associated with adult BMI but was associated with LEPROT gene expression levels, especially in subcutaneous fat (P<2x10-51). We identify strong positive genetic correlations between early growth and later adiposity traits, and analysis of the full discovery stage results for Age at AR revealed enrichment for insulin-like growth factor 1 (IGF-1) signaling and apolipoprotein pathways. This genome-wide association study suggests mechanistic links between early childhood growth and adiposity in later childhood and adulthood, highlighting these early growth phenotypes as potential targets for the prevention of obesity.

genomics