bioRxiv Science⌕ Search

Biology subjects

Tabet, D.

Publications and source records attributed to Tabet, D..

2 recordsLinked to original sources

Pacybara: Accurate long-read sequencing for barcoded mutagenized allelic libraries

SummaryLong read sequencing technologies, an attractive solution for many applications, often suffer from higher error rates. Alignment of multiple reads can improve base-calling accuracy, but some applications, e.g. sequencing mutagenized libraries where multiple distinct clones differ by one or few variants, require the use of barcodes or unique molecular identifiers. Unfortunately, sequencing errors can interfere with correct barcode identification, and a given barcode sequence may be linked to multiple independent clones within a given library. Here we focus on the target application of sequencing mutagenized libraries in the context of multiplexed assays of variant effects (MAVEs). MAVEs are increasingly used to create comprehensive genotype-phenotype maps that can aid clinical variant interpretation. Many MAVE methods use long-read sequencing of barcoded mutant libraries for accurate association of barcode with genotype. Existing long-read sequencing pipelines do not account for inaccurate sequencing or non-unique barcodes. Here, we describe Pacybara, which handles these issues by clustering long reads based on the similarities of (error-prone) barcodes while also detecting barcodes that have been associated with multiple genotypes. Pacybara also detects recombinant (chimeric) clones and reduces false positive indel calls. In three example applications, we show that Pacybara identifies and correctly resolves these issues. Availability and ImplementationPacybara, freely available at https://github.com/rothlab/pacybara, is implemented using R, Python and bash for Linux. It has both a single-threaded implementation and, for GNU/Linux clusters that use Slurm, PBS, or GridEngine schedulers, a multi-node version. Supplementary MaterialSupplementary materials are available at Bioinformatics online.

bioinformatics↗

Genome-scale mapping of DNA damage suppressors identifies GNB1L as essential for ATM and ATR biogenesis

To maintain genome integrity, cells must avoid DNA damage by ensuring the accurate duplication of the genome and by having efficient repair and signaling systems that counteract the genome-destabilizing potential of DNA lesions. To uncover genes and pathways that suppress DNA damage in human cells, we undertook genome-scale CRISPR/Cas9 screens that monitored the levels of DNA damage in the absence or presence of DNA replication stress. We identified 160 genes in RKO cells whose mutation caused high levels of DNA damage in the absence of exogenous genotoxic treatment. This list was highly enriched in essential genes, highlighting the importance of genomic integrity for cellular fitness. Furthermore, the majority of these 160 genes are involved in a limited set of biological processes related to DNA replication and repair, nucleotide biosynthesis, RNA metabolism and iron sulfur cluster biogenesis, suggesting that genome integrity may be insulated from a wide range of cellular processes. Among the many genes identified and validated in this study, we discovered that GNB1L, a schizophrenia/autism-susceptibility gene implicated in 22q11.2 syndrome, protects cells from replication catastrophe promoted by mild DNA replication stress. We show that GNB1L is involved in the biogenesis of ATR and related phosphatidylinositol 3-kinase-related kinases (PIKKs) through its interaction with the TTT co-chaperone complex. These results implicate PIKK biogenesis as a potential root cause for the neuropsychiatric phenotypes associated with 22q11.2 syndrome. The phenotypic mapping of genes that suppress DNA damage in human cells therefore provides a powerful approach to probe genome maintenance mechanisms.

molecular biology↗