bioRxiv Science⌕ Search

Biology subjects

Weller, J.

Publications and source records attributed to Weller, J..

5 recordsLinked to original sources

Predicting efficiency of writing short sequences into the genome using prime editing

Short sequences can be precisely written into a selected genomic target using prime editing. This ability facilitates protein tagging, correction of pathogenic deletions, and many other exciting applications. However, it remains unclear what types of sequences prime editors can easily insert, and how to choose optimal reagents for a desired outcome. To characterize features that influence insertion efficiency, we designed a library of 2,666 sequences up to 69 nt in length and measured the frequency of their insertion into four genomic sites in three human cell lines, using different prime editor systems. We discover that insertion sequence length, nucleotide composition and secondary structure all affect insertion rates, and that mismatch repair proficiency is a strong determinant for the shortest insertions. Combining the sequence and repair features into a machine learning model, we can predict insertion frequency for new sequences with R = 0.69. The tools we provide allow users to choose optimal constructs for DNA insertion using prime editing.

molecular biology↗

Predicting base editing outcomes using position-specific sequence determinants

Nucleotide-level control over DNA sequences is poised to power functional genomics studies and lead to new therapeutics. CRISPR/Cas base editors promise to achieve this ability, but the determinants of their activity remain incompletely understood. We measured base editing frequencies in two human cell lines for two cytosine and two adenine base editors at [~]14,000 target sequences. Base editing activity is sequence-biased, with largest effects from nucleotides flanking the target base, and is correlated with measures of Cas9 guide RNA efficiency. Whether a base is edited depends strongly on the combination of its position in the target and the preceding base, with a preceding thymine in both editor types leading to a wider editing window, while a preceding guanine in cytosine editors and preceding adenine in adenine editors to a narrower one. The impact of features on editing rate depends on the position, with guide RNA efficacy mainly influencing bases around the centre of the window, and sequence biases away from it. We use these observations to train a machine learning model to predict editing activity per position for both adenine and cytosine editors, with accuracy ranging from 0.49 to 0.72 between editors, and with better generalization performance across datasets than existing tools. We demonstrate the usefulness of our model by predicting the efficacy of potential disease mutation correcting guides, and find that most of them suffer from more unwanted editing than corrected outcomes. This work unravels the position-specificity of base editing biases, and provides a solution to account for them, thus allowing more efficient planning of base edits in experimental and therapeutic contexts.

genetics↗

Transcriptomics data availability and reusability in the transition from microarray to next-generation sequencing

Over the last two decades, molecular biology has been changed by the introduction of high-throughput technologies. Data sharing requirements have prompted the establishment of persistent data archives. A standardized approach for recording and managing these data was first proposed in the Minimal Information About a Microarray Experiment (MIAME) guidelines. The Minimal Information about a high throughput nucleotide Sequencing Experiment (MINSEQE) proposal was introduced in 2008 as a logical extension of the guidelines to next-generation sequencing (NGS) technologies used for transcriptome analysis. We present a historical snapshot of the data-sharing situation focusing on transcriptomics data from both microarray and RNA-sequencing experiments published between 2009 and 2013, a period during which RNA-seq studies became increasingly popular for transcriptome analysis. We assess how much data from RNA-seq based experiments is actually available in persistent data archives, compared to data derived from microarray based experiments, and evaluate how these types of data differ. Based on this analysis, we provide recommendations to improve RNA-seq data availability, reusability, and reproducibility.

bioinformatics↗

Efficient design of maximally active and specific nucleic acid diagnostics for thousands of viruses

Diagnostics, particularly for rapidly evolving viruses, stand to benefit from a principled, measurement-driven design that harnesses machine learning and vast genomic data--yet the capability for such design has not been previously built. Here, we develop and extensively validate an approach to designing viral diagnostics that applies a learned model within a combinatorial optimization framework. Concentrating on CRISPR-based diagnostics, we screen a library of 19,209 diagnostic-target pairs and train a deep neural network that predicts, from RNA sequence alone, diagnostic signal better than contemporary techniques. Our model then makes it possible to design assays that are maximally sensitive over the spectrum of a viruss genomic variation. We introduce ADAPT (https://adapt.guide), a system for fully-automated design, and use ADAPT to design optimal diagnostics for the 1,933 vertebrate-infecting viral species within 2 hours for most species and 24 hours for all but 3. We experimentally show ADAPTs designs are sensitive and specific down to the lineage level, including against viruses that pose challenges involving genomic variation and specificity. ADAPTs designs exhibit significantly higher fluorescence and permit lower limits of detection, across a viruss entire variation, than the outputs of standard design techniques. Our model-based optimization strategy has applications broadly to viral nucleic acid diagnostics and other sequence-based technologies, and, paired with clinical validation, could enable a critically-needed, proactive resource of assays for surveilling and responding to pathogens.

genomics↗

AutoRELACS: Automated Generation And Analysis Of Ultra-parallel ChIP-seq

Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is a method used to profile protein-DNA interactions genome-wide. RELACS (Restriction Enzyme-based Labeling of Chromatin in Situ) is a recently developed ChIP-seq protocol that deploys a chromatin barcoding strategy to enable standardized and high-throughput generation of ChIP-seq data. The manual implementation of RELACS is constrained by human processivity in both data generation and data analysis. To overcome these limitations, we have developed AutoRELACS, an automated implementation of the RELACS protocol using the liquid handler Biomek i7 workstation. We match the unprecedented processivity in data generation allowed by AutoRELACS with the automated computation pipelines offered by snakePipes. In doing so, we build a continuous workflow that streamlines epigenetic profiling, from sample collection to biological interpretation. Here, we show that AutoRELACS successfully automates chromatin barcode integration, and is able to generate high-quality ChIP-seq data comparable with the standards of the manual protocol, also for limited amounts of biological samples.

genomics↗