bioRxiv Science⌕ Search

Biology subjects

Luebbert, L.

Publications and source records attributed to Luebbert, L..

10 recordsLinked to original sources

Delphy: scalable, near-real-time Bayesian phylogenetics for outbreaks

Pathogen genomic analysis is central to tracking, understanding, and containing outbreaks, but complexity and high costs of state-of-the-art (SOTA) phylogenetic tools limit global access and impact. We introduce Delphy, an exact reformulation of Bayesian phylogenetics designed to transform its speed, scalability and accessibility while retaining SOTA accuracy. Delphys central data structure, an Explicit Mutation Annotated Tree, exploits the high sequence similarity in large-scale epidemic datasets for efficient tree exploration and convergence. By reproducing key analyses from recent major epidemics (Ebola, Zika, SARS-CoV-2, mpox, and H5N1), we demonstrate SOTA accuracy with up to 1,000x speedups. Assessing Delphys scalability, we show that a simulated dataset of 100,000 sequences can be analyzed in under a day-the largest such computation to date. We distribute Delphy as a client-side web application, enabling users worldwide to turn raw data into interactive results within minutes, without the data ever leaving the users machine. Delphy automatically identifies key viral lineages and mutations, as well as their emergence and prevalence through time, all with quantified uncertainties derived from a solid theoretical foundation. Delphy shows the power of Bayesian phylogenetics as a fast, accessible frontline tool for tackling future outbreaks.

genomics↗

The impact of package selection and versioning on single-cell RNA-seq analysis

Standard single-cell RNA-sequencing analysis (scRNA-seq) workflows consist of converting raw read data into cell-gene count matrices through sequence alignment, followed by analyses including filtering, highly variable gene selection, dimensionality reduction, clustering, and differential expression analysis. Seurat and Scanpy are the most widely-used packages implementing such workflows, and are generally thought to implement individual steps similarly. We investigate in detail the algorithms and methods underlying Seurat and Scanpy and find that there are, in fact, considerable differences in the outputs of Seurat and Scanpy. The extent of differences between the programs is approximately equivalent to the variability that would be introduced in benchmarking scRNA-seq datasets by sequencing less than 5% of the reads or analyzing less than 20% of the cell population. Additionally, distinct versions of Seurat and Scanpy can produce very different results, especially during parts of differential expression analysis. Our analysis highlights the need for users of scRNA-seq to carefully assess the tools on which they rely, and the importance of developers of scientific software to prioritize transparency, consistency, and reproducibility for their tools.

bioinformatics↗

Efficient and accurate detection of viral sequences at single-cell resolution reveals novel viruses perturbing host gene expression

There are an estimated 300,000 mammalian viruses from which infectious diseases in humans may arise. They inhabit human tissues such as the lungs, blood, and brain and often remain undetected. Efficient and accurate detection of viral infection is vital to understanding its impact on human health and to make accurate predictions to limit adverse effects, such as future epidemics. The increasing use of high-throughput sequencing methods in research, agriculture, and healthcare provides an opportunity for the cost-effective surveillance of viral diversity and investigation of virus-disease correlation. However, existing methods for identifying viruses in sequencing data rely on and are limited to reference genomes or cannot retain single-cell resolution through cell barcode tracking. We introduce a method that accurately and rapidly detects viral sequences in bulk and single-cell transcriptomics data based on highly conserved amino acid domains, which enables the detection of RNA viruses covering over 100,000 virus species. The analysis of viral presence and host gene expression in parallel at single-cell resolution allows for the characterization of host viromes and the identification of viral tropism and host responses. We applied our method to identify putative novel viruses in rhesus macaque PBMC data that display cell type specificity and whose presence correlates with altered host gene expression.

bioinformatics↗

kallisto, bustools, and kb-python for quantifying bulk, single-cell, and single-nucleus RNA-seq

The term "RNA-seq" refers to a collection of assays based on sequencing experiments that involve quantifying RNA species from bulk tissue, from single cells, or from single nuclei. The kallisto, bustools, and kb-python programs are free, open-source software tools for performing this analysis that together can produce gene expression quantification from raw sequencing reads. The quantifications can be individualized for multiple cells, multiple samples, or both. Additionally, these tools allow gene expression values to be classified as originating from nascent RNA species or mature RNA species, making this workflow amenable to both cell-based and nucleus-based assays. This protocol describes in detail how to use kallisto and bustools in conjunction with a wrapper, kb-python, to preprocess RNA-seq data.

bioinformatics↗

Fast and scalable querying of eukaryotic linear motifs with gget elm

MotivationEukaryotic linear motifs (ELMs), or Short Linear Motifs (SLiMs), are protein interaction modules that play an essential role in cellular processes and signaling networks and are often involved in diseases like cancer. The ELM database is a collection of manually curated motif knowledge from scientific papers. It has become a crucial resource for cataloging motif biology and recognizing candidate ELMs in novel amino acid sequences. Users can search amino acid sequences or UniProt IDs on the ELM resource web interface. However, as with many web services, there are limitations in the swift processing of large-scale queries through the ELM web interface or API calls, and, therefore, integration into protein function analysis pipelines is limited. ResultsTo allow swift, large-scale motif analyses on protein sequences using ELMs curated on the ELM database, we have developed a Python and command line tool, gget elm, which relies on local computations for efficiently finding candidate ELMs in user-submitted amino acid sequences and UniProt identifiers. gget elm increases accessibility to the information stored in the ELM database and allows scalable searches for motif-mediated interaction sites in the amino acid sequences. Availability and implementationThe manual and source code are available at https://github.com/pachterlab/gget.

bioinformatics↗

Voyager: exploratory single-cell genomics data analysis with geospatial statistics

Exploratory spatial data analysis (ESDA) can be a powerful approach to understanding single-cell genomics datasets, but it is not yet part of standard data analysis workflows. In particular, geospatial analyses, which have been developed and refined for decades, have yet to be fully adapted and applied to spatial single-cell analysis. We introduce the Voyager platform, which systematically brings the geospatial ESDA tradition to (spatial) -omics, with local, bivariate, and multivariate spatial methods not yet commonly applied to spatial -omics, united by a uniform user interface. Using Voyager, we showcase biological insights that can be derived with its methods, such as biologically relevant negative spatial autocorrelation. Underlying Voyager is the SpatialFeatureExperiment data structure, which combines Simple Feature with SingleCellExperiment and AnnData to represent and operate on geometries bundled with gene expression data. Voyager has comprehensive tutorials demonstrating ESDA built on GitHub Actions to ensure reproducibility and scalability, using data from popular commercial technologies. Voyager is implemented in both R/Bioconductor and Python/PyPI, and features compatibility tests to ensure that both implementations return consistent results.

bioinformatics↗

Recovery of a learned behavior despite partial restoration of neuronal dynamics after chronic inactivation of inhibitory neurons

Maintaining motor behaviors throughout life is crucial for an individuals survival and reproductive success. The neuronal mechanisms that preserve behavior are poorly understood. To address this question, we focused on the zebra finch, a bird that produces a highly stereotypical song after learning it as a juvenile. Using cell-specific viral vectors, we chronically silenced inhibitory neurons in the pre-motor song nucleus called the high vocal center (HVC), which caused drastic song degradation. However, after producing severely degraded vocalizations for around 2 months, the song rapidly improved, and animals could sing songs that highly resembled the original. In adult birds, single-cell RNA sequencing of HVC revealed that silencing interneurons elevated markers for microglia and increased expression of the Major Histocompatibility Complex I (MHC I), mirroring changes observed in juveniles during song learning. Interestingly, adults could restore their songs despite lesioning the lateral magnocellular nucleus of the anterior neostriatum (LMAN), a brain nucleus crucial for juvenile song learning. This suggests that while molecular mechanisms may overlap, adults utilize different neuronal mechanisms for song recovery. Chronic and acute electrophysiological recordings within HVC and its downstream target, the robust nucleus of the archistriatum (RA), revealed that neuronal activity in the circuit permanently altered with higher spontaneous firing in RA and lower in HVC compared to control even after the song had fully recovered. Together, our findings show that a complex learned behavior can recover despite extended periods of perturbed behavior and permanently altered neuronal dynamics. These results show that loss of inhibitory tone can be compensated for by recovery mechanisms partly local to the perturbed nucleus and do not require circuits necessary for learning.

neuroscience↗

Selective Serotonin Reuptake Inhibitors Within Cells: Temporal Resolution in Cytoplasm, Endoplasmic Reticulum, and Membrane

Selective serotonin reuptake inhibitors (SSRIs) are the most prescribed treatment for individuals experiencing major depressive disorder (MDD). The therapeutic mechanisms that take place before, during, or after SSRIs bind the serotonin transporter (SERT) are poorly understood, partially because no studies exist of the cellular and subcellular pharmacokinetic properties of SSRIs in living cells. We studied escitalopram and fluoxetine using new intensity- based drug-sensing fluorescent reporters ("iDrugSnFRs") targeted to the plasma membrane (PM), cytoplasm, or endoplasmic reticulum (ER) of cultured neurons and mammalian cell lines. We also employed chemical detection of drug within cells and phospholipid membranes. The drugs attain equilibrium in neuronal cytoplasm and ER, at approximately the same concentration as the externally applied solution, with time constants of a few s (escitalopram) or 200-300 s (fluoxetine). Simultaneously, the drugs accumulate within lipid membranes by [≥] 18-fold (escitalopram) or 180-fold (fluoxetine), and possibly by much larger factors. Both drugs leave cytoplasm, lumen, and membranes just as quickly during washout. We synthesized membrane-impermeant quaternary amine derivatives of the two SSRIs. The quaternary derivatives are substantially excluded from membrane, cytoplasm, and ER for > 2.4 h. They inhibit SERT transport-associated currents 6- or 11-fold less potently than the SSRIs (escitalopram or fluoxetine derivative, respectively), providing useful probes for distinguishing compartmentalized SSRI effects. Although our measurements are orders of magnitude faster than the "therapeutic lag" of SSRIs, these data suggest that SSRI-SERT interactions within organelles or membranes may play roles during either the therapeutic effects or the "antidepressant discontinuation syndrome". SIGNIFICANCE STATEMENTSelective serotonin reuptake inhibitors stabilize mood in several disorders. In general, these drugs bind to the serotonin (5-hydroxytryptamine) transporter (SERT), which clears serotonin from CNS and peripheral tissues. SERT ligands are effective and relatively safe; primary care practitioners often prescribe them. However, they have several side effects and require 2 to 6 weeks of continuous administration until they act effectively. How they work remains perplexing, contrasting with earlier assumptions that the therapeutic mechanism involves SERT inhibition followed by increased extracellular serotonin levels. This study establishes that two SERT ligands, fluoxetine and escitalopram, enter neurons within minutes, while simultaneously accumulating in many membranes. Such knowledge will motivate future research, hopefully revealing where and how SERT ligands "engage" their therapeutic target(s).

neuroscience↗

Efficient querying of genomic databases for single-cell RNA-seq with gget

MotivationA recurring challenge in interpreting genomic data is the assessment of results in the context of existing reference databases. Currently, there is no tool implementing automated, easy programmatic access to curated reference information stored in a diverse collection of large, public genomic databases. Resultsgget is a free and open-source command-line tool and Python package that enables efficient querying of genomic reference databases, such as Ensembl. gget consists of a collection of separate but interoperable modules, each designed to facilitate one type of database querying required for genomic data analysis in a single line of code. AvailabilityThe manual and source code are available at https://github.com/pachterlab/gget. Contactlpachter@caltech.edu

bioinformatics↗

Fluorescence Activation Mechanism and Imaging of Drug Permeation with New Sensors for Smoking-Cessation Ligands

Nicotinic partial agonists provide an accepted aid for smoking cessation and thus contribute to decreasing tobacco-related disease. Improved drugs constitute a continued area of study. However, there remains no reductionist method to examine the cellular and subcellular pharmacokinetic properties of these compounds in living cells. Here, we developed new intensity-based drug sensing fluorescent reporters ("iDrugSnFRs") for the nicotinic partial agonists dianicline, cytisine, and two cytisine derivatives - 10-fluorocytisine and 9-bromo-10-ethylcytisine. We report the first atomic-scale structures of liganded periplasmic binding protein-based biosensors, accelerating development of iDrugSnFRs and also explaining the activation mechanism. The nicotinic iDrugSnFRs detect their drug partners in solution, as well as at the plasma membrane (PM) and in the endoplasmic reticulum (ER) of cell lines and mouse hippocampal neurons. At the PM, the speed of solution changes limits the growth and decay rates of the fluorescence response in almost all cases. In contrast, we found that rates of membrane crossing differ among these nicotinic drugs by > 30 fold. The new nicotinic iDrugSnFRs provide insight into the real-time pharmacokinetic properties of nicotinic agonists and provide a methodology whereby iDrugSnFRs can inform both pharmaceutical neuroscience and addiction neuroscience.

pharmacology and toxicology↗