bioRxiv ScienceSearch

Biology subjects

Bahlo, M.

Publications and source records attributed to Bahlo, M..

7 recordsLinked to original sources

SIS-seq, a molecular ‘time machine’, connects single cell fate with gene programs

Conventional single cell RNA-seq methods are destructive, such that a given cell cannot also then be tested for fate and function, without a time machine. Here, we develop a clonal method SIS-seq, whereby single cells are allowed to divide, and progeny cells are assayed separately in SISter conditions; some for fate, others by RNA-seq. By cross-correlating progenitor gene expression with mature cell fate within a clone, and doing this for many clones, we can identify the earliest gene expression signatures of dendritic cell subset development. SIS-seq could be used to study other populations harboring clonal heterogeneity, including stem, reprogrammed and cancer cells to reveal the transcriptional origins of fate decisions.

systems biology

dtangle: accurate and fast cell-type deconvolution

MotivationUnderstanding cell type composition is important to understanding many biological processes. Furthermore, in gene expression studies cell type composition can confound differential expression analysis (DEA). To aid understanding cell type composition, methods of estimating (deconvolving) cell type proportions from gene expression data have been developed.\n\nResultsWe propose dtangle, a new cell-type deconvolution method. dtangle works on a range of DNA microarray and bulk RNA-seq platforms. It estimates cell-type proportions using publicly available, often cross-platform, reference data. To comprehensively evaluate dtangle, we assemble ten benchmark data sets. Here, dtangle is competitive with published deconvolution methods, is robust to selection of tuning parameters and is quicker than other methods. As a case study, we investigate the human immune response to Lyme disease. dtangles estimates reveal a temporal trend consistent with previous findings and are important covariates for DEA across disease status.\n\nAvailabilitydtangle is on CRAN (cran.r-project.org/package=dtangle) or github (dtangle.github.io).\n\nContactgjhunt@umich.edu

bioinformatics

Functional analysis of a hypomorphic allele shows that MMP14 catalytic activity is the prime determinant of the Winchester syndrome phenotype

Winchester syndrome (WS, MIM #277950) is an extremely rare autosomal recessive skeletal dysplasia characterized by progressive joint destruction and osteolysis. To date, only one missense mutation in MMP14, encoding the membrane-bound matrix metalloprotease 14, has been reported in WS patients. Here, we report a novel hypomorphic MMP14 p.Arg111His (R111H) allele, associated with a mitigated form of WS. Functional analysis demonstrated that this mutation, in contrast to previously reported human and murine MMP14 mutations, does not affect MMP14s transport to the cell membrane. Instead, it partially impairs MMP14s proteolytic activity. This residual activity likely accounts for the mitigated phenotype observed in our patients. Based on our observations as well as previously published data, we hypothesize that MMP14s catalytic activity is the prime determinant of disease severity. Given the limitations of our in vitro assays in addressing the consequences of MMP14 dysfunction, we generated a novel mmp14a/b knockout zebrafish model. The fish accurately reflected key aspects of the WS phenotype including craniofacial malformations, kyphosis, short-stature and reduced bone density due to defective collagen remodeling. Notably, the zebrafish model will be a valuable tool for developing novel therapeutic approaches to a devastating bone disorder.

genetics

Cluster Headache: Comparing Clustering Tools for 10X Single Cell Sequencing Data

The commercially available 10X Genomics protocol to generate droplet-based single cell RNA-seq (scRNA-seq) data is enjoying growing popularity among researchers. Fundamental to the analysis of such scRNA-seq data is the ability to cluster similar or same cells into non-overlapping groups. Many competing methods have been proposed for this task, but there is currently little guidance with regards to which method offers most accuracy. Answering this question is complicated by the fact that 10X Genomics data lack cell labels that would allow a direct performance evaluation. Thus in this review, we focused on comparing clustering solutions of a dozen methods for three datasets on human peripheral mononuclear cells generated with the 10X Genomics technology. While clustering solutions appeared robust, we found that solutions produced by different methods have little in common with each other. They also failed to replicate cell type assignment generated with supervised labeling approaches. Furthermore, we demonstrate that all clustering methods tested clustered cells to a large degree according to the amount of genes coding for ribosomal protein genes in each cell.

bioinformatics

Detecting known repeat expansions with standard protocol next generation sequencing, towards developing a single screening test for neurological repeat expansion disorders

Repeat expansions cause over 30, predominantly neurogenetic, inherited disorders. These can present with overlapping clinical phenotypes, making molecular diagnosis challenging. Single gene or small panel PCR-based methods are employed to identify the precise genetic cause, but can be slow and costly, and often yield no result. Genomic analysis via whole exome and whole genome sequencing (WES and WGS) is being increasingly performed to diagnose genetic disorders. However, until recently analysis protocols could not identify repeat expansions in these datasets.\n\nA new method, called exSTRa (expanded Short Tandem Repeat algorithm) for the identification of repeat expansions using either WES or WGS was developed and performance of exSTRa was assessed in a simulation study. In addition, four retrospective cohorts of individuals with eleven different known repeat expansion disorders were analysed with the new method. Results were assessed by comparing to known disease status. Performance was also compared to three other analysis methods (ExpansionHunter, STRetch and TREDPARSE), which were developed specifically for WGS data. Expansions in the STR loci assessed were successfully identified in WES and WGS datasets by all four methods, with high specificity and sensitivity, excepting the FRAXA STR where expansions were unlikely to be detected. Overall exSTRa demonstrated more robust/superior performance for WES data in comparison to the other three methods. exSTRa can be applied to existing WES or WGS data to identify likely repeat expansions and can be used to investigate any STR of interest, by specifying location and repeat motif. We demonstrate that methods such as exSTRa can be effectively utilized as a screening tool to interrogate WES data generated with PCR-based library preparations and WGS data generated using either PCR-based or PCR-free library protocols, for repeat expansions which can then be followed up with specific diagnostic tests. exSTRa is available via GitHub (https://github.com/bahlolab/exSTRa).

bioinformatics

Long-term sustained malaria control leads to inbreeding and fragmentation of Plasmodium vivax populations

The human malaria parasite Plasmodium vivax is resistant to malaria control strategies maintaining high genetic diversity even when transmission is low. To investigate whether declining P. vivax transmission leads to increasing P. vivax population structure that would facilitate elimination, we genotyped samples from a wide range of transmission intensities and spatial scales in the Southwest Pacific, including two time points at one site (Tetere, Solomon Islands) during intensified control. Analysis of 887 P. vivax microsatellite haplotypes from hyperendemic Papua New Guinea (PNG, n = 443), meso-hyperendemic Solomon Islands (n= 420), and hypoendemic Vanuatu (n=24) revealed increasing population structure and multilocus linkage disequilibrium and a modest decline in diversity as transmission decreases over space and time. In Solomon Islands, which has had sustained control efforts for 20 years, and Vanuatu, which has experienced sustained low transmission for many years, significant population structure was observed at different spatial scales. We conclude that control efforts will eventually impact P. vivax population structure and with sustained pressure, populations may eventually fragment into a limited number of clustered foci that could be targeted for elimination.

epidemiology

Detecting Selection Signals In Plasmodium falciparum Using Identity-By-Descent Analysis

Identification of genomic regions that are identical by descent (IBD) has proven useful for human genetic studies where analyses have led to the discovery of familial relatedness and fine-mapping of disease critical regions. Unfortunately however, IBD analyses have been underutilized inanalysis of other organisms, including human pathogens. This is in part due to the lack of statistical methodologies for non-diploid genomes in addition to the added complexity of multiclonal infections. As such, we have developed an IBD methodology, called isoRelate, for analysis of haploid recombining microorganisms in the presence of multiclonal infections. Using the inferred IBD status at genomic locations, we have also developed a novel statistic for identifying loci under positive selection and propose relatedness networks as a means of exploring shared haplotypes within populations. We evaluate the performance of our methodologies for detecting IBD and selection, including comparisons with existing tools, then perform an exploratory analysis of whole genome sequencing data from a global Plasmodium falciparum dataset of more than 2500 genomes. This analysis identifies Southeast Asia as havingmany highly related isolates, possibly as a result of both reduced transmission from intensified control efforts and population bottlenecks following the emergence of antimalarial drug resistance. Many signals of selection are also identified, most of which overlap genes that are known to be associated with drug resistance, in addition to two novel signals observed in multiple countries that have yet to be explored in detail. Additionally, we investigate relatedness networks over the selected loci and determine that one of these sweeps has spread between continents while the other has arisen independently in different countries. IBD analysis of microorganisms using isoRelate can be used for exploring population structure, positive selection and haplotype distributions, and will be a valuable tool for monitoring disease control and elimination efforts of many diseases.

bioinformatics