bioRxiv ScienceSearch

Biology subjects

Lagergren, J.

Publications and source records attributed to Lagergren, J..

4 recordsLinked to original sources

Charting Tissue Expression Anatomy by Spatial Transcriptome Deconvolution

We create data-driven maps of transcriptomic anatomy with a probabilistic framework for unsupervised pattern discovery in spatial gene expression data. With convolved negative binomial regression we discover patterns which correspond to cell types, microenvironments, or tissue components, and that consist of gene expression profiles and spatial activity maps. Expression profiles quantify how strongly each gene is expressed in a given pattern, and spatial activity maps reflect where in space each pattern is active. Arbitrary covariates and prior hierarchies are supported to leverage complex experimental designs.\n\nWe demonstrate the method with Spatial Transcriptomics data of mouse brain and olfactory bulb. The discovered transcriptomic patterns correspond to neuroanatomically distinct cell layers. Moreover, batch effects are successfully addressed, leading to consistent pattern inference for multi-sample analyses. On this basis, we identify known and uncharacterized genes that are spatially differentially expressed in the hippocampal field between Ammons horn and the dentate gyrus.

bioinformatics

SCuPhr: A Probabilistic Framework for Cell Lineage Tree Reconstruction

Cell lineage tree reconstruction methods are developed for various tasks, such as investigating the development, differentiation, and cancer progression. Single-cell sequencing technologies enable more thorough analysis with higher resolution. We present Scuphr, a distance-based cell lineage tree reconstruction method using bulk and single-cell DNA sequencing data from healthy tissues. Common challenges of single-cell DNA sequencing, such as allelic dropouts and amplification errors, are included in Scuphr. Scuphr computes the distance between cell pairs and reconstructs the lineage tree using the neighbor-joining algorithm. With its embarrassingly parallel design, Scuphr can do faster analysis than the state-of-the-art methods while obtaining better accuracy. The methods robustness is investigated using various synthetic datasets and a biological dataset of 18 cells. Author summaryCell lineage tree reconstruction carries a significant potential for studies of development and medicine. The lineage tree reconstruction task is especially challenging for cells taken from healthy tissue due to the scarcity of mutations. In addition, the single-cell whole-genome sequencing technology introduces artifacts such as amplification errors, allelic dropouts, and sequencing errors. We propose Scuphr, a probabilistic framework to reconstruct cell lineage trees. We designed Scuphr for single-cell DNA sequencing data; it accounts for technological artifacts in its graphical model and uses germline heterozygous sites to improve its accuracy. Scuphr is embarrassingly parallel; the speed of the computational analysis is inversely proportional to the number of available computational nodes. We demonstrated that Scuphr is fast, robust, and more accurate than the state-of-the-art method with the synthetic data experiments. Moreover, in the biological data experiment, we showed Scuphr successfully identifies different clones and further obtains more support on closely related cells within clones.

bioinformatics

Species tree-aware simultaneous reconstruction of gene and domain evolution

Most genes are composed of multiple domains, with a common evolutionary history, that typically perform a specific function in the resulting protein. As witnessed by many studies of key gene families, it is important to understand how domains have been duplicated, lost, transferred between genes, and rearranged. Analogously to the case of evolutionary events affecting entire genes, these domain events have large consequences for phylogenetic reconstruction and, in addition, they create considerable obstacles for gene sequence alignment algorithms, a prerequisite for phylogenetic reconstruction.\n\nWe introduce the DomainDLRS model, a hierarchical, generative probabilistic model containing three levels corresponding to species, genes, and domains, respectively. From a dated species tree, a gene tree is generated according to the DL model, which is a birth-death model generalized to occur in a dated tree. Then, from the dated gene tree, a pre-specified number of dated domain trees are generated using the DL model and the molecular clock is relaxed, effectively converting edge times to edge lengths. Finally, for each domain tree and its lengths, domain sequences are generated for the leaves based on a selected model of sequence evolution.\n\nFor this model, we present a MCMC-based inference framework called DomainDLRS that takes a dated species tree together with a multiple sequence alignment for each domain family as input and outputs an estimated posterior distribution over reconciled gene and domain trees. By requiring aligned domains rather than genes, our framework evades the problem of aligning full-length genes that have been exposed to domain duplications, in particular non-tandem domain duplications. We show that DomainDLRS performs better than MrBayes on synthetic data and that it outperforms MrBayes on biological data. We analyse several zincfinger genes and show that most domain duplications have been tandem duplications, some involving two or more domains, but non-tandem duplications have also been common.

bioinformatics

Patient-specific detection of cancer genes reveals recurrently perturbed processes in esophageal adenocarcinoma

The identification of somatic alterations with a cancer promoting role is challenging in highly unstable and heterogeneous cancers, such as esophageal adenocarcinoma (EAC). Here we developed a machine learning algorithm to identify cancer genes in individual patients considering all types of damaging alterations simultaneously (mutations, copy number alterations and structural rearrangements). Analysing 261 EACs from the OCCAMS Consortium, we discovered a large number of novel cancer genes that, together with well-known drivers, help promote cancer. Validation using 107 additional EACs confirmed the robustness of the approach. Unlike known drivers whose alterations recur across patients, the large majority of the newly discovered cancer genes are rare or patient-specific. Despite this, they converge towards perturbing cancer-related processes, including intracellular signalling, cell cycle regulation, proteasome activity and Toll-like receptor signalling. Recurrence of process perturbation, rather than individual genes, divides EACs into six clusters that differ in their molecular and clinical features and suggest patient stratifications for personalised treatments. By experimentally mimicking or reverting alterations of predicted cancer genes, we validated their contribution to cancer progression and revealed EAC acquired dependencies, thus demonstrating their potential as therapeutic targets.

cancer biology