bioRxiv Science⌕ Search

Biology subjects

Das Adhikari, S.

Publications and source records attributed to Das Adhikari, S..

3 recordsLinked to original sources

SPACE: Spatially variable gene clustering adjusting for cell type effect for improved spatial domain detection.

Recent advances in spatial transcriptomics have significantly deepened our understanding of biology. A primary focus has been identifying spatially variable genes (SVGs) which are crucial for downstream tasks like spatial domain detection. Traditional methods often use all or a set number of top SVGs for this purpose. However, in diverse datasets with many SVGs, this approach may not ensure accurate results. Instead, grouping SVGs by expression patterns and using all SVG groups in downstream analysis can improve accuracy. Furthermore, classifying SVGs in this manner is akin to identifying cell type marker genes, offering valuable biological insights. The challenge lies in accurately categorizing SVGs into relevant clusters, aggravated by the absence of prior knowledge regarding the number and spectrum of spatial gene patterns. Addressing this challenge, we propose SPACE, SPatially variable gene clustering Adjusting for Cell type Effect, a framework that classifies SVGs based on their spatial patterns by adjusting for confounding effects caused by shared cell types, to improve spatial domain detection. This method does not require prior knowledge of gene cluster numbers, spatial patterns, or cell type information. Our comprehensive simulations and real data analyses demonstrate that SPACE is an efficient and promising tool for spatial transcriptomics analysis. Key PointsO_LISPACE eliminates the need for prior knowledge about the number of gene clusters, known cell types, or the quantity of SVGs to identify clusters for downstream analysis. C_LIO_LISPACE offers a method to effectively leverage SVGs for low-dimensional embedding within each cluster to improve the accuracy of spatial domain detection. C_LIO_LIThe efficiency and utility of the SPACE algorithm have been validated across multiple datasets and simulations, demonstrating its effectiveness in producing meaningful and interpretable results. C_LI

genomics↗

BayesKAT: Bayesian Optimal Kernel-based Test for genetic association studies reveals joint genetic effects in complex diseases

GWAS methods have identified individual SNPs significantly associated with specific phenotypes. Nonetheless, many complex diseases are polygenic and are controlled by multiple genetic variants that are usually non-linearly dependent. These genetic variants are marginally less effective and remain undetected in GWAS analysis. Kernel-based tests (KBT), which evaluate the joint effect of a group of genetic variants, are therefore critical for complex disease analysis. However, choosing different kernel functions in KBT can significantly influence the type I error control and power, and selecting the optimal kernel remains a statistically challenging task. A few existing methods suffer from inflated type 1 errors, limited scalability, inferior power, or issues of ambiguous conclusions. Here, we present a new Bayesian framework, BayesKAT(https://github.com/wangjr03/BayesKAT), which overcomes these kernel specification issues by selecting the optimal composite kernel adaptively from the data while testing genetic associations simultaneously. Furthermore, BayesKAT implements a scalable computational strategy to boost its applicability, especially for high-dimensional cases where other methods become less effective. Based on a series of performance comparisons using both simulated and real large-scale genetics data, BayesKAT outperforms the available methods in detecting complex group-level associations and controlling type I errors simultaneously. Applied on a variety of groups of functionally related genetic variants based on biological pathways, co-expression gene modules, and protein complexes, BayesKAT deciphers the complex genetic basis and provides mechanistic insights into human diseases.

genetics↗

A data-driven evaluation of Arabidopsis-centric research and the model species concept

The selection of Arabidopsis as a model organism played a pivotal role in advancing genomic science, firmly establishing the cornerstone of today s plant molecular biology. Competing frameworks to select an agricultural- or ecological-based model species, or to decentralize plant science and study a multitude of diverse species, were selected against in favor of building core knowledge in a species that would facilitate genome-enabled research that could assumedly be transferred to other plants. Here, we examine the ability of models based on Arabidopsis gene expression data to predict tissue identity in other flowering plant species. Comparing different machine learning algorithms, models trained and tested on Arabidopsis data achieved near perfect precision and recall values using the K-Nearest Neighbor method, whereas when tissue identity is predicted across the flowering plants using models trained on Arabidopsis data, precision values range from 0.69 to 0.74 and recall from 0.54 to 0.64, depending on the algorithm used. Below-ground tissue is more predictable than other tissue types, and the ability to predict tissue identity is not correlated with phylogenetic distance from Arabidopsis. This suggests that gene expression signatures rather than marker genes are more valuable to create models for tissue and cell type prediction in plants. Our data-driven results highlight that, in hindsight, the assertion that knowledge from Arabidopsis is translatable to other plants is not always true. Considering the current landscape of abundant sequencing data and computational resources, it may be prudent to reevaluate the scientific emphasis on Arabidopsis and to prioritize the exploration of plant diversity.

plant biology↗