bioRxiv Science⌕ Search

Biology subjects

Carilli, M.

Publications and source records attributed to Carilli, M..

5 recordsLinked to original sources

Hybrid crosses reveal a cell-type-specific landscape of mouse regulatory variation

Understanding the genetic architecture of gene expression is fundamental to evolutionary biology and medicine. As part of the IGVF Consortium, we present a single-nucleus RNA-seq resource of 5.3 million nuclei across eight tissue groups, featuring seven F1 hybrids from C57BL/6J dams crossed with the other Collaborative Cross founder strains for comparison against parental strains. We identify 25,864 genes (91% of those detected) exhibiting non-conserved regulatory behavior in at least one of 92 cell types in one or more crosses. Our results show that while cis-acting variation primarily drives divergence, trans-acting effects are substantially more cell-type specific and sensitive to tissue environment. Notably, bulk tissue analyses frequently mask these signals, particularly in smaller populations such as astrocytes. Furthermore, increasing genetic divergence primarily expands the landscape of cis-acting variation, while trans-acting effects remain stable across genetic distances within species. This atlas establishes a foundational framework for decoding the complex interplay between genetic variation and cell-type-specific regulation across the mammalian body.

genomics↗

Foundation Models Improve Perturbation Response Prediction

Predicting cellular responses to genetic or chemical perturbations has been a long-standing goal in biology. Recent applications of foundation models to this task have yielded contradictory results regarding their superiority over simple baselines. We conducted an extensive analysis of over 600 different models across various prediction tasks and evaluation metrics, demonstrating that while some foundation models fail to outperform simple baselines, others significantly improve predictions for both genetic and chemical perturbations. Furthermore, we developed and evaluated methods for integrating multiple foundation models for perturbation prediction. Our results show that with sufficient data, these models approach fundamental performance limits, confirming that foundation models can improve cellular response simulations. Code and Data: https://github.com/genbio-ai/foundation-models-perturbation

bioinformatics↗

Systematic cell-type resolved transcriptomes of 8 tissues in 8 lab and wild-derived mouse strains captures global and local expression variation

Mapping the impact of genomic variation on gene expression facilitates an understanding of the molecular basis of complex phenotypic traits and disease predisposition. Mouse models provide a controlled and reproducible framework for capturing the breadth of genomic variation observed in different genotypes across a wide variety of tissues. As part of the IGVF consortiums effort to catalog the effects of genetic variation, we uniformly characterized the transcriptomes of eight tissues from each mouse founder strain used to derive the Collaborative Cross strains, comprising five classical laboratory inbred strains and three wild-derived inbred strains. We sequenced samples from four male and four female replicates per tissue using single-nucleus RNA-seq to generate an "8-cube" dataset of 5.2 million nuclei across 106 cell types and cell states. As expected, the overall extent of transcriptome variation correlates positively with genetic divergence across the strains with the greatest differential between PWK/PhJ and CAST/EiJ. At the individual tissue level, heart and brain are relatively more similar across strains compared with gonads, adrenal, skeletal muscle, kidney, and liver. Further analyses revealed substantial strain variation, often concentrated in a few cell types as well as cell-state signatures that especially reflect strain-associated immune and metabolic trait differences. The founder 8-cube dataset provides rich transcriptome variation signatures to help explain strain-specific phenotypic traits and disease states, as illustrated by examples in tissue-resident immune cells, muscle degeneration, kidney sex differences, and the hypothalamicpituitary-adrenal axis. This data further provides a systematic foundation for the analysis of these tissues in the founder strains as well as the Collaborative Cross.

genomics↗

Estimating cis and trans contributions todifferences in gene regulation

We describe a coordinate system and associated hypothesis testing framework for determining whether cis or trans regulation is responsible for differences in gene expression between two homozygous strains or species. We apply our framework to data from single replicate studies on yeast strains and human-chimpanzee hybrid cells, as well as to data from a mouse study with replicates, showing marked differences between our gene regulatory assignments and those previously reported. We also show how our multi-sample framework can determine the context dependency of cis and trans effects as well as explicitly model different hypotheses regarding the underlying mechanism of trans regulation.

genetics↗

Efficient and accurate detection of viral sequences at single-cell resolution reveals novel viruses perturbing host gene expression

There are an estimated 300,000 mammalian viruses from which infectious diseases in humans may arise. They inhabit human tissues such as the lungs, blood, and brain and often remain undetected. Efficient and accurate detection of viral infection is vital to understanding its impact on human health and to make accurate predictions to limit adverse effects, such as future epidemics. The increasing use of high-throughput sequencing methods in research, agriculture, and healthcare provides an opportunity for the cost-effective surveillance of viral diversity and investigation of virus-disease correlation. However, existing methods for identifying viruses in sequencing data rely on and are limited to reference genomes or cannot retain single-cell resolution through cell barcode tracking. We introduce a method that accurately and rapidly detects viral sequences in bulk and single-cell transcriptomics data based on highly conserved amino acid domains, which enables the detection of RNA viruses covering over 100,000 virus species. The analysis of viral presence and host gene expression in parallel at single-cell resolution allows for the characterization of host viromes and the identification of viral tropism and host responses. We applied our method to identify putative novel viruses in rhesus macaque PBMC data that display cell type specificity and whose presence correlates with altered host gene expression.

bioinformatics↗