bioRxiv Science⌕ Search

Biology subjects

Filimban, G.

Publications and source records attributed to Filimban, G..

3 recordsLinked to original sources

Hybrid crosses reveal a cell-type-specific landscape of mouse regulatory variation

Understanding the genetic architecture of gene expression is fundamental to evolutionary biology and medicine. As part of the IGVF Consortium, we present a single-nucleus RNA-seq resource of 5.3 million nuclei across eight tissue groups, featuring seven F1 hybrids from C57BL/6J dams crossed with the other Collaborative Cross founder strains for comparison against parental strains. We identify 25,864 genes (91% of those detected) exhibiting non-conserved regulatory behavior in at least one of 92 cell types in one or more crosses. Our results show that while cis-acting variation primarily drives divergence, trans-acting effects are substantially more cell-type specific and sensitive to tissue environment. Notably, bulk tissue analyses frequently mask these signals, particularly in smaller populations such as astrocytes. Furthermore, increasing genetic divergence primarily expands the landscape of cis-acting variation, while trans-acting effects remain stable across genetic distances within species. This atlas establishes a foundational framework for decoding the complex interplay between genetic variation and cell-type-specific regulation across the mammalian body.

genomics↗

Systematic cell-type resolved transcriptomes of 8 tissues in 8 lab and wild-derived mouse strains captures global and local expression variation

Mapping the impact of genomic variation on gene expression facilitates an understanding of the molecular basis of complex phenotypic traits and disease predisposition. Mouse models provide a controlled and reproducible framework for capturing the breadth of genomic variation observed in different genotypes across a wide variety of tissues. As part of the IGVF consortiums effort to catalog the effects of genetic variation, we uniformly characterized the transcriptomes of eight tissues from each mouse founder strain used to derive the Collaborative Cross strains, comprising five classical laboratory inbred strains and three wild-derived inbred strains. We sequenced samples from four male and four female replicates per tissue using single-nucleus RNA-seq to generate an "8-cube" dataset of 5.2 million nuclei across 106 cell types and cell states. As expected, the overall extent of transcriptome variation correlates positively with genetic divergence across the strains with the greatest differential between PWK/PhJ and CAST/EiJ. At the individual tissue level, heart and brain are relatively more similar across strains compared with gonads, adrenal, skeletal muscle, kidney, and liver. Further analyses revealed substantial strain variation, often concentrated in a few cell types as well as cell-state signatures that especially reflect strain-associated immune and metabolic trait differences. The founder 8-cube dataset provides rich transcriptome variation signatures to help explain strain-specific phenotypic traits and disease states, as illustrated by examples in tissue-resident immune cells, muscle degeneration, kidney sex differences, and the hypothalamicpituitary-adrenal axis. This data further provides a systematic foundation for the analysis of these tissues in the founder strains as well as the Collaborative Cross.

genomics↗

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput "short-read" technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive fulllength, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. "Long-read" sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

bioinformatics↗