bioRxiv Science⌕ Search

Biology subjects

Bromberek, S.

Publications and source records attributed to Bromberek, S..

3 recordsLinked to original sources

Integration of cell-specific gene expression and chromatin accessibility facilitates localization of neurodegenerative risk in microglia

Genome-wide association studies (GWAS) have identified many loci that contribute to the risk of neurodegenerative diseases. However, a persistent challenge in interpretation of GWAS is to break loci down to specific genes, variants, and cell types, and thus nominate disease mechanisms. Here, we used iPSC-derived cells containing population-level variation to examine GWAS loci across NDDs including Alzheimer's disease, Parkinson's disease and Lewy body dementia. We differentiated a set of 135 iPSC donor lines into two cell types relevant to neurodegeneration, neurons and microglia, and completed single cell gene expression and chromatin accessibility profiling. Meta-analysis of these data with published human brain snRNAseq for QTL mapping identified multiple loci associated with NDDs that are restricted to either neurons or microglia. Colocalization of GWAS and these QTL supports microglia as having a strong contribution to disease risk. We tested peaks nominated at the BIN1 locus for enhancer activity using a perturb-seq-based method in microglia. Our results show one of the nominated peaks controls BIN1 expression in microglia but also modifies expression of other genes at the locus. These results support the hypothesis that common variants affecting gene expression specifically in microglia can contribute directly to NDD risk rather than functioning solely as a secondary response to neurodegeneration. These data also show that iPSC-derived cells are a useful model to experimentally dissect GWAS loci that colocalize with QTL.

genomics↗

Long-read sequencing maps transposable element variation and its regulatory and epigenetic effects in the human brain

Transposable elements (TEs) are mobile DNA sequences that shape genome architecture and gene regulation, yet their roles in the human brain remain largely unresolved. Short-read sequencing lacks the resolution to accurately map TE insertions, detect associated structural variants, and resolve highly repetitive regions. Here, we leverage long-read whole-genome sequencing to profile germline TE insertions in postmortem brain tissue from two ancestrally diverse cohorts: the North American Brain Expression Consortium (NABEC; European ancestry, n = 205) and the Human Brain Collection Core (HBCC; African and African-admixed ancestry, n = 146). We identified 2,842 and 1,660 high-confidence non-reference insertions in HBCC and NABEC, respectively, spanning Alu, LINE-1, and SVA elements. We then also further characterized complex short tandem repeat and variable number tandem repeat variation within reference SVA and Alu loci. Reference TEs were also found to mediate complex structural variants at loci implicated in brain development and neurodegenerative disease, with several showing ancestry-specific patterns. Integration of bulk RNA-sequencing data identified TE expression quantitative trait loci, including insertions that modulate neuronal gene expression. Single-nucleus RNA sequencing revealed cell-type-specific effects of TE regulation across cortical populations. Long-read methylation profiling further demonstrated age-associated epigenetic regulation of both reference and non-reference Alu elements. As a community resource, we release a catalog of TE insertions, allele frequencies, and ancestry-specific distributions to enable future functional and disease-focused investigations. Together, these findings highlight the widespread regulatory and epigenetic influence of TEs in the human brain and establish long-read sequencing as a powerful approach for uncovering cell-type- and population-specific TE dynamics.

genomics↗

Accurate strand-specific long-read transcript isoform discovery and quantification at bulk, single-cell, and single-nucleus resolution

Recent advances in long-read transcriptome sequencing enable high-throughput profiling of full-length RNA isoforms in bulk, single-cell, and single-nucleus samples. However, long-read datasets typically contain a mixture of complete and partial transcripts, leading to pervasive ambiguity in read-to-isoform assignment and complicating accurate isoform identification and quantification, particularly in the absence of reliable reference annotations. These challenges are further amplified in single-cell and single-nucleus samples, where coverage is sparse and transcriptional heterogeneity is high. Here, we present the Long Read Alignment Assembler (LRAA), a unified and versatile computational framework for isoform identification and quantification from long-read RNA sequencing data across bulk, single-cell, and single-nucleus transcriptomic samples. LRAA combines splice-graph based structural modeling with expectation maximization based optimization to probabilistically resolve ambiguous read assignments and improve isoform abundance estimation. The framework supports quantification-only, reference-guided, and fully reference-free (de novo) modes of analysis within a single methodological paradigm. We benchmarked LRAA using both simulated and genuine long-read datasets spanning sequencing standards and whole transcriptomes. Central to this evaluation is a novel benchmarking strategy based on Multiplexed Overexpression of Regulatory Factors (MORFs), which provides biologically expressed, barcoded isoforms with unambiguous read-level ground truth. Across all benchmarks, including MORFs, synthetic spike-ins, and whole-transcriptome datasets, LRAA consistently outperformed state-of-the-art methods in isoform identification accuracy, sensitivity, and expression quantification. Finally, we demonstrate the biological utility of LRAA by resolving cell-type-specific isoform usage across peripheral blood immune cell populations and by detecting a pathogenic cryptic isoform of STMN2 with associated transcriptional changes in single-nucleus RNA-seq data from frontal cortex tissue of an individual with frontotemporal dementia (FTD). Together, these results establish LRAA as a robust and general solution for resolving transcript diversity in complex biological systems, from development to disease.

bioinformatics↗