bioRxiv Science⌕ Search

Biology subjects

Paquette, K.

Publications and source records attributed to Paquette, K..

5 recordsLinked to original sources

SALRR: Scalable Analysis of Long-Read RNA-Seq Enables Comprehensive Transcriptome Profiling in Human Brain

Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.

genomics↗

Long-read sequencing maps transposable element variation and its regulatory and epigenetic effects in the human brain

Transposable elements (TEs) are mobile DNA sequences that shape genome architecture and gene regulation, yet their roles in the human brain remain largely unresolved. Short-read sequencing lacks the resolution to accurately map TE insertions, detect associated structural variants, and resolve highly repetitive regions. Here, we leverage long-read whole-genome sequencing to profile germline TE insertions in postmortem brain tissue from two ancestrally diverse cohorts: the North American Brain Expression Consortium (NABEC; European ancestry, n = 205) and the Human Brain Collection Core (HBCC; African and African-admixed ancestry, n = 146). We identified 2,842 and 1,660 high-confidence non-reference insertions in HBCC and NABEC, respectively, spanning Alu, LINE-1, and SVA elements. We then also further characterized complex short tandem repeat and variable number tandem repeat variation within reference SVA and Alu loci. Reference TEs were also found to mediate complex structural variants at loci implicated in brain development and neurodegenerative disease, with several showing ancestry-specific patterns. Integration of bulk RNA-sequencing data identified TE expression quantitative trait loci, including insertions that modulate neuronal gene expression. Single-nucleus RNA sequencing revealed cell-type-specific effects of TE regulation across cortical populations. Long-read methylation profiling further demonstrated age-associated epigenetic regulation of both reference and non-reference Alu elements. As a community resource, we release a catalog of TE insertions, allele frequencies, and ancestry-specific distributions to enable future functional and disease-focused investigations. Together, these findings highlight the widespread regulatory and epigenetic influence of TEs in the human brain and establish long-read sequencing as a powerful approach for uncovering cell-type- and population-specific TE dynamics.

genomics↗

Haplotype-Resolved DNA Methylation at the APOE Locus identifies Allele-Specific Epigenetic Signatures Relevant to Alzheimer's Disease Risk

The APOE gene encodes a key lipid transport protein and plays a central role in Alzheimers disease (AD) pathogenesis. Three common APOE alleles, {varepsilon}2 (rs7412(C>T), {varepsilon}3 (reference), and {varepsilon}4 (rs429358(T>C)), arise from two coding variants in exon 4 and confer distinct AD risk profiles, with {varepsilon}4 increasing risk and {varepsilon}2 providing protection. The {varepsilon}3-linked APOE variant rs769455[T] has also been associated with elevated AD risk in individuals of African ancestry carrying both rs769455[T] and {varepsilon}4 alleles. These single nucleotide variants (SNVs) reside in a cytosine-phosphate-guanine (CpG) island, which is a region with a higher frequency of CpG sites compared to the rest of the genome. CpG sites are subject to 5-methylcytosine (5mC) methylation by DNA methyltransferases which add a methyl group to the fifth carbon on the cytosine residue of a CpG site. The presence of SNVs can disrupt this process, making these regions prime targets for differential methylation; however, allele-specific methylation patterns in APOE remain poorly resolved due to technical limitations of conventional bisulfite and methylation array based methods, including degraded DNA quality, sparse CpG coverage, and lack of haplotype phasing. Here, we leverage high-accuracy long-read sequencing data to generate haplotype-resolved methylation profiles of the APOE locus in 332 postmortem brain samples from two ancestrally different cohorts. This includes 201 individuals of European ancestry from the North American Brain Expression Consortium (NABEC), comprising 402 haplotypes (48 {varepsilon}2 and 58 {varepsilon}4 alleles), and 131 individuals of African and African admixed ancestry from the Human Brain Core Collection (HBCC), comprising 262 haplotypes (25 {varepsilon}2, 64 {varepsilon}4, and 7 rs769455 alleles). A linear regression analysis identified 18 novel differentially methylated CpG sites (DMCs) associated with APOE {varepsilon}2, {varepsilon}4, and rs769455 within a gene cluster spanning TOMM40, APOE, APOC1, and APOC4-APOC2. This represents the most comprehensive haplotype-resolved methylation study of APOE in human brain tissue to date. Our results uncover distinct allele-specific methylation signatures and demonstrate the power of long-read sequencing for resolving epigenetic variation relevant to AD risk.

genomics↗

Dissecting the biological impact of GBA1 mutations using multi-omics in an isogenic setting

GBA1 is a risk gene for multiple neurodegenerative diseases, including Lewy Body Dementia and Parkinsons disease, and biallelic pathogenic variants in the gene result in the lysosomal storage disorder Gaucher disease. GBA1 encodes the enzyme glucocerebrosidase (GCase), and alterations in the gene result in reduced enzymatic activity, which affects lysosome function downstream. Induced pluripotent stem cells (iPSCs) are a useful tool for testing the functional consequences of gene variants in an isogenic setting. Additionally, they can be used to perform multiomic studies to explore biological effects independent of disease mechanisms. Using CRISPR-edited isogenic KOLF2.1J iPSC lines containing pathogenic GBA1 variants D409H (p.D448H), D409V (p.D448V) and GBA1 knockout line generated by the iPSC Neurodegenerative Disease Initiative (iNDI), we examined potential molecular mechanisms and downstream consequences of GCase reduction. In this study, we confirm that this isogenic series behaves as expected for loss of function variants, despite the known difficulties with GBA1 editing. We identified that there are limited overlapping results across cell types suggesting potential different downstream effects caused by GBA1 variants. Additionally, we note that RNA-based quantitation may not be the best method to characterize GCase mechanisms, but protein and metabolomic analyses may be used to evaluate differences across genotypes.

genomics↗

Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural variation on gene expression and DNA methylation

Structural variants (SVs) drive gene expression in the human brain and are causative of many neurological conditions. However, most existing genetic studies have been based on short-read sequencing methods, which capture fewer than half of the SVs present in any one individual. Long-read sequencing (LRS) enhances our ability to detect disease-associated and functionally relevant structural variants; however, its application in large-scale genomic studies has been limited by challenges in sample preparation and high costs. Here, we leverage a new scalable wet-lab protocol and computational pipeline for whole-genome Oxford Nanopore Technologies sequencing and apply it to neurologically normal control samples from the North American Brain Expression Consortium (NABEC) (European ancestry) and Human Brain Collection Core (HBCC) (African or African admixed ancestry) cohorts. Through this work, we present a publicly available long-read resource from 351 human brain samples (median N50: 27 Kbp and at an average depth of ~40x genome coverage). We discover approximately 234,905 SVs and produce locally phased assemblies that cover 95% of all protein-coding genes in GRCh38. To resolve cis-regulatory effects, we develop ASM-LR, a method for allele-specific methylation analysis from long-read data, revealing both strong and subtle regulatory effects, including numerous novel methylation QTLs masked in unphased models. Our results highlight the power of haplotype-resolved methylation to uncover regulatory mechanisms and establish a foundational resource for exploring how genetic variation shapes gene expression and epigenetic architecture across diverse ancestries.

genomics↗