bioRxiv Science⌕ Search

Biology subjects

Daida, K.

Publications and source records attributed to Daida, K..

6 recordsLinked to original sources

Long-read sequencing maps transposable element variation and its regulatory and epigenetic effects in the human brain

Transposable elements (TEs) are mobile DNA sequences that shape genome architecture and gene regulation, yet their roles in the human brain remain largely unresolved. Short-read sequencing lacks the resolution to accurately map TE insertions, detect associated structural variants, and resolve highly repetitive regions. Here, we leverage long-read whole-genome sequencing to profile germline TE insertions in postmortem brain tissue from two ancestrally diverse cohorts: the North American Brain Expression Consortium (NABEC; European ancestry, n = 205) and the Human Brain Collection Core (HBCC; African and African-admixed ancestry, n = 146). We identified 2,842 and 1,660 high-confidence non-reference insertions in HBCC and NABEC, respectively, spanning Alu, LINE-1, and SVA elements. We then also further characterized complex short tandem repeat and variable number tandem repeat variation within reference SVA and Alu loci. Reference TEs were also found to mediate complex structural variants at loci implicated in brain development and neurodegenerative disease, with several showing ancestry-specific patterns. Integration of bulk RNA-sequencing data identified TE expression quantitative trait loci, including insertions that modulate neuronal gene expression. Single-nucleus RNA sequencing revealed cell-type-specific effects of TE regulation across cortical populations. Long-read methylation profiling further demonstrated age-associated epigenetic regulation of both reference and non-reference Alu elements. As a community resource, we release a catalog of TE insertions, allele frequencies, and ancestry-specific distributions to enable future functional and disease-focused investigations. Together, these findings highlight the widespread regulatory and epigenetic influence of TEs in the human brain and establish long-read sequencing as a powerful approach for uncovering cell-type- and population-specific TE dynamics.

genomics↗

The complete genome of the KOLF2.1J reference iPSC line

While induced pluripotent stem cells (iPSCs) have gained popularity in studying neurodegenerative diseases, the heterogeneity of stem cells used across studies impacts cross-study comparison. The iPSC Neurodegenerative Disease Initiative (iNDI) selected the KOLF2.1J cell line and prioritized its use as a reference standard for studying the effects of pathogenic variants on cell biology due to its stability and neutral neurodegenerative disease genetic risk. This cell line, and its derivatives expressing over 100 variants related to Alzheimers disease, Parkinsons disease, and other neurological diseases, are available for academic and industry access. Current genomic data analyses are limited by the use of a human reference genome that does not capture the complete genetic background of a given iPSC line. While in the future this issue may be partially mitigated by the creation of a comprehensive human pangenome, previous work has shown that generating custom genomes is of value both to characterize the variation present and to serve as a more appropriate genomic reference. Here, we generated and characterized a custom complete genome assembly from KOLF2.1J. Mapping of sequencing reads to a personalized diploid assembly results in more comprehensive mapping compared to traditional linear references (i.e GRCh38). In addition, we provide a comprehensive custom gene annotation along with isoform expression and differential methylation analyses across multiple cell types. The assembly and all additional data is browsable and publicly available. This resource will enable more accurate investigation of the KOLF2.1J cell line and any genomics data generated compared to using traditional generalized references, while also serving as a foundational approach for establishing custom reference assemblies for other high-value iPSC lines.

genomics↗

Haplotype-Resolved DNA Methylation at the APOE Locus identifies Allele-Specific Epigenetic Signatures Relevant to Alzheimer's Disease Risk

The APOE gene encodes a key lipid transport protein and plays a central role in Alzheimers disease (AD) pathogenesis. Three common APOE alleles, {varepsilon}2 (rs7412(C>T), {varepsilon}3 (reference), and {varepsilon}4 (rs429358(T>C)), arise from two coding variants in exon 4 and confer distinct AD risk profiles, with {varepsilon}4 increasing risk and {varepsilon}2 providing protection. The {varepsilon}3-linked APOE variant rs769455[T] has also been associated with elevated AD risk in individuals of African ancestry carrying both rs769455[T] and {varepsilon}4 alleles. These single nucleotide variants (SNVs) reside in a cytosine-phosphate-guanine (CpG) island, which is a region with a higher frequency of CpG sites compared to the rest of the genome. CpG sites are subject to 5-methylcytosine (5mC) methylation by DNA methyltransferases which add a methyl group to the fifth carbon on the cytosine residue of a CpG site. The presence of SNVs can disrupt this process, making these regions prime targets for differential methylation; however, allele-specific methylation patterns in APOE remain poorly resolved due to technical limitations of conventional bisulfite and methylation array based methods, including degraded DNA quality, sparse CpG coverage, and lack of haplotype phasing. Here, we leverage high-accuracy long-read sequencing data to generate haplotype-resolved methylation profiles of the APOE locus in 332 postmortem brain samples from two ancestrally different cohorts. This includes 201 individuals of European ancestry from the North American Brain Expression Consortium (NABEC), comprising 402 haplotypes (48 {varepsilon}2 and 58 {varepsilon}4 alleles), and 131 individuals of African and African admixed ancestry from the Human Brain Core Collection (HBCC), comprising 262 haplotypes (25 {varepsilon}2, 64 {varepsilon}4, and 7 rs769455 alleles). A linear regression analysis identified 18 novel differentially methylated CpG sites (DMCs) associated with APOE {varepsilon}2, {varepsilon}4, and rs769455 within a gene cluster spanning TOMM40, APOE, APOC1, and APOC4-APOC2. This represents the most comprehensive haplotype-resolved methylation study of APOE in human brain tissue to date. Our results uncover distinct allele-specific methylation signatures and demonstrate the power of long-read sequencing for resolving epigenetic variation relevant to AD risk.

genomics↗

Long-read sequencing of hundreds of diverse brains provides insight into the impact of structural variation on gene expression and DNA methylation

Structural variants (SVs) drive gene expression in the human brain and are causative of many neurological conditions. However, most existing genetic studies have been based on short-read sequencing methods, which capture fewer than half of the SVs present in any one individual. Long-read sequencing (LRS) enhances our ability to detect disease-associated and functionally relevant structural variants; however, its application in large-scale genomic studies has been limited by challenges in sample preparation and high costs. Here, we leverage a new scalable wet-lab protocol and computational pipeline for whole-genome Oxford Nanopore Technologies sequencing and apply it to neurologically normal control samples from the North American Brain Expression Consortium (NABEC) (European ancestry) and Human Brain Collection Core (HBCC) (African or African admixed ancestry) cohorts. Through this work, we present a publicly available long-read resource from 351 human brain samples (median N50: 27 Kbp and at an average depth of ~40x genome coverage). We discover approximately 234,905 SVs and produce locally phased assemblies that cover 95% of all protein-coding genes in GRCh38. To resolve cis-regulatory effects, we develop ASM-LR, a method for allele-specific methylation analysis from long-read data, revealing both strong and subtle regulatory effects, including numerous novel methylation QTLs masked in unphased models. Our results highlight the power of haplotype-resolved methylation to uncover regulatory mechanisms and establish a foundational resource for exploring how genetic variation shapes gene expression and epigenetic architecture across diverse ancestries.

genomics↗

CNV-Finder: Streamlining Copy Number Variation Discovery

Copy Number Variations (CNVs) play pivotal roles in the etiology of complex diseases and are variable across diverse populations. Understanding the association between CNVs and disease susceptibility is significant in disease genetics research and often requires analysis of large sample sizes. One of the most cost-effective and scalable methods for detecting CNVs is based on normalized signal intensity values, such as Log R Ratio (LRR) and B Allele Frequency (BAF), from Illumina genotyping arrays. In this study, we present CNV-Finder, a novel pipeline integrating deep learning techniques on array data, specifically a Long Short-Term Memory (LSTM) network, to expedite the large-scale identification of CNVs within predefined genomic regions. This facilitates efficient prioritization of samples for time-consuming or costly subsequent analyses such as Multiplex Ligation-dependent Probe Amplification (MLPA), short-read, and long-read whole genome sequencing. We incorporate four genes to establish our methods--Parkin (PRKN), Leucine Rich Repeat And Ig Domain Containing 2 (LINGO2), Microtubule Associated Protein Tau (MAPT), and alpha-Synuclein (SNCA)--which may be relevant to neurological diseases such as Alzheimers disease (AD), Parkinsons disease (PD), Progressive Supranuclear Palsy (PSP), or related disorders such as essential tremor (ET). By training our models on expert-annotated samples and validating them across diverse cohorts, including those from the Global Parkinsons Genetics Program (GP2) and additional dementia-specific databases, we demonstrate the efficacy of CNV-Finder in accurately detecting deletions and duplications. Our pipeline outputs app-compatible files for visualization within CNV-Finders interactive web application. This interface enables researchers to review predictions and filter displayed samples by model prediction values, LRR range, and variant count in order to explore or confirm results. Our pipeline integrates this human feedback to enhance model performance and reduce false positive rates. Through a series of comprehensive analyses and validations using visual inspection, MLPA, short-read, and long-read sequencing data, we demonstrate the robustness and adaptability of CNV-Finder in identifying CNVs with regions of varied size, probe density, and noise. Our findings highlight the significance of contextual understanding and human expertise in enhancing the precision of CNV identification, particularly in complex genomic regions like 17q21.31. The CNV-Finder pipeline is a scalable, publicly available resource for the scientific community, available on GitHub (https://github.com/GP2code/CNV-Finder; DOI 10.5281/zenodo.14182563). CNV-Finder not only expedites accurate candidate identification but also significantly reduces the manual workload for researchers, enabling future targeted validation and downstream analyses in regions or phenotypes of interest.

bioinformatics↗

Scalable Nanopore sequencing of human genomes provides a comprehensive view of haplotype-resolved variation and methylation

Long-read sequencing technologies substantially overcome the limitations of short-reads but to date have not been considered as feasible replacement at scale due to a combination of being too expensive, not scalable enough, or too error-prone. Here, we develop an efficient and scalable wet lab and computational protocol for Oxford Nanopore Technologies (ONT) long-read sequencing that seeks to provide a genuine alternative to short-reads for large-scale genomics projects. We applied our protocol to cell lines and brain tissue samples as part of a pilot project for the NIH Center for Alzheimers and Related Dementias (CARD). Using a single PromethION flow cell, we can detect SNPs with F1-score better than Illumina short-read sequencing. Small indel calling remains difficult within homopolymers and tandem repeats, but is comparable to Illumina calls elsewhere. Further, we can discover structural variants with F1-score comparable to state-of-the-art methods involving Pacific Biosciences HiFi sequencing and trio information (but at a lower cost and greater throughput). Using ONT-based phasing, we can then combine and phase small and structural variants at megabase scales. Our protocol also produces highly accurate, haplotype-specific methylation calls. Overall, this makes large-scale long-read sequencing projects feasible; the protocol is currently being used to sequence thousands of brain-based genomes as a part of the NIH CARD initiative. We provide the protocol and software as open-source integrated pipelines for generating phased variant calls and assemblies.

bioinformatics↗