bioRxiv Science⌕ Search

Biology subjects

Blue, E. E.

Publications and source records attributed to Blue, E. E..

4 recordsLinked to original sources

Synchronized long-read genome, methylome, epigenome, and transcriptome for resolving a Mendelian condition

Resolving the molecular basis of a Mendelian condition (MC) remains challenging owing to the diverse mechanisms by which genetic variants cause disease. To address this, we developed a synchronized long-read genome, methylome, epigenome, and transcriptome sequencing approach, which enables accurate single-nucleotide, insertion-deletion, and structural variant calling and diploid de novo genome assembly, and permits the simultaneous elucidation of haplotype-resolved CpG methylation, chromatin accessibility, and full-length transcript information in a single long-read sequencing run. Application of this approach to an Undiagnosed Diseases Network (UDN) participant with a chromosome X;13 balanced translocation of uncertain significance revealed that this translocation disrupted the functioning of four separate genes (NBEA, PDK3, MAB21L1, and RB1) previously associated with single-gene MCs. Notably, the function of each gene was disrupted via a distinct mechanism that required integration of the four omes to resolve. These included nonsense-mediated decay, fusion transcript formation, enhancer adoption, transcriptional readthrough silencing, and inappropriate X chromosome inactivation of autosomal genes. Overall, this highlights the utility of synchronized long-read multi-omic profiling for mechanistically resolving complex phenotypes.

genetics↗

Full-length isoform sequencing for resolving the molecular basis of Charcot-Marie-Tooth 2A

ObjectivesTranscript sequencing of patient derived samples has been shown to improve the diagnostic yield for solving cases of likely Mendelian disorders, yet the added benefit of full-length long-read transcript sequencing is largely unexplored. MethodsWe applied short-read and full-length isoform cDNA sequencing and mitochondrial functional studies to a patient-derived fibroblast cell line from an individual with neuropathy that previously lacked a molecular diagnosis. ResultsWe identified an intronic homozygous MFN2 c.600-31T>G variant that disrupts a branch point critical for intron 6 spicing. Full-length long-read isoform cDNA sequencing after treatment with a nonsense-mediated mRNA decay (NMD) inhibitor revealed that this variant creates five distinct altered splicing transcripts. All five altered splicing transcripts have disrupted open reading frames and are subject to NMD. Furthermore, a patient-derived fibroblast line demonstrated abnormal lipid droplet formation, consistent with MFN2 dysfunction. Although correctly spliced full-length MFN2 transcripts are still produced, this branch point variant results in deficient MFN2 protein levels and autosomal recessive Charcot-Marie-Tooth disease, axonal, type 2A (CMT2A). DiscussionThis case highlights the utility of full-length isoform sequencing for characterizing the molecular mechanism of undiagnosed rare diseases and expands our understanding of the genetic basis for CMT2A.

genetics↗

FAVOR: Functional Annotation of Variants Online Resource and Annotator for Variation across the Human Genome

Large-scale whole genome sequencing (WGS) studies and biobanks are rapidly generating a multitude of coding and non-coding variants. They provide an unprecedented resource for illuminating the genetic basis of human diseases. Variant functional annotations play a critical role in WGS analysis, result interpretation, and prioritization of disease- or trait-associated causal variants. Existing functional annotation databases have limited scope to perform online queries or are unable to functionally annotate the genotype data of large WGS studies and biobanks for downstream analysis. We develop the Functional Annotation of Variants Online Resources (FAVOR) to meet these pressing needs. FAVOR provides a comprehensive online multi-faceted portal with summarization and visualization of all possible 9 billion single nucleotide variants (SNVs) across the genome, and allows for rapid variant-, gene-, and region-level online queries. It integrates variant functional information from multiple sources to describe the functional characteristics of variants and facilitates prioritizing plausible causal variants influencing human phenotypes. Furthermore, a scalable annotation tool, FAVORannotator, is provided for functionally annotating and efficiently storing the genotype and variant functional annotation data of a large-scale sequencing study in an annotated GDS file format to facilitate downstream analysis. FAVOR and FAVORannotator are available at https://favor.genohub.org.

genetics↗

Leveraging TOPMed Imputation Server and Constructing a Cohort-Specific Imputation Reference Panel to Enhance Genotype Imputation among Cystic Fibrosis Patients

Cystic fibrosis (CF) is a severe genetic disorder that can cause multiple comorbidities affecting the lungs, the pancreas, the luminal digestive system and beyond. In our previous genome-wide association studies (GWAS), we genotyped [~]8,000 CF samples using a mixture of different genotyping platforms. More recently, the Cystic Fibrosis Genome Project (CFGP) performed deep ([~]30x) whole genome sequencing (WGS) of 5,095 samples to better understand the genetic mechanisms underlying clinical heterogeneity among CF patients. For mixtures of GWAS array and WGS data, genotype imputation has proven effective in increasing effective sample size. Therefore, we first performed imputation for the [~]8,000 CF samples with GWAS array genotype using the TOPMed freeze 8 reference panel. Our results demonstrate that TOPMed can provide high-quality imputation for CF patients, boosting genomic coverage from [~]0.3 - 4.2 million genotyped markers to [~]11 - 43 million well-imputed markers, and significantly improving Polygenic Risk Score (PRS) prediction accuracy. Furthermore, we built a CF-specific CFGP reference panel based on WGS data of CF patients. We demonstrate that despite having [~]3% the sample size of TOPMed, our CFGP reference panel can still outperform TOPMed when imputing some CF disease-causing variants, likely due to allele and haplotype differences between CF patients and general populations. We anticipate our imputed data for 4,656 samples without WGS data will benefit our subsequent genetic association studies, and the CFGP reference panel built from CF WGS samples will benefit other investigators studying CF.

genetics↗