bioRxiv Science⌕ Search

Biology subjects

Khan, A. T.

Publications and source records attributed to Khan, A. T..

2 recordsLinked to original sources

Full-length isoform sequencing for resolving the molecular basis of Charcot-Marie-Tooth 2A

ObjectivesTranscript sequencing of patient derived samples has been shown to improve the diagnostic yield for solving cases of likely Mendelian disorders, yet the added benefit of full-length long-read transcript sequencing is largely unexplored. MethodsWe applied short-read and full-length isoform cDNA sequencing and mitochondrial functional studies to a patient-derived fibroblast cell line from an individual with neuropathy that previously lacked a molecular diagnosis. ResultsWe identified an intronic homozygous MFN2 c.600-31T>G variant that disrupts a branch point critical for intron 6 spicing. Full-length long-read isoform cDNA sequencing after treatment with a nonsense-mediated mRNA decay (NMD) inhibitor revealed that this variant creates five distinct altered splicing transcripts. All five altered splicing transcripts have disrupted open reading frames and are subject to NMD. Furthermore, a patient-derived fibroblast line demonstrated abnormal lipid droplet formation, consistent with MFN2 dysfunction. Although correctly spliced full-length MFN2 transcripts are still produced, this branch point variant results in deficient MFN2 protein levels and autosomal recessive Charcot-Marie-Tooth disease, axonal, type 2A (CMT2A). DiscussionThis case highlights the utility of full-length isoform sequencing for characterizing the molecular mechanism of undiagnosed rare diseases and expands our understanding of the genetic basis for CMT2A.

genetics↗

A system for phenotype harmonization in the NHLBI Trans-Omics for Precision Medicine (TOPMed) Program

Genotype-phenotype association studies often combine phenotype data from multiple studies to increase power. Harmonization of the data usually requires substantial effort due to heterogeneity in phenotype definitions, study design, data collection procedures, and data set organization. Here we describe a centralized system for phenotype harmonization that includes input from phenotype domain and study experts, quality control, documentation, reproducible results, and data sharing mechanisms. This system was developed for the National Heart, Lung and Blood Institutes Trans-Omics for Precision Medicine (TOPMed) program, which is generating genomic and other omics data for >80 studies with extensive phenotype data. To date, 63 phenotypes have been harmonized across thousands of participants from up to 17 TOPMed studies per phenotype. We discuss the challenges faced in this undertaking and how they were addressed. The harmonized phenotype data and associated documentation have been submitted to National Institutes of Health data repositories for controlled-access by the scientific community. We also provide materials to facilitate future harmonization efforts by the community, which include (1) the code used to generate the 63 harmonized phenotypes, enabling others to reproduce, modify or extend these harmonizations to additional studies; and (2) results of labeling thousands of phenotype variables with controlled vocabulary terms.

genetics↗