bioRxiv Science⌕ Search

Biology subjects

Gunturkun, M. H.

Publications and source records attributed to Gunturkun, M. H..

4 recordsLinked to original sources

Private and sub-family specific mutations of founder haplotypes in the BXD family reveal phenotypic consequences relevant to health and disease

The BXD family of recombinant inbred mice were developed by crossing and inbreeding progeny of C57BL/6J and DBA/2J strains. This family is the largest and most extensively phenotyped mammalian experimental genetic resource. Although used in genetics for 52 years, we do not yet have comprehensive data on DNA variants segregating in the BXDs. Using linked-read whole-genome sequencing, we sequenced 152 members of the family at about 40X coverage and quantified most variants. We identified 6.25 million polymorphism segregating at a near-optimal minor allele frequency of 0.42. We also defined two other major variants: strain-specific de novo singleton mutations and epoch-specific de novo polymorphism shared among subfamilies of BXDs. We quantified per-generation mutation rates of de novo variants and demonstrate how founder-derived, strain-specific, and epoch-specific variants can be analyzed jointly to model genome-phenome causality. This integration enables forward and reverse genetics at scale, rapid production of any of more than 10,000 diallel F1 hybrid progeny to test predictions across diverse environments or treatments. Combined with five decades of phenome data, the BXD family and F1 hybrids are a major resource for systems genetics and experimental precision medicine.

genomics↗

SVJAM: Joint Analysis of Structural Variants Using Linked Read Sequencing Data

Linked-read whole genome sequencing methods, such as the 10x Chromium, attach a unique molecular barcode to each high molecular weight DNA molecule. The samples are then sequenced using short-read technology. During analysis, sequence reads sharing the same barcode are aligned to adjacent genomic locations. The pattern of barcode sharing between genomic regions allows the discovery of large structural variants (SVs) in the range of 1 Kb to a few Mb. Most SV calling methods for these data, such as LongRanger, analyze one sample at a time and often produces inconsistent results for the same genomic location across multiple samples. We developed a method, SVJAM, for joint calling of SVs, using data from 152 members of the BXD family of recombinant inbred strains of mice. Our method first collects candidate SV regions from single sample analysis, such as those produced by LongRanger. We then retrieve barcode overlapping data from all samples for each region. These data are organized as a high dimensional matrix. The dimension of this matrix is then reduced using principal component analysis. Samples projected onto a two dimensional space formed by the first two principal components forms two or three clusters based on their genotype, representing the reference, alternative, or heterozygotic alleles. We developed a novel distance measure for hierarchical clustering and rotating the axes to find the optimal clustering results. We also developed an algorithm to decide whether the pattern of sample distribution is best fitted with one, two, or three genotypes. For each sample, we calculate its membership score for each genotype. We compared results produced by SVJAM with LongRanger and few methods that rely on PacBio or Oxford Nanopore data. In a comparison of SVJAM with SV detected using long-read sequencing data for the DBA/2J strain, we found that our results recovered many SVs missed by LongRanger. We also found many SVs called by LongRanger were assigned with an incorrect SV type. Our algorithm also consistently identified heterozygotic regions.

genomics↗

Genome-wide association study of open fieldbehavior in outbred heterogeneous stock ratsidentifies multiple loci implicated inpsychiatric disorders

Many personality traits are influenced by genetic factors. Rodents models provide an efficient system for analyzing genetic contribution to these traits. Using 1,246 adolescent heterogeneous stock (HS) male and female rats, we conducted a genome-wide association study (GWAS) of behaviors measured in an open field, including locomotion, novel object interaction, and social interaction. We identified 30 genome-wide significant quantitative trait loci (QTL). Using multiple criteria, including the presence of high impact genomic variants and co-localization of cis-eQTL, we identified 13 candidate genes (Adarb2, Ankrd26, Cacna1c, Clock, Crhr1, Ctu2, Cyp26b1, Eva1a, Fam114a1, Kcnj9, Mlf2, Rab27b, Sec11a) for these traits. Most of these genes have been implicated by human GWAS of various psychiatric traits. For example, Cacna1c, a gene known to be critical for social behavior in rodents and implicated in human schizophrenia and bipolar disorder, is a candidate gene for distance to the social zone. In addition, the QTL region for total distance to the novel object zone, on Chr1 at 144 Mb, is syntenic to a hotspot on human Chr15 (82.5-90.8 Mb) that contains 14 genes associated with psychiatric or substance abuse traits. Although some of the genes identified by this study appear to replicate findings from prior human GWAS, others likely represent novel findings that can be the catalyst for future molecular and genetic insights into human psychiatric diseases. Together, these findings provide strong support for the use of the HS population to study psychiatric disorders.

genetics↗

RatsPub: a webservice aided by deep learning to mine PubMed for addiction-related genes

Interpreting and integrating results from omics studies typically requires a comprehensive and time consuming survey of extant literature. Here, we introduce GeneCup, an easy to use literature mining web service that searches all PubMed abstracts for user-provided gene symbols in conjunction with a set of custom keywords organized into a customized ontology, as well as results from human genome-wide association studies (GWAS). As an example, we organized over 300 keywords related to drug addiction into seven categories. The literature search is conducted by querying the NIH PubMed server using a programming interface, which is followed by retrieving abstracts from a local copy of the PubMed archive. The main results presented to the user are individual sentences containing the gene symbol, organized by the keywords they also contain. These sentences are presented through an interactive graphical interface or as tables. GWAS results are displayed using a similar method. All results are linked to the original abstract in PubMed. In addition, a convolutional neural network is employed to distinguish sentences describing systemic stress from those describing cellular stress. The automated and comprehensive search strategy provided by GeneCup facilitates the integration of new discoveries from omic studies with existing literature. GeneCup is free and open source software. The source code of GeneCup and the link to a running instance is available at https://github.com/hakangunturkun/GeneCup

bioinformatics↗