bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Predicting functional neuroanatomical maps from fusing brain networks with genetic information

A central aim, from basic neuroscience to psychiatry, is to resolve how genes control brain circuitry and behavior. This is experimentally hard, since most brain functions and behaviors are controlled by multiple genes. In low throughput, one gene at a time, experiments, it is therefore difficult to delineate the neural circuitry through which these sets of genes express their behavioral effects. The increasing amount of publicly available brain and genetic data offers a rich source that could be mined to address this problem computationally. However, most computational approaches are not tailored to reflect functional synergies in brain circuitry accumulating within sets of genes. Here, we developed an algorithm that fuses gene expression and connectivity data with functional genetic meta data and exploits such cumulative effects to predict neuroanatomical maps for multigenic functions. These maps recapture known functional anatomical annotations from literature and functional MRI data. When applied to meta data from mouse QTLs and human neuropsychiatric databases, our method predicts functional maps underlying behavioral or psychiatric traits. We show that it is possible to predict functional neuroanatomy from mouse and human genetic meta data and provide a discovery tool for high throughput functional exploration of brain anatomy in silico.

Neuroscience

Population genetic history and polygenic risk biases in 1000 Genomes populations

The vast majority of genome-wide association studies are performed in Europeans, and their transferability to other populations is dependent on many factors (e.g. linkage disequilibrium, allele frequencies, genetic architecture). As medical genomics studies become increasingly large and diverse, gaining insights into population history and consequently the transferability of disease risk measurement is critical. Here, we disentangle recent population history in the widely-used 1000 Genomes Project reference panel, with an emphasis on populations underrepresented in medical studies. To examine the transferability of single-ancestry GWAS, we used published summary statistics to calculate polygenic risk scores for six well-studied traits and diseases. We identified directional inconsistencies in all scores; for example, height is predicted to decrease with genetic distance from Europeans, despite robust anthropological evidence that West Africans are as tall as Europeans on average. To gain deeper quantitative insights into GWAS transferability, we developed a complex trait coalescent-based simulation framework considering effects of polygenicity, causal allele frequency divergence, and heritability. As expected, correlations between true and inferred risk were typically highest in the population from which summary statistics were derived. We demonstrated that scores inferred from European GWAS were biased by genetic drift in other populations even when choosing the same causal variants, and that biases in any direction were possible and unpredictable. This work cautions that summarizing findings from large-scale GWAS may have limited portability to other populations using standard approaches, and highlights the need for generalized risk prediction methods and the inclusion of more diverse individuals in medical genomics.

Genomics

Evolutionary Genetics of Insecticide Resistance and the Effects of Chemical Rotation

Repeated use of the same class of pesticides to control a target pest is a form of artificial selection that leads to pesticide resistance. We studied insecticide resistance and cross-resistance to five commercial insecticides in each of six populations of the red flour beetle, Tribolium castaneum. We estimated the dosage response curves for lethality in each parent population for each insecticide and found an 800-fold difference among populations in resistance to insecticides. As expected, a naive laboratory population was among the most sensitive of populations to most insecticides. We then used inbred lines derived from five of these populations to estimate the heritability (h2) of resistance for each pesticide and the genetic correlation (rG) of resistance among pesticides in each population. These quantitative genetic parameters allow insight into the adaptive potential of populations to further evolve insecticide resistance. Lastly, we use our estimates of the genetic variance and covariance of resistance and stochastic simulations to evaluate the efficacy of \"windowing\" as an insecticide resistance management strategy, where the application of several insecticides is rotated on a periodic basis.

Evolutionary Biology

Stochastic dynamics of genetic broadcasting networks

The complex genetic programs of eukaryotic cells are often regulated by key transcription factors occupying or clearing out of a large number of genomic locations. Orchestrating the residence times of these factors is therefore important for the well organized functioning of a large network. The classic models of genetic switches sidestep this timing issue by assuming the binding of transcription factors to be governed entirely by thermodynamic protein-DNA affinities. Here we show that relying on passive thermodynamics and random release times can lead to a \"time-scale crisis\" of master genes that broadcast their signals to large number of binding sites. We demonstrate that this \"time-scale crisis\" can be resolved by actively regulating residence times through molecular stripping. We illustrate these ideas by studying the stochastic dynamics of the genetic network of the central eukaryotic master regulator NF{kappa}B which broadcasts its signals to many downstream genes that regulate immune response, apoptosis etc.

Systems Biology

Variation in olfactory neuron repertoires is genetically controlled and environmentally modulated

The mouse olfactory sensory neuron (OSN) repertoire is composed of 10 million cells and each expresses one olfactory receptor (OR) gene from a pool of over 1000. Thus, the nose is sub-stratified into more than a thousand OSN subtypes. Here, we employ and validate an RNA-sequencing based method to quantify the abundance of all OSN subtypes in parallel, and investigate the genetic and environmental factors that contribute to neuronal diversity. We find that the OSN subtype distribution is stereotyped in genetically identical mice, but varies extensively between different strains. Further, we identify cis-acting genetic variation as the greatest component influencing OSN composition and demonstrate independence from OR function. However, we show that olfactory stimulation with particular odorants results in modulation of dozens of OSN subtypes in a subtle but reproducible, specific and time-dependent manner. Together, these mechanisms generate a highly individualized olfactory sensory system by promoting neuronal diversity.

Neuroscience

The genetic basis and fitness consequences of sperm midpiece size in deer mice

An extraordinary array of reproductive traits vary among species, yet the genetic mechanisms that enable divergence, often over short evolutionary timescales, remain elusive. Here we examine two sister-species of Peromyscus mice with divergent mating systems. We find that the promiscuous species produces sperm with longer midpiece than the monogamous species, and midpiece size correlates positively with competitive ability and swimming performance. Using forward genetics, we identify a gene associated with midpiece length: Prkar1a, which encodes the R1 regulatory subunit of PKA. R1 localizes to midpiece in Peromyscus and is differentially expressed in mature sperm of the two species yet is similarly abundant in the testis. We also show that genetic variation at this locus accurately predicts male reproductive success. Our findings suggest that rapid evolution of reproductive traits can occur through cell type-specific changes to ubiquitously expressed genes and have an important effect on fitness.

Evolutionary Biology

An Optimized Approach for Annotation of Large Eukaryotic Genomic Sequences using Genetic Algorithm

Detection of important functional and/or structural elements and identifying their positions in a large eukaryotic genome is an active research area. Gene is an important functional and structural unit of DNA. The computation of gene prediction is essential for detailed genome annotation. In this paper, we propose a new gene prediction technique based on Genetic Algorithm (GA) for determining the optimal positions of exons of a gene in a chromosome or genome. The correct identification of the coding and non-coding regions are difficult and computationally demanding. The proposed genetic-based method, named Gene Prediction with Genetic Algorithm (GPGA), reduces this problem by searching only one exon at a time instead of all exons along with its introns. The advantage of this representation is that it can break the entire gene-finding problem into a number of smaller subspaces and thereby reducing the computational complexity. We tested the performance of the GPGA with some benchmark datasets and compared the results with the well-known and relevant techniques. The comparison shows the better or comparable performance of the proposed method (GPGA). We also used GPGA for annotating the human chromosome 21 (HS21) using cross species comparison with the mouse orthologs.

bioinformatics

Deciphering the genic basis of environmental adaptations by simultaneous forward and reverse genetics in Saccharomyces cerevisiae

The budding yeast Saccharomyces cerevisiae is the best studied eukaryote in molecular and cell biology, but its utility for understanding the genetic basis of natural phenotypic variation is limited by the inefficiency of association mapping owing to strong and complex population structure. To facilitate association mapping, we analyzed 190 high-quality genomes of diverse strains, including 85 newly sequenced ones, to uncover yeasts population structure that varies substantially among genomic regions. We identified 181 yeast genes that are absent from the reference genome and demonstrated their expression and role in important functions such as drug resistance. We then simultaneously measured the growth rates of over 4500 lab strains each deficient of a nonessential gene and 81 natural strains across multiple environments using unique DNA barcode present in each strain. We combined the genome-wide reverse genetic information with genome-wide association analysis to determine potential genomic regions of importance to environmental adaptations, and for a subset experimentally validated their role by reciprocal hemizygosity tests. The resources provided permit efficient and reliable association mapping in yeast and significantly enhances its value as a model for understanding the genetic mechanisms of phenotypic polymorphism and evolution.

genomics

Understanding genetic changes underlying the molybdate resistance and the glutathione production in Saccharomyces cerevisiae wine strains using an evolution-based strategy

In this work we have investigated the genetic changes underlying the high glutathione (GSH) production showed by the evolved Saccharomyces cerevisiae strain UMCC 2581, selected in a molybdate-enriched environment after sexual recombination of the parental wine strain UMCC 855. To reach our goal, we first generated strains with the desired phenotype, and then we mapped changes underlying adaptation to molybdate by using a whole-genome sequencing. Moreover, we carried out the RNA-seq that allowed an accurate measurement of gene expression and an effective comparison between the transcriptional profiles of parental and evolved strains, in order to investigate the relationship between genotype and high GSH production phenotype.\n\nAmong all genes evaluated only two genes, MED2 and RIM15 both related to oxidative stress response, presented new mutations in the UMCC 2581 strain sequence and were potentially related to the evolved phenotype.\n\nRegarding the expression of high GSH production phenotype, it included over-expression of amino acids permeases and precursor biosynthetic enzymes rather than the two GSH metabolic enzymes, whereas GSH production and metabolism, transporter activity, vacuolar detoxification and oxidative stress response enzymes were probably added resulting in the molybdate resistance phenotype. This work provides an example of a combination of an evolution-based strategy to successful obtain yeast strain with desired phenotype and inverse engineering approach to genetic characterize the evolved strain. The obtained genetic information could be useful for further optimization of the evolved strains and for providing an even more rapid approach to identify new strains, with a high GSH production, through a marked-assisted selection strategy.

genomics

Shared activity patterns arising at genetic susceptibility loci reveal underlying genomic and cellular architecture of human disease.

Genetic variants underlying complex traits, including disease susceptibility, are enriched within the transcriptional regulatory elements, promoters and enhancers. There is emerging evidence that regulatory elements associated with particular traits or diseases share patterns of transcriptional regulation. Accordingly, shared transcriptional regulation (coexpression) may help prioritise loci associated with a given trait, and help to identify the biological processes underlying it. Using cap analysis of gene expression (CAGE) profiles of promoter and enhancer-derived RNAs across 1824 human samples, we have quantified coexpression of RNAs originating from trait-associated regulatory regions using a novel analytical method (network density analysis; NDA). For most traits studied, sequence variants in regulatory regions were linked to tightly coexpressed networks that are likely to share important functional characteristics. These networks implicate particular cell types and tissues in disease pathogenesis; for example, variants associated with ulcerative colitis are linked to expression in gut tissue, whereas Crohns disease variants are restricted to immune cells. We show that this coexpression signal provides additional independent information for fine mapping likely causative variants. This approach identifies additional genetic variants associated with specific traits, including an association between the regulation of the OCT1 cation transporter and genetic variants underlying circulating cholesterol levels. This approach enables a deeper biological understanding of the causal basis of complex traits.\n\nONE SENTENCE SUMMARYWe discover that variants associated with a specific disease share expression profiles across tissues and cell types, enabling fine mapping and identification of new disease-associated variants, illuminating key cell types involved in disease pathogenesis.

genomics

Infectious Disease Dynamics Inferred from Genetic Data via Sequential Monte Carlo

Genetic sequences from pathogens can provide information about infectious disease dynamics that may supplement or replace information from other epidemiological observations. Currently available methods first estimate phylogenetic trees from sequence data, then estimate a transmission model conditional on these phylogenies. Outside limited classes of models, existing methods are unable to enforce logical consistency between the model of transmission and that underlying the phylogenetic reconstruction. Such conflicts in assumptions can lead to bias in the resulting inferences. Here, we develop a general, statistically efficient, plug-and-play method to jointly estimate both disease transmission and phylogeny using genetic data and, if desired, other epidemiological observations. This method explicitly connects the model of transmission and the model of phylogeny so as to avoid the aforementioned inconsistency. We demonstrate the feasibility of our approach through simulation and apply it to estimate stage-specific infectiousness in a subepidemic of HIV in Detroit, Michigan. In a supplement, we prove that our approach is a valid sequential Monte Carlo algorithm. While we focus on how these methods may be applied to population-level models of infectious disease, their scope is more general. These methods may be applied in other biological systems where one seeks to infer population dynamics from genetic sequences, and they may also find application for evolutionary models with phenotypic rather than genotypic data.

epidemiology

Preserving microsatellites? Conservation genetics of the giant Galapagos tortoise.

This preprint has been reviewed and recommended by Peer Community In Evolutionary Biology (http://dx.doi.org/10.24072/pci.evolbiol.100031).\n\nConservation policy in the giant Galapagos tortoise, an iconic endangered animal, has been assisted by genetic markers for [~]15 years: a dozen loci have been used to delineate thirteen (sub)species, between which hybridization is prevented. Here, comparative reanalysis of a previously published NGS data set reveals a conflict with traditional markers. Genetic diversity and population substructure in the giant Galapagos tortoise are found to be particularly low, questioning the genetic relevance of current conservation practices. Further examination of giant Galapagos tortoise population genomics is critically needed.

evolutionary biology

Spontaneous mutations and transmission distortions of genic copy number variants shape the standing genetic variation in Picea glauca

Copy number variations (CNVs) are large genetic variations detected among the individuals of every multicellular organism examined so far. These variations are believed to play an important role in the evolution and adaptation of species. In plants, little is known about the characteristics of CNVs, particularly regarding the rates at which they are generated and the mechanics of their transmission from a generation to the next. Here, we used SNP-array raw intensity data for 55 two-generations families (3663 individuals) to scan the gene space of the conifer tree Picea glauca (Moench) Voss for CNVs. We were particularly interested in the abundance, inheritance, spontaneous mutation rate spectrum and the evolutionary consequences they may have on the standing genetic variation of white spruce. Our findings show that CNVs affect a small proportion of the gene space and are predominantly copy number losses. CNVs were either inherited or generated through de novo events. De novo CNVs present high rates of spontaneous mutations that vary for different genes and alleles and are correlated with gene expression levels. Most of the inherited CNVs (70%) are transmitted from the parents in violation of Mendelian expectations. These transmission distortions can cause considerable frequency changes between generations and be dependent on whether the heterozygote parents contribute as male or female. Transmission distortions were also influenced by the partner genotype and the parents genetic background. This study provides new insights into the effects of different evolutionary forces on copy number variations based on the analysis of a perennial plant.

genomics

Genetic equidistance at the nucleotide level

The genetic equidistance phenomenon was first discovered in 1963 by Margoliash and shows complex taxa to be all approximately equidistant to a less complex species in amino acid percentage identity. The result has been mis-interpretated by the ad hoc universal molecular clock hypothesis, and the much overlooked mystery was finally solved by the maximum genetic diversity hypothesis (MGD). Here, we studied 15 proteomes and their coding DNA sequences (CDS) to see if the equidistance phenomenon also holds at the CDS level. We performed DNA alignments for a total of 5 groups with 3 proteomes per group and found that in all cases the outgroup taxon was equidistant to the two more complex taxa species at the DNA level. Also, when two sister taxa (snake and bird) were compared to human as the outgroup, the more complex taxon bird was closer to human, confirming species complexity rather than time to be the primary determinant of MGD. Finally, we found the fraction of overlap sites where coincident substitutions occur to be inversely correlated with CDS conservation, indicating saturation to be more common in less conserved DNAs. These results establish the genetic equidistance phenomenon to be universal at the DNA level and provide additional evidence for the MGD theory.

evolutionary biology

A maternal-effect genetic incompatibility in Caenorhabditis elegans

Selfish genetic elements spread in natural populations and have an important role in genome evolution. We discovered a selfish element causing a genetic incompatibility between strains of the nematode Caenorhabditis elegans. The element is made up of sup-35, a maternal-effect toxin that kills developing embryos, and pha-1, its zygotically expressed antidote. pha-1 has long been considered essential for pharynx development based on its mutant phenotype, but this phenotype in fact arises from a loss of suppression of sup-35 toxicity. Inactive copies of the sup-35/pha-1 element show high sequence divergence from active copies, and phylogenetic reconstruction suggests that they represent ancestral stages in the evolution of the element. Our results suggest that other essential genes identified by genetic screens may turn out to be components of selfish elements.

evolutionary biology

MOSAIC: a chemical-genetic interaction data repository and web resource for exploring chemical modes of action

SummaryChemical-genomic approaches that map interactions between small molecules and genetic perturbations offer a promising strategy for functional annotation of uncharacterized bioactive compounds. We recently developed a new high-throughput platform for mapping chemical-genetic (CG) interactions in yeast that can be scaled to screen large compound collections, and we applied this system to generate CG interaction profiles for more than 13,000 compounds. When integrated with the existing global yeast genetic interaction network, CG interaction profiles can enable mode-of-action prediction for previously uncharacterized compounds as well as discover unexpected secondary effects for known drugs. To facilitate future analysis of these valuable data, we developed a public database and web interface named MOSAIC. The website provides a convenient interface for querying compounds, bioprocesses (GO terms), and genes for CG information including direct CG interactions, bioprocesses, and gene-level target predictions. MOSAIC also provides access to chemical structure information of screened molecules, chemical-genomic profiles, and the ability to search for compounds sharing structural and functional similarity. This resource will be of interest to chemical biologists for discovering new small molecule probes with specific modes-of-action as well as computational biologists interested in analyzing CG interaction networks.\n\nAvailabilityMOSAIC is available at http://mosaic.cs.umn.edu.\n\nContactchadm@umn.edu, charlie.boone@utoronto.ca, yoshidam@riken.jp, or hisyo@riken.jp

bioinformatics

Evolution and clinical impact of genetic epistasis within EGFR-mutant lung cancers

Introductory paragraph Introductory paragraph Main text Methods Author Contributions Author Information References The current understanding of tumorigenesis is largely centered on a monogenic driver oncogene model. This paradigm is incompatible with the prevailing clinical experience in most solid malignancies: monotherapy with a drug directed against an individual oncogenic driver typically results in incomplete clinical responses and eventual tumor progression1-7. By profiling the somatic genetic alterations present in over 2,000 cases of lung cancer, the leading cause of cancer mortality worldwide8,9, we show that combinations of functional genetic alterations, i.e. genetic collectives dominate the landscape of ...

cancer biology

Intra- and Inter-individual genetic variation in human ribosomal RNAs

The ribosome is an ancient RNA-protein complex essential for translating DNA to protein. At its core are the ribosomal RNAs (rRNAs), the most abundant RNA in the cell. To support high levels of transcription, repetitive arrays of ribosomal DNA (rDNA) are necessary and the long-standing hypothesis is that they undergo sequence homogenization towards rDNA uniformity.\n\nHere I present evidence of the rich genetic diversity in human rDNA, both within and between individuals. Using state-of-the-art genome sequencing data revealed an average of 192.7 intra-individual variants, including some deeply penetrating the rDNA copies, such as the bi-allelically expressed 28S.59A>G. From 104 diverse genomes, 947 high-confidence variants were identified and unmask a hidden genetic diversity of humans.\n\nThese findings support the emerging concept that ribosomes are heterogeneous within cells and extends the heterogeneity into the realm of population genetics. Fundamentally, do our ribosomal variants determine how our cells interpret the genome?

genomics