bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

Genetic Diversity Patterns and Domestication Origin of Soybean

Understanding diversity and evolution of a crop is an essential step to implement a strategy to expand its germplasm base for crop improvement research. Samples intensively collected from Korea, which is a small but central region in the distribution geography of soybean, were genotyped to provide sufficient data to underpin genome-wide population genetic questions. After removing natural hybrids and duplicated or redundant accessions, we obtained a non-redundant set comprising 1,957 domesticated and 1,079 wild accessions to perform population structure analyses. Our analysis demonstrates that while wild soybean germplasm will require additional sampling from diverse indigenous areas to expand the germplasm base, the current domesticated soybean germplasm is saturated in terms of genetic diversity. We then showed that our genome-wide polymorphism map enabled us to detect genetic loci underling flower color, seed-coat color, and domestication syndrome. A representative soybean set consisting of 194 accessions were divided into one domesticated subpopulation and four wild subpopulations that could be traced back to their geographic collection areas. Population genomics analyses suggested that the monophyletic group of domesticated soybeans was originated in eastern Japan. The results were further substantiated by a phylogenetic tree constructed from domestication-associated single nucleotide polymorphisms identified in this study.

plant biology

High Genetic Potential for Proteolytic Decomposition in Northern Peatland Ecosystems

AbstractNitrogen (N) is a scarce nutrient commonly limiting primary productivity. Microbial decomposition of complex carbon (C) into small organic molecules (e.g., free amino acids) has been suggested to supplement biologically-fixed N in high latitude peatlands. We evaluated the microbial (fungal, bacterial, and archaeal) genetic potential for organic N depolymerization in peatlands at Marcell Experimental Forest (MEF) in northern Minnesota. We used guided gene assembly to examine the abundance and diversity of protease genes; and further compared to those of N-fixing (nifH) genes in shotgun metagenomic data collected across depth at two distinct peatland environments (bogs and fens). Microbial proteases greatly outnumbered nifH genes with the most abundant gene families (archaeal M1 and bacterial Trypsin) each containing more sequences than all sequences attributed to nifH. Bacterial protease gene assemblies were diverse and abundant across depth profiles, indicating a role for bacteria in releasing free amino acids from peptides through depolymerization of older organic material and contrasting the paradigm of fungal dominance in depolymerization in forest soils. Although protease gene assemblies for fungi were much less abundant overall than for bacteria, fungi were prevalent in surface samples and therefore may be vital in degrading large soil polymers from fresh plant inputs during early stage of depolymerization. In total, we demonstrate that depolymerization enzymes from a diverse suite of microorganisms, including understudied bacterial and archaeal lineages, are likely to play a substantial role in C and N cycling within northern peatlands.\n\nImportanceNitrogen (N) is a common limitation on primary productivity, and its source remains unresolved in northern peatlands that are vulnerable to environmental change. Decompositionof complex organic matter into free amino acids has been proposed as an important N source, but the genetic potential of microorganisms mediating this process has not been examined. Such information can elucidate possible responses of northern peatlands to environmental change. We show high genetic potential for microbial production of free amino acids across a range of microbial guilds. In particular, the abundance and diversity of bacterial genes encoding proteolytic activity suggests a predominant role for bacteria in regulating productivity and contrasts a paradigm of fungal dominance of organic N decomposition. Our results expand our current understanding of coupled carbon and nitrogen cycles in north peatlands and indicate that understudied bacterial and archaeal lineages may be central in this ecosystems response to environmental change.

ecology

Building comprehensive MS-friendly databases for proteomic analysis of bacterial species of unknown genetic background

In proteomics, peptide information within mass spectrometry data from a specific organism sample is routinely challenged against a protein sequence database that best represent such organism. However, if the species/strain in the sample is unknown or poorly genetically characterized, it becomes challenging to determine a database which can represent such sample. Building customized protein sequence databases merging multiple strains for a given species has become a strategy to overcome such restrictions. However, as more genetic information is publicly available and interesting genetic features such as the existence of pan- and core genes within a species are revealed, we questioned how efficient such merging strategies are to report relevant information. To test this assumption, we constructed databases containing conserved and unique sequences for ten different species. Features that are relevant for probabilistic-based protein identification by proteomics were then monitored. As expected, increase in database complexity correlates with pangenomic complexity. However, Mycobacterium tuberculosis and Bortedella pertusis generated very complex databases even having low pangenomic complexity or no pangenome at all. This suggests that discrepancies in gene annotation is higher than average between strains of those species. We further tested database performance by using mass spectrometry data from eight clinical strains from Mycobacterium tuberculosis, and from two published datasets from Staphylococcus aureus. We show that by using an approach where database size is controlled by removing repeated identical tryptic sequences across strains/species, computational time can be reduced drastically as database complexity increases.

bioinformatics

An automated, high-throughput image analysis pipeline enables genetic studies of shoot and root morphology in carrot (Daucus carota L.)

Carrot is a globally important crop, yet efficient and accurate methods for quantifying its most important agronomic traits are lacking. To address this problem, we developed an automated analysis platform that extracts components of size and shape for carrot shoots and roots, which are necessary to advance carrot breeding and genetics. This method reliably measured variation in shoot size and shape, leaf number, petiole length, and petiole width as evidenced by high correlations with hundreds of manual measurements. Similarly, root length and biomass were accurately measured from the images. This platform quantified shoot and root shapes in terms of principal components, which do not have traditional, manually-measurable equivalents. We applied the pipeline in a study of a six-parent diallel population and an F2 mapping population consisting of 316 individuals. We found high levels of repeatability within a growing environment, with low to moderate repeatability across environments. We also observed co-localization of quantitative trait loci for shoot and root characteristics on chromosomes 1, 2, and 7, suggesting these traits are controlled by genetic linkage and/or pleiotropy. By increasing the number of individuals and phenotypes that can be reliably quantified, the development of a high-throughput image analysis pipeline to measure carrot shoot and root morphology will expand the scope and scale of breeding and genetic studies.

plant biology

Transforming summary statistics from logistic regression to the liability scale: application to genetic and environmental risk scores

1. Abstract1.1. ObjectiveStratified medicine requires models of disease risk incorporating genetic and environmental factors. These may combine estimates from different studies and models must be easily updatable when new estimates become available. The logit scale is often used in genetic and environmental association studies however the liability scale is used for polygenic risk scores and measures of heritability, but combining parameters across studies requires a common scale for the estimates.\n\n1.2. MethodsWe present equations to approximate the relationship between univariate effect size estimates on the logit scale and the liability scale, allowing model parameters to be translated between scales.\n\n1.3. ResultsThese equations are used to build a risk score on the liability scale, using effect size estimates originally estimated on the logit scale. Such a score can then be used in a joint effects model to estimate the risk of disease, and this is demonstrated for schizophrenia using a polygenic risk score and environmental risk factors.\n\n1.4. ConclusionThis straightforward method allows conversion of model parameters between the logit and liability scales, and may be a key tool to integrate risk estimates into a comprehensive risk model, particularly for joint models with environmental and genetic risk factors.

epidemiology

Human pancreatic islet 3D chromatin architecture provides insights into the genetics of type 2 diabetes

Genetic studies promise to provide insight into the molecular mechanisms underlying type 2 diabetes (T2D). Variants associated with T2D are often located in tissue-specific enhancer regions (enhancer clusters, stretch enhancers or super-enhancers). So far, such domains have been defined through clustering of enhancers in linear genome maps rather than in 3D-space. Furthermore, their target genes are generally unknown. We have now created promoter capture Hi-C maps in human pancreatic islets. This linked diabetes-associated enhancers with their target genes, often located hundreds of kilobases away. It further revealed sets of islet enhancers, super-enhancers and active promoters that form 3D higher-order hubs, some of which show coordinated glucose-dependent activity. Hub genetic variants impact the heritability of insulin secretion, and help identify individuals in whom genetic variation of islet function is important for T2D. Human islet 3D chromatin architecture thus provides a framework for interpretation of T2D GWAS signals.

genomics

Genetic structure of invasive babys breath (Gypsophila paniculata) populations in a freshwater Michigan dune system

Coastal sand dunes are dynamic ecosystems with elevated levels of disturbance, and as such they are highly susceptible to plant invasions. One such invasion that is of major concern to the Great Lakes dune systems is that of perennial babys breath (Gypsophila paniculata). The invasion of babys breath negatively impacts native species such as the federal threatened Pitchers thistle (Cirsium pitcheri) that occupy the open sand habitat of the Michigan dune system. Our research goals were to (1) quantify the genetic diversity of invasive babys breath populations in the Michigan dune system, and (2) estimate the genetic structure of these invasive populations. We analyzed 12 populations at 14 nuclear and 2 chloroplast microsatellite loci. We found strong genetic structure among populations of babys breath sampled along Michigans dunes (global FST = 0.228), and also among two geographic regions that are separated by the Leelanau peninsula. Pairwise comparisons using the nSSR data among all 12 populations yielded significant FST values. Results from a Bayesian clustering analysis suggest two main population clusters. Isolation by distance was found over all 12 populations (R = 0.755, P < 0.001) and when only cluster 2 populations were included (R = 0.523, P = 0.030); populations within cluster 1 revealed no significant relationship (R = 0.205, P = 0.494). Private nSSR alleles and cpSSR haplotypes within each cluster suggest the possibility of at least two separate introduction events to Michigan.

plant biology

GADMA: Genetic Algorithm for Automatic Inferring Joint Demographic History of Multiple Populations from Allele Frequency Spectrum

The demographic history of any population is imprinted in the genomes of the individuals that make up the population. One of the most popular and convenient representations of genetic information is the allele frequency spectrum or AFS, the distribution of allele frequencies in populations. The joint allele frequency spectrum is commonly used to reconstruct the demographic history of multiple populations and several methods based on diffusion approximation (e.g.,{partial} a{partial}i) and ordinary differential equations (e.g., moments) have been developed and applied for demographic inference. These methods provide an opportunity to simulate AFS under a variety of researcher-specified demographic models and to estimate the best model and associated parameters using likelihood-based local optimizations. However, there are no known algorithms to perform global searches of demographic models with a given AFS. Here, we introduce a new method that implements a global search using a genetic algorithm for the automatic and unsupervised inference of demographic history from joint allele frequency spectrum data. Our method is implemented in the software GADMA (Genetic Algorithm for Demographic Analysis, https://github.com/ctlab/GADMA). We demonstrate the performance of GADMA by applying it to sequence data from humans and non-model organisms and show that it is able to automatically infer a demographic model close to or even better than the one that was previously obtained manually. Moreover, GADMA is able to infer demographic models at different local optima close to the global one, making it is possible to detect more biology corrected model during further research.

evolutionary biology

Genetic correlation between sea age at maturity and iteroparity in Atlantic salmon.

Genetic correlations in life history traits may result in unpredictable evolutionary trajectories if not accounted for in life-history models. Iteroparity (the reproductive strategy of reproducing more than once) in Atlantic salmon (Salmo salar) is a fitness trait with substantial variation within and among populations. In the Teno River in northern Europe, iteroparous individuals constitute an important component of many populations and have experienced a sharp increase in abundance in the last 20 years, partly overlapping with a general decrease in age structure. The physiological basis of iteroparity bears similarities to that of age at first maturity, another life history trait with substantial fitness effects in salmon. Sea age at maturity in Atlantic salmon is controlled by a major locus around the vgll3 gene, and we used this opportunity demonstrate that the two traits are genetically correlated around this genome region. The odds ratio of survival until second reproduction was up to 2.4 (1.8-3.5 90% CI) times higher for fish with the early-maturing vgll3 genotype (EE) compared to fish with the late-maturing genotype (LL). The association had a dominance architecture, although the dominant allele was reversed in the late-maturing group compared to younger groups that stayed only one year at sea before maturation. Post hoc analysis indicated that iteroparous fish with the EE genotype had accelerated growth prior to first reproduction compared to first-time spawners, across all age groups, while this effect was not detected in fish with the LL genotype. These results broaden the functional link around the vgll3 genome region and help us understand constraints in the evolution of life history variation in salmon. Our results further highlight the need to account for genetic correlations between fitness traits when predicting demographic changes in changing environments.

evolutionary biology

Assessment Of Genetic Structure Of The Endangered Forest Species Boswellia Serrata Roxb. Population In Central India

Boswellia serrata Roxb., a commercially important species for its pulp and pharmaceutical properties was sampled from three locations representing its natural distribution in central India for genetic characterization through 56 RAPD + 42 ISSR loci. The wood fiber dimensions measured for morphometric characterization confirmed 11.36% of the variation in the length and 8.75% of the variation in the width indicating its fitness for local adaptation. Bayesian and non-Bayesian approach based diversity measures resulted moderate within population gene diversity (0.26{+/-}0.17), Shannons information index (0.40{+/-}0.22) and panmictic heterozygosity (0.28{+/-}0.01). A high estimate for genetic differentiation measures i.e. GST (0.31), GST-B (0.33{+/-}0.02) and {theta}-II (0.45) led to the distinct clusters of the sampled genotypes representing their regional variability due to limited gene flow and total absence of natural regeneration. We report the first investigation of the species for its molecular characterization emphasizing the urgent need for the genetic improvement program for the In-situ/Ex-situ conservation and sustainable commercialization.

plant biology

Repeated phenotypic evolution by different genetic routes: the evolution of colony switching in Pseudomonas fluorescens SBW25

Repeated evolution of functionally similar phenotypes is observed throughout the tree of life. The extent to which the underlying genetics are conserved remains an area of considerable interest. Previously, we reported the evolution of colony switching in two independent lineages of Pseudomonas fluorescens SBW25 (Beaumont et al., 2009). The phenotypic and genotypic bases of colony switching in the first lineage (Line 1) have been described elsewhere (Beaumont et al., 2009; Gallie et al., 2015). Here, we deconstruct the evolution of colony switching in the second lineage (Line 6). We show that, as for Line 1, Line 6 colony switching results from an increase in the expression of a colanic acid-like polymer (CAP). At the genetic level, nine mutations occur in Line 6. Only one of these - a non-synonymous point mutation in the housekeeping sigma factor rpoD - is required for colony switching. In contrast, the genetic basis of colony switching in Line 1 is a mutation in the metabolic gene carB (Beaumont et al., 2009). A molecular model has recently been proposed whereby the carB mutation increases capsulation by redressing the intracellular balance of positive (ribosomes) and negative (RsmAE/CsrA) regulators of a positive feedback loop in capsule expression (Remigi et al., 2018). We show that Line 6 colony switching is consistent with this model; the rpoD mutation generates an increase in ribosome expression, and ultimately an increase in CAP expression.

evolutionary biology

A phylogenetic framework of the legume genus Aeschynomene for comparative genetic analysis of the Nod-dependent and Nod-independent symbioses

SUMMARYO_LISome Aeschynomene legume species have the property of being nodulated by photosynthetic Bradyrhizobium lacking the nodABC genes. Knowledge of this unique Nod (factor)-independent symbiosis has been gained from the model A. evenia but our understanding remains limited due to the lack of comparative genetics with related taxa using a Nod-dependent process.\nC_LIO_LITo fill this gap, this study significantly broadened previous taxon sampling, including in allied genera, to construct a comprehensive phylogeny. This backbone tree was matched with data on chromosome number, genome size, low-copy nuclear genes and strengthened by nodulation tests and a comparison of the diploid species.\nC_LIO_LIThe phylogeny delineated five main lineages that all contained diploid species while polyploid groups were clustered in a polytomy and were found to originate from a single paleo-allopolyploid event. In addition, new nodulation behaviours were revealed and Nod-dependent diploid species were shown to be tractable.\nC_LIO_LIThe extended knowledge of the genetics and biology of the different lineages in the legume genus Aeschynomene provides a solid research framework. Notably, it enabled the identification of A. americana and A. patula as the most suitable species to undertake a comparative genetic study of the Nod-independent and Nod-dependent symbioses.\nC_LI

plant biology

An assessment of the interactions between climatic conditions and genetic characteristic on the agricultural performance of soybeans grown in Northeast Asia

Glycine max, commonly known as soybean or soya bean, is a species of legume native to East Asia. The interactions between climatic conditions and genetic characteristic affect the agricultural performance of soybean. Therefore, an investigation to identify the main elements affecting the agricultural performances of 11 soybeans was conducted in Northeast Asia, China [Harbin (45{degrees}12'N) Yanji (42{degrees}53'N) Dalian (39{degrees}30'N) Qingdao (36{degrees}26'N)] Republic of Korea [Suwon (37{degrees}16'N) and Jeonju (35{degrees}49'N)]. The days to flowering (DTF) of soybeans with the e1-nf and e1-as alleles and the E1e2e3e4 genotype, except Keumgangkong, Tawonkong, and Duyoukong, was relatively short compared to soybeans with other alleles. Although DTF of the soybeans was highly correlated to all climatic conditions, days to maturity (DTM) and 100-seed weight (HSW) of the soybeans showed no significant correlation with any climatic conditions. The soybeans with a dominant Dt1 allele, except Tawonkong, had the longest stem length (STL). Moreover, the STL of the soybeans grown at the test fields showed a positive correlation with only day length (DL) although the results of our chamber test showed that STL of soybean was positively affected by average temperature (AVT) and DL. Soybean yield (YLD) showed positive correlations with latitude and DL (except L62-667, OT89-5, and OT89-6) although the response of YLD to the climatic conditions was cultivar-specific. Our results show that DTF and STL of soybeans grown in Northeast Asia are highly affected by DL although AVT and genetic characteristic also affect DTF and STL. Along with these results, we confirmed that the DTM, HSW, and YLD of the soybeans vary in relation to their genetic characteristic.

plant biology

Genetic signatures of human cytomegalovirus variants acquired by seronegative glycoprotein B vaccinees

Human cytomegalovirus (HCMV) is the most common congenital infection worldwide, and a frequent cause of hearing loss or debilitating neurologic disease in newborn infants. Thus, a vaccine to prevent HCMV-associated congenital disease is a public health priority. One potential strategy is vaccination of women of child-bearing age to prevent maternal HCMV acquisition during pregnancy. The glycoprotein B (gB) + MF59 adjuvant subunit vaccine is the most efficacious tested clinically to date, demonstrating approximately 50% protection against HCMV infection of seronegative women in multiple phase 2 trials. Yet, the impact of gB/MF59-elicited immune responses on the population of viruses acquired by trial participants has not been assessed. In this analysis, we employed quantitative PCR as well as multiple sequencing methodologies to interrogate the magnitude and genetic composition of HCMV populations infecting gB/MF59 vaccinees and placebo recipients. We identified several differences between the viral dynamics of acutely-infected vaccinees and placebo recipients. First, there was reduced magnitude viral shedding in the saliva of gB vaccinees. Additionally, employing a panel of tests for genetic compartmentalization, we noted tissue-specific gB haplotypes in the majority of vaccinees though only in a single placebo recipient. Finally, we observed reduced acquisition of genetically-related gB1, gB2, and gB4 genotype \"supergroup\" HCMV variants among vaccine recipients, suggesting that the gB1 genotype vaccine construct may have elicited partial protection against HCMV viruses with antigenically-similar gB sequences. These findings indicate that gB immunization may have had a measurable impact on viral intrahost population dynamics and support future analysis of a larger cohort.\n\nAuthor SummaryThough not a household name like Zika virus, human cytomegalovirus (HCMV) causes permanent neurologic disability in one newborn child every hour in the United States - more than Down syndrome, fetal alcohol syndrome, and neural tube defects combined. There are currently no established effective preventative measures to inhibit congenital HCMV transmission following acute or chronic HCMV infection of a pregnant mother. However, the glycoprotein B (gB) vaccine is the most effective HCMV vaccine tried clinically to date. Here, we utilized high-throughput, next-generation sequencing of viral DNA isolated from patients enrolled in a gB vaccine trial, and identified several impacts that this vaccine had on the size, distribution, and composition of the in vivo viral population. These results have increased our understanding of why the gB/MF59 vaccine was partially efficacious and will inform future rational design of a vaccine to prevent congenital HCMV.

microbiology

Genetic analysis of the Komagataella phaffii centromeres by a color-based plasmid stability assay

The yeast Komagataella phaffii is widely used as a microbial host for heterologous protein production. However, molecular tools for this yeast are basically restricted to a few integrative and replicative plasmids. Four sequences that have recently been proposed as the K. phaffii centromeres could be used to develop a new class of mitotically stable vectors. In this work we designed a color-based genetic assay to investigate genetic stability in K. phaffii. Plasmids bearing K. phaffii centromeres and the ADE3 marker were evaluated in terms of mitotic stability in an ade2/ade3 auxotrophic strain which allows plasmid screening through colony color. Plasmid copy number was verified through qPCR. Our results confirmed that the centromeric plasmids were maintained at low copy number as a result of typical chromosome-like segregation during cell division. These features, combined with high transformation efficiency and in vivo assembly possibilities, prompt these plasmids as a new addition to the K. phaffii genetic toolbox.

molecular biology

One Health genomic surveillance of Escherichia coli demonstrates distinct lineages and mobile genetic elements in isolates from humans versus livestock

Livestock have been proposed as a reservoir for drug-resistant Escherichia coli that infect humans. We isolated and sequenced 431 E. coli (including 155 ESBL-producing isolates) from cross-sectional surveys of livestock farms and retail meat in the East of England. These were compared with the genomes of 1517 E. coli associated with bloodstream infection in the United Kingdom. Phylogenetic core genome comparisons demonstrated that livestock and patient isolates were genetically distinct, indicating that E. coli causing serious human infection do not directly originate from livestock. By contrast, we observed highly related isolates from the same animal species on different farms. Analysis of accessory (variable) genomes identified a virulence cassette associated previously with cystitis and neonatal meningitis that was only present in isolates from humans. Screening all 1948 isolates for accessory genes encoding antibiotic resistance revealed 41 different genes present in variable proportions of humans and livestock isolates. We identified a low prevalence of shared antimicrobial resistance genes between livestock and humans based on analysis of mobile genetic elements and long-read sequencing. We conclude that in this setting, there was limited evidence to support the suggestion that antimicrobial resistant pathogens that cause serious infection in humans originate from livestock.\n\nImportanceThe increasing prevalence of E. coli bloodstream infections is a serious public health problem. We used genomic epidemiology in a One Health study conducted in the East of England to examine putative sources of E. coli associated with serious human disease. E. coli from 1517 patients with bloodstream infection were compared with 431 isolates from livestock farms and meat. Livestock-associated and bloodstream isolates were genetically distinct populations based on core genome and accessory genome analyses. Identical antimicrobial resistance genes were found in livestock and human isolates, but there was little overlap in the mobile elements carrying these genes. In addition, a virulence cassette found in humans isolates was not identified in any livestock-associated isolate. Our findings do not support the idea that E. coli causing invasive disease or their resistance genes are commonly acquired from livestock.

genomics

Genetic variability in response to Aβ deposition influences Alzheimer’s risk

Genetic analysis of late-onset Alzheimers disease risk has previously identified a network of largely microglial genes that form a transcriptional network. In transgenic mouse models of amyloid deposition we have previously shown that the expression of many of the mouse orthologs of these genes are co-ordinately up-regulated by amyloid deposition. Here we investigate whether systematic analysis of other members of this mouse amyloid-responsive network predicts other Alzheimers risk loci. This statistical comparison of the mouse amyloid-response network with Alzheimers disease genome-wide association studies identifies 5 other genetic risk loci for the disease (OAS1, CXCL10, LAPTM5, ITGAM and LILRB4). This work suggests that genetic variability in the microglial response to amyloid deposition is a major determinant for Alzheimers risk.\n\nOne Sentence SummaryIdentification of 5 new risk loci for Alzheimers by statistical comparison of mouse A{beta} microglial response with gene-based SNPs from human GWAS

neuroscience

Evidence for bias of genetic ancestry in resting state functional MRI

Resting state functional magnetic resonance imaging (rs-fMRI) is a popular imaging modality for mapping the functional connectivity of the brain. Rs-fMRI is, just like other neuroimaging modalities, subject to a series of technical and subject level biases that change the inferred connectivity pattern. In this work we predicted genetic ancestry from rs-fMRI connectivity data at very high performance (area under the ROC curve of 0.93). Thereby, we demonstrated that genetic ancestry is encoded in the functional connectivity pattern of the brain at rest. Consequently, genetic ancestry constitutes a bias that should be accounted for in the analysis of rs-fMRI data.

neuroscience