bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 955 records · Page 53Linked to original sources

Bone Morphology is Regulated Modularly by Global and Regional Genetic Programs

During skeletogenesis, a variety of protrusions of different shapes and sizes develop on the surfaces of long bones. These superstructures provide stable anchoring sites for ligaments and tendons during the assembly of the musculoskeletal system. Despite their importance, the mechanism by which superstructures are patterned and ultimately give rise to the unique morphology of each long bone is far from understood. In this work, we provide further evidence that long bones form modularly from Sox9+ cells, which contribute to their substructure, and from Sox9+/Scx+ progenitors that give rise to superstructures. Moreover, we identify components of the genetic program that controls the patterning of Sox9+/Scx+ progenitors and show that this program includes both global and regional regulatory modules.\n\nUsing light sheet fluorescence microscopy combined with genetic lineage labeling, we mapped the broad contribution of the Sox9+/Scx+ progenitors to the formation of bone superstructures. Additionally, by combining literature-based evidence and comparative transcriptomic analysis of different Sox9+/Scx+ progenitor populations, we identified genes potentially involved in patterning of bone superstructures. We present evidence indicating that Gli3 is a global regulator of superstructure patterning, whereas Pbx1, Pbx2, Hoxa11 and Hoxd11 act as proximal and distal regulators, respectively. Moreover, by demonstrating a dose-dependent pattern regulation in Gli3 and Pbx1 compound mutations, we show that the global and regional regulatory modules work coordinately. Collectively, our results provide strong evidence for genetic regulation of superstructure patterning that further supports the notion that long bone development is a modular process.

developmental biology

Fast estimation of genetic relatedness between members of heterogeneous populations of closely related genomic variants

Many biological analysis tasks require extraction of families of genetically similar sequences from large datasets produced by Next-generation Sequencing (NGS). Such tasks include detection of viral transmissions by analysis of all genetically close pairs of sequences from viral datasets sampled from infected individuals or studying of evolution of viruses or immune repertoires by analysis of network of intra-host viral variants or antibody clonotypes formed by genetically close sequences. The most obvious na{iota}eve algorithms to extract such sequence families are impractical in light of the massive size of modern NGS datasets. In this paper, we present fast and scalable k-mer-based framework to perform such sequence similarity queries efficiently, which specifically targets data produced by deep sequencing of heterogeneous populations such as viruses. The tool is freely available for download at https://github.com/vyacheslav-tsivina/signature-sj

bioinformatics

crossword: A data-driven simulation language for the design of genetic-mapping experiments and breeding strategies

The simulation of genetic systems can save time and resources by optimizing the logistics of an experiment. Current tools are difficult to use by those unfamiliar with programming, and these tools rarely address the actual genetic structure of the population under study. Here, we introduce crossword, which utilizes the widely available results of re-sequencing and genomics data to create more realistic simulations and to simplify user input. The software was written in R, making installation and implementation straightforward. Because crossword is a domain-specific language, it allows complex and unique simulations to be performed, but the language is supported by a graphical interface that guides users through functions and options. We first show crosswords utility in QTL-seq design, where its output accurately reflects empirical data. By introducing the concept of levels to reflect family relatedness, crossword is suitable to a broad range of breeding programs and crops. Using levels, we further illustrate crosswords capabilities by examining the effect of family size and number of selfing generations on phenotyping accuracy and genomic selection. Additionally, we explore the ramifications of effect polarity among parents in a mapping cross, a scenario that is common in crop genetics but often difficult to simulate. Given the ease of use and apparent realism, we anticipate crossword will quickly become a \"bicycle for the [geneticists] mind\".

bioinformatics

Longitudinal studies at birth and age 7 reveal strong effects of genetic variation on ancestry-associated DNA methylation patterns in blood cells from ethnically admixed children

Epigenetic architecture is influenced by genetic and environmental factors, but little is known about their relative contributions or longitudinal dynamics. Here, we studied DNA methylation (DNAm) at over 750,000 CpG sites in mononuclear blood cells collected at birth and age 7 from 196 children of primarily self-reported Black and Hispanic ethnicities to study race-associated DNAm patterns. We developed a novel Bayesian method for high dimensional longitudinal data and showed that race-associated DNAm patterns at birth and age 7 are nearly identical. Additionally, we estimated that up to 51% of all self-reported race-associated CpGs had race-dependent DNAm levels that were mediated through local genotype and, quite surprisingly, found that genetic factors explained an overwhelming majority of the variation in DNAm levels at other, previously identified, environmentally-associated CpGs. These results not only indicate that race-associated DNAm patterns in blood are present at birth and are primarily genetically, and not environmentally, determined, but also that DNAm in blood cells overall is robust to many environmental exposures during the first 7 years of life.

genomics

The probiotic effectiveness in experimental colitis is correlated with gut microbiome and host genetic features

Current evidence to support extensive use of probiotics in inflammatory bowel disease is limited and factors contribute to the inconsistent effectiveness of clinical probiotic therapy are not completely known. Here, as a proof-of-concept, we utilized Bifidobacterium longum JDM 301, a widely used commercial probiotic strain in China, to study potential factors that may influence the beneficial effect of probiotics in experimental colitis. We found that the probiotic therapeutic effect was varied across individual mouse even with the same genetic background and consuming the same type of food. The different probiotic efficacy was highly correlated with different microbiome features in each mouse. Consumption of a diet rich in fat can change the host sensitivity to mucosal injury-induced colitis but did not change the host responsiveness to probiotic therapy. Finally, the host genetic factor TLR2 was required for a therapeutic effect of B. longum JDM 301. Together, our results suggest that personalized microbiome and genetic features may modify the probiotic therapeutic effect.

microbiology

Genetic Diversity Study of Fusarium culmorum: Causal agent of wheat crown rot in Iraq

Fusarium crown rot (FCR), caused by Fusarium culmorum (Wm.G.Sm) Sacc., is an important disease of wheat both in Iraq and other regions of wheat production worldwide. Changes in environmental conditions and cultural practices such as crop rotation generate stress on pathogen populations leading to the evolution of new strains that can tolerate more stressful environments. This study aims to investigate the genetic diversity among isolates of F. culmorum in Iraq. Twenty-nine samples were collected from different regions of wheat cultivation in Iraq to investigate the pathogenicity and genetic diversity of F. culmorum using the REP-PCR technique. Among the twenty-nine isolates of F. culmorum examined for pathogenicity, 96% were pathogenic to wheat at the seedling stage. The most aggressive isolate, from Baghdad, was IF 0021 at 0.890 on the FCR severity index. Three primer sets were used to assess the genotypic diversity via REP, ERIC and BOX elements. The amplicon sizes ranged from 200-800 bp for BOX-ERIC2, 110-1100 bp for ERIC-ERIC2 and 200-1300 bp for REP. In total, 410 markers were polymorphic, including 106 for BOX, 175 for ERIC and 129 for the REP. Genetic similarity was calculated by comparing markers according to minimum variance (Squared Euclidean). Clustering analysis generated two major groups, group 1 with two subgroup 1a and 1b with 5 and 12 isolates respectively, and group 2 with two subgroups 2a and 2b with 3 and 9 isolates respectively. This is the first study in this field that has been reported in Iraq.

microbiology

High-resolution genetic map and QTL analysis of growth-related traits of Hevea brasiliensis cultivated under suboptimal temperature and humidity conditions

Rubber tree (Hevea brasiliensis) cultivation is the main source of natural rubber worldwide and has been extended to areas with suboptimal climates and lengthy drought periods; this transition affects growth and latex production. High-density genetic maps with reliable markers support precise mapping of quantitative trait loci (QTL), which can help reveal the complex genome of the species, provide tools to enhance molecular breeding, and shorten the breeding cycle. In this study, QTL mapping of the stem diameter, tree height, and number of whorls was performed for a full-sibling population derived from a GT1 and RRIM701 cross. A total of 225 simple sequence repeats (SSRs) and 186 single-nucleotide polymorphism (SNP) markers were used to construct a base map with 18 linkage groups and to anchor 671 SNPs from genotyping by sequencing (GBS) to produce a very dense linkage map with small intervals between loci. The final map was composed of 1,079 markers, spanned 3,779.7 cM with an average marker density of 3.5 cM, and showed collinearity between markers from previous studies. Significant variation in phenotypic characteristics was found over a 59-month evaluation period with a total of 38 QTLs being identified through a composite interval mapping method. Linkage group 4 showed the greatest number of QTLs (7), with phenotypic explained values varying from 7.67% to 14.07%. Additionally, we estimated segregation patterns, dominance, and additive effects for each QTL. A total of 53 significant effects for stem diameter were observed, and these effects were mostly related to additivity in the GT1 clone. Associating accurate genome assemblies and genetic maps represents a promising strategy for identifying the genetic basis of phenotypic traits in rubber trees. Then, further research can benefit from the QTLs identified herein, providing a better understanding of the key determinant genes associated with growth of Hevea brasiliensis under limiting water conditions.

plant biology

Genetic Diversity Patterns and Domestication Origin of Soybean

Understanding diversity and evolution of a crop is an essential step to implement a strategy to expand its germplasm base for crop improvement research. Samples intensively collected from Korea, which is a small but central region in the distribution geography of soybean, were genotyped to provide sufficient data to underpin genome-wide population genetic questions. After removing natural hybrids and duplicated or redundant accessions, we obtained a non-redundant set comprising 1,957 domesticated and 1,079 wild accessions to perform population structure analyses. Our analysis demonstrates that while wild soybean germplasm will require additional sampling from diverse indigenous areas to expand the germplasm base, the current domesticated soybean germplasm is saturated in terms of genetic diversity. We then showed that our genome-wide polymorphism map enabled us to detect genetic loci underling flower color, seed-coat color, and domestication syndrome. A representative soybean set consisting of 194 accessions were divided into one domesticated subpopulation and four wild subpopulations that could be traced back to their geographic collection areas. Population genomics analyses suggested that the monophyletic group of domesticated soybeans was originated in eastern Japan. The results were further substantiated by a phylogenetic tree constructed from domestication-associated single nucleotide polymorphisms identified in this study.

plant biology

High Genetic Potential for Proteolytic Decomposition in Northern Peatland Ecosystems

AbstractNitrogen (N) is a scarce nutrient commonly limiting primary productivity. Microbial decomposition of complex carbon (C) into small organic molecules (e.g., free amino acids) has been suggested to supplement biologically-fixed N in high latitude peatlands. We evaluated the microbial (fungal, bacterial, and archaeal) genetic potential for organic N depolymerization in peatlands at Marcell Experimental Forest (MEF) in northern Minnesota. We used guided gene assembly to examine the abundance and diversity of protease genes; and further compared to those of N-fixing (nifH) genes in shotgun metagenomic data collected across depth at two distinct peatland environments (bogs and fens). Microbial proteases greatly outnumbered nifH genes with the most abundant gene families (archaeal M1 and bacterial Trypsin) each containing more sequences than all sequences attributed to nifH. Bacterial protease gene assemblies were diverse and abundant across depth profiles, indicating a role for bacteria in releasing free amino acids from peptides through depolymerization of older organic material and contrasting the paradigm of fungal dominance in depolymerization in forest soils. Although protease gene assemblies for fungi were much less abundant overall than for bacteria, fungi were prevalent in surface samples and therefore may be vital in degrading large soil polymers from fresh plant inputs during early stage of depolymerization. In total, we demonstrate that depolymerization enzymes from a diverse suite of microorganisms, including understudied bacterial and archaeal lineages, are likely to play a substantial role in C and N cycling within northern peatlands.\n\nImportanceNitrogen (N) is a common limitation on primary productivity, and its source remains unresolved in northern peatlands that are vulnerable to environmental change. Decompositionof complex organic matter into free amino acids has been proposed as an important N source, but the genetic potential of microorganisms mediating this process has not been examined. Such information can elucidate possible responses of northern peatlands to environmental change. We show high genetic potential for microbial production of free amino acids across a range of microbial guilds. In particular, the abundance and diversity of bacterial genes encoding proteolytic activity suggests a predominant role for bacteria in regulating productivity and contrasts a paradigm of fungal dominance of organic N decomposition. Our results expand our current understanding of coupled carbon and nitrogen cycles in north peatlands and indicate that understudied bacterial and archaeal lineages may be central in this ecosystems response to environmental change.

ecology

Building comprehensive MS-friendly databases for proteomic analysis of bacterial species of unknown genetic background

In proteomics, peptide information within mass spectrometry data from a specific organism sample is routinely challenged against a protein sequence database that best represent such organism. However, if the species/strain in the sample is unknown or poorly genetically characterized, it becomes challenging to determine a database which can represent such sample. Building customized protein sequence databases merging multiple strains for a given species has become a strategy to overcome such restrictions. However, as more genetic information is publicly available and interesting genetic features such as the existence of pan- and core genes within a species are revealed, we questioned how efficient such merging strategies are to report relevant information. To test this assumption, we constructed databases containing conserved and unique sequences for ten different species. Features that are relevant for probabilistic-based protein identification by proteomics were then monitored. As expected, increase in database complexity correlates with pangenomic complexity. However, Mycobacterium tuberculosis and Bortedella pertusis generated very complex databases even having low pangenomic complexity or no pangenome at all. This suggests that discrepancies in gene annotation is higher than average between strains of those species. We further tested database performance by using mass spectrometry data from eight clinical strains from Mycobacterium tuberculosis, and from two published datasets from Staphylococcus aureus. We show that by using an approach where database size is controlled by removing repeated identical tryptic sequences across strains/species, computational time can be reduced drastically as database complexity increases.

bioinformatics

An automated, high-throughput image analysis pipeline enables genetic studies of shoot and root morphology in carrot (Daucus carota L.)

Carrot is a globally important crop, yet efficient and accurate methods for quantifying its most important agronomic traits are lacking. To address this problem, we developed an automated analysis platform that extracts components of size and shape for carrot shoots and roots, which are necessary to advance carrot breeding and genetics. This method reliably measured variation in shoot size and shape, leaf number, petiole length, and petiole width as evidenced by high correlations with hundreds of manual measurements. Similarly, root length and biomass were accurately measured from the images. This platform quantified shoot and root shapes in terms of principal components, which do not have traditional, manually-measurable equivalents. We applied the pipeline in a study of a six-parent diallel population and an F2 mapping population consisting of 316 individuals. We found high levels of repeatability within a growing environment, with low to moderate repeatability across environments. We also observed co-localization of quantitative trait loci for shoot and root characteristics on chromosomes 1, 2, and 7, suggesting these traits are controlled by genetic linkage and/or pleiotropy. By increasing the number of individuals and phenotypes that can be reliably quantified, the development of a high-throughput image analysis pipeline to measure carrot shoot and root morphology will expand the scope and scale of breeding and genetic studies.

plant biology

Transforming summary statistics from logistic regression to the liability scale: application to genetic and environmental risk scores

1. Abstract1.1. ObjectiveStratified medicine requires models of disease risk incorporating genetic and environmental factors. These may combine estimates from different studies and models must be easily updatable when new estimates become available. The logit scale is often used in genetic and environmental association studies however the liability scale is used for polygenic risk scores and measures of heritability, but combining parameters across studies requires a common scale for the estimates.\n\n1.2. MethodsWe present equations to approximate the relationship between univariate effect size estimates on the logit scale and the liability scale, allowing model parameters to be translated between scales.\n\n1.3. ResultsThese equations are used to build a risk score on the liability scale, using effect size estimates originally estimated on the logit scale. Such a score can then be used in a joint effects model to estimate the risk of disease, and this is demonstrated for schizophrenia using a polygenic risk score and environmental risk factors.\n\n1.4. ConclusionThis straightforward method allows conversion of model parameters between the logit and liability scales, and may be a key tool to integrate risk estimates into a comprehensive risk model, particularly for joint models with environmental and genetic risk factors.

epidemiology

Human pancreatic islet 3D chromatin architecture provides insights into the genetics of type 2 diabetes

Genetic studies promise to provide insight into the molecular mechanisms underlying type 2 diabetes (T2D). Variants associated with T2D are often located in tissue-specific enhancer regions (enhancer clusters, stretch enhancers or super-enhancers). So far, such domains have been defined through clustering of enhancers in linear genome maps rather than in 3D-space. Furthermore, their target genes are generally unknown. We have now created promoter capture Hi-C maps in human pancreatic islets. This linked diabetes-associated enhancers with their target genes, often located hundreds of kilobases away. It further revealed sets of islet enhancers, super-enhancers and active promoters that form 3D higher-order hubs, some of which show coordinated glucose-dependent activity. Hub genetic variants impact the heritability of insulin secretion, and help identify individuals in whom genetic variation of islet function is important for T2D. Human islet 3D chromatin architecture thus provides a framework for interpretation of T2D GWAS signals.

genomics

Genetic structure of invasive babys breath (Gypsophila paniculata) populations in a freshwater Michigan dune system

Coastal sand dunes are dynamic ecosystems with elevated levels of disturbance, and as such they are highly susceptible to plant invasions. One such invasion that is of major concern to the Great Lakes dune systems is that of perennial babys breath (Gypsophila paniculata). The invasion of babys breath negatively impacts native species such as the federal threatened Pitchers thistle (Cirsium pitcheri) that occupy the open sand habitat of the Michigan dune system. Our research goals were to (1) quantify the genetic diversity of invasive babys breath populations in the Michigan dune system, and (2) estimate the genetic structure of these invasive populations. We analyzed 12 populations at 14 nuclear and 2 chloroplast microsatellite loci. We found strong genetic structure among populations of babys breath sampled along Michigans dunes (global FST = 0.228), and also among two geographic regions that are separated by the Leelanau peninsula. Pairwise comparisons using the nSSR data among all 12 populations yielded significant FST values. Results from a Bayesian clustering analysis suggest two main population clusters. Isolation by distance was found over all 12 populations (R = 0.755, P < 0.001) and when only cluster 2 populations were included (R = 0.523, P = 0.030); populations within cluster 1 revealed no significant relationship (R = 0.205, P = 0.494). Private nSSR alleles and cpSSR haplotypes within each cluster suggest the possibility of at least two separate introduction events to Michigan.

plant biology

GADMA: Genetic Algorithm for Automatic Inferring Joint Demographic History of Multiple Populations from Allele Frequency Spectrum

The demographic history of any population is imprinted in the genomes of the individuals that make up the population. One of the most popular and convenient representations of genetic information is the allele frequency spectrum or AFS, the distribution of allele frequencies in populations. The joint allele frequency spectrum is commonly used to reconstruct the demographic history of multiple populations and several methods based on diffusion approximation (e.g.,{partial} a{partial}i) and ordinary differential equations (e.g., moments) have been developed and applied for demographic inference. These methods provide an opportunity to simulate AFS under a variety of researcher-specified demographic models and to estimate the best model and associated parameters using likelihood-based local optimizations. However, there are no known algorithms to perform global searches of demographic models with a given AFS. Here, we introduce a new method that implements a global search using a genetic algorithm for the automatic and unsupervised inference of demographic history from joint allele frequency spectrum data. Our method is implemented in the software GADMA (Genetic Algorithm for Demographic Analysis, https://github.com/ctlab/GADMA). We demonstrate the performance of GADMA by applying it to sequence data from humans and non-model organisms and show that it is able to automatically infer a demographic model close to or even better than the one that was previously obtained manually. Moreover, GADMA is able to infer demographic models at different local optima close to the global one, making it is possible to detect more biology corrected model during further research.

evolutionary biology

Genetic correlation between sea age at maturity and iteroparity in Atlantic salmon.

Genetic correlations in life history traits may result in unpredictable evolutionary trajectories if not accounted for in life-history models. Iteroparity (the reproductive strategy of reproducing more than once) in Atlantic salmon (Salmo salar) is a fitness trait with substantial variation within and among populations. In the Teno River in northern Europe, iteroparous individuals constitute an important component of many populations and have experienced a sharp increase in abundance in the last 20 years, partly overlapping with a general decrease in age structure. The physiological basis of iteroparity bears similarities to that of age at first maturity, another life history trait with substantial fitness effects in salmon. Sea age at maturity in Atlantic salmon is controlled by a major locus around the vgll3 gene, and we used this opportunity demonstrate that the two traits are genetically correlated around this genome region. The odds ratio of survival until second reproduction was up to 2.4 (1.8-3.5 90% CI) times higher for fish with the early-maturing vgll3 genotype (EE) compared to fish with the late-maturing genotype (LL). The association had a dominance architecture, although the dominant allele was reversed in the late-maturing group compared to younger groups that stayed only one year at sea before maturation. Post hoc analysis indicated that iteroparous fish with the EE genotype had accelerated growth prior to first reproduction compared to first-time spawners, across all age groups, while this effect was not detected in fish with the LL genotype. These results broaden the functional link around the vgll3 genome region and help us understand constraints in the evolution of life history variation in salmon. Our results further highlight the need to account for genetic correlations between fitness traits when predicting demographic changes in changing environments.

evolutionary biology

Assessment Of Genetic Structure Of The Endangered Forest Species Boswellia Serrata Roxb. Population In Central India

Boswellia serrata Roxb., a commercially important species for its pulp and pharmaceutical properties was sampled from three locations representing its natural distribution in central India for genetic characterization through 56 RAPD + 42 ISSR loci. The wood fiber dimensions measured for morphometric characterization confirmed 11.36% of the variation in the length and 8.75% of the variation in the width indicating its fitness for local adaptation. Bayesian and non-Bayesian approach based diversity measures resulted moderate within population gene diversity (0.26{+/-}0.17), Shannons information index (0.40{+/-}0.22) and panmictic heterozygosity (0.28{+/-}0.01). A high estimate for genetic differentiation measures i.e. GST (0.31), GST-B (0.33{+/-}0.02) and {theta}-II (0.45) led to the distinct clusters of the sampled genotypes representing their regional variability due to limited gene flow and total absence of natural regeneration. We report the first investigation of the species for its molecular characterization emphasizing the urgent need for the genetic improvement program for the In-situ/Ex-situ conservation and sustainable commercialization.

plant biology

Repeated phenotypic evolution by different genetic routes: the evolution of colony switching in Pseudomonas fluorescens SBW25

Repeated evolution of functionally similar phenotypes is observed throughout the tree of life. The extent to which the underlying genetics are conserved remains an area of considerable interest. Previously, we reported the evolution of colony switching in two independent lineages of Pseudomonas fluorescens SBW25 (Beaumont et al., 2009). The phenotypic and genotypic bases of colony switching in the first lineage (Line 1) have been described elsewhere (Beaumont et al., 2009; Gallie et al., 2015). Here, we deconstruct the evolution of colony switching in the second lineage (Line 6). We show that, as for Line 1, Line 6 colony switching results from an increase in the expression of a colanic acid-like polymer (CAP). At the genetic level, nine mutations occur in Line 6. Only one of these - a non-synonymous point mutation in the housekeeping sigma factor rpoD - is required for colony switching. In contrast, the genetic basis of colony switching in Line 1 is a mutation in the metabolic gene carB (Beaumont et al., 2009). A molecular model has recently been proposed whereby the carB mutation increases capsulation by redressing the intracellular balance of positive (ribosomes) and negative (RsmAE/CsrA) regulators of a positive feedback loop in capsule expression (Remigi et al., 2018). We show that Line 6 colony switching is consistent with this model; the rpoD mutation generates an increase in ribosome expression, and ultimately an increase in CAP expression.

evolutionary biology