bioRxiv ScienceSearch

Biology subjects

Hwang, S.

Publications and source records attributed to Hwang, S..

5 recordsLinked to original sources

ForestQC: quality control on genetic variants from next-generation sequencing data using random forest

Next-generation sequencing technology (NGS) enables discovery of nearly all genetic variants present in a genome. A subset of these variants, however, may have poor sequencing quality due to limitations in sequencing technology or in variant calling algorithms. In genetic studies that analyze a large number of sequenced individuals, it is critical to detect and remove those variants with poor quality as they may cause spurious findings. In this paper, we present a statistical approach for performing quality control on variants identified from NGS data by combining a traditional filtering approach and a machine learning approach. Our method uses information on sequencing quality such as sequencing depth, genotyping quality, and GC contents to predict whether a certain variant is likely to contain errors. To evaluate our method, we applied it to two whole-genome sequencing datasets where one dataset consists of related individuals from families while the other consists of unrelated individuals. Results indicate that our method outperforms widely used methods for performing quality control on variants such as VQSR of GATK by considerably improving the quality of variants to be included in the analysis. Our approach is also very efficient, and hence can be applied to large sequencing datasets. We conclude that combining a machine learning algorithm trained with sequencing quality information and the filtering approach is an effective approach to perform quality control on genetic variants from sequencing data.\n\nAuthor SummaryGenetic disorders can be caused by many types of genetic mutations, including common and rare single nucleotide variants, structural variants, insertions and deletions. Nowadays, next generation sequencing (NGS) technology allows us to identify various genetic variants that are associated with diseases. However, variants detected by NGS might have poor sequencing quality due to biases and errors in sequencing technologies and analysis tools. Therefore, it is critical to remove variants with low quality, which could cause spurious findings in follow-up analyses. Previously, people applied either hard filters or machine learning models for variant quality control (QC), which failed to filter out those variants accurately. Here, we developed a statistical tool, ForestQC, for variant QC by combining a filtering approach and a machine learning approach. We applied ForestQC to one family-based whole genome sequencing (WGS) dataset and one general case-control WGS dataset, to evaluate our method. Results show that ForestQC outperforms widely used methods for variant QC by considerably improving the quality of variants. Also, ForestQC is very efficient and scalable to large-scale sequencing datasets. Our study indicates that combining filtering approaches and machine learning approaches enables effective variant QC.

bioinformatics

The Effect of Aldehyde Dehydrogenase Activator, Alda-1(R), on the Ethanol-induced Brain Damage in a Rat of Binge Ethanol Intoxication.

AimsThis study aimed to investigate whether an aldehyde dehydrogenase (ALDH) activator (Alda-1(R)) reduces neuronal damage in a rat model of binge ethanol exposure.\n\nMethodsThirty-six adolescent male rats (130-150 g) were randomly assigned into three groups: sham, ethanol-only group (25% ethanol intragastrically thrice daily for four days, approximately 10 g/kg/day) and ethanol with Alda-1(R) group (10 mg/kg thrice daily for four days). The ALDH activity at baseline and 90 min after the last infusion in each group was measured. Brain damage was investigated using Luxol fast blue-Cresyl violet staining in the hippocampus, CA1 and CA2/3. The activation of astrocytes and microglia was examined using immunohistochemistry for antiglial fibrillary acidic protein (GFAP) and anti-ionized calcium-binding adapter molecule 1 (Iba-1).\n\nResultsAfter a four-day binge, the ALDH activity level was doubled in the ethanol with Alda-1(R) group (mean: 7.87, SD: 0.67), whereas the levels in the sham group (mean: 4.07, SD: 0.53) and ethanol-only group (mean: 3.77, SD: 0.36) were slightly decreased. More significant neuronal shrinkage, fewer neurons, and loss of Nissl in the hippocampus were observed in the ethanol-only group compared to the ethanol with Alda-1(R) group. Astrocytosis and microgliosis of the hippocampus also showed increased activation in the ethanol only group compared with the ethanol with Alda-1(R) group.\n\nConclusionAlda-1(R) administration reduces cytotoxic damage to the hippocampus in adolescent rats with binge ethanol exposure.

neuroscience

Mutation supply and the repeatability of selection for antibiotic resistance

Whether evolution can be predicted is a key question in evolutionary biology. Here we set out to better understand the repeatability of evolution, which is a necessary condition for predictability. We explored experimentally the effect of mutation supply and the strength of selective pressure on the repeatability of selection from standing genetic variation. Different sizes of mutant libraries of an antibiotic resistance gene, TEM-1 {beta}-lactamase in Escherichia coli, were subjected to different antibiotic concentrations. We determined whether populations went extinct or survived, and sequenced the TEM gene of the surviving populations. The distribution of mutations per allele in our mutant libraries--generated by error-prone PCR--followed a Poisson distribution. Extinction patterns could be explained by a simple stochastic model that assumed the sampling of beneficial mutations was key for survival. In most surviving populations, alleles containing at least one known large-effect beneficial mutation were present. These genotype data also support a model which only invokes sampling effects to describe the occurrence of alleles containing large-effect driver mutations. Hence, evolution is largely predictable given cursory knowledge of mutational fitness effects, the mutation rate and population size. There were no clear trends in the repeatability of selected mutants when we considered all mutations present. However, when only known large-effect mutations were considered, the outcome of selection is less repeatable for large libraries, in contrast to expectations. Furthermore, we show experimentally that alleles carrying multiple mutations selected from large libraries confer higher resistance levels relative to alleles with only a known large-effect mutation, suggesting that the scarcity of high-resistance alleles carrying multiple mutations may contribute to the decrease in repeatability at large library sizes.

evolutionary biology

WheatNet: A genome-scale functional network for hexaploid bread wheat, Triticum aestivum

Gene networks provide a system-level overview of genetic organizations and enable the dissection of functional modules underlying complex traits. Here we report the generation of WheatNet, the first genome-scale functional network for T. aestivum and a companion web server (www.inetbio.org/wheatnet). WheatNet was constructed by integrating 20 distinct genomics datasets, including 156,000 wheat-specific co-expression links mined from 1,929 microarray data. A unique feature of WheatNet is that each network node represents either a single gene or a group of genes. We computationally partitioned gene groups mimicking homeologous genes by clustering 99,386 wheat genes, resulting in 20,248 gene groups comprising 63,401 genes and 35,985 individual genes. Thus, WheatNet was constructed using 56,233 nodes, and the final integrated network has 20,230 nodes and 567,000 edges. The edge information of the integrated WheatNet and all 20 component networks are available for download.

genetics

Genotypic complexity of Fisher’s geometric model

Fishers geometric model was originally introduced to argue that complex adaptations must occur in small steps because of pleiotropic constraints. When supplemented with the assumption of additivity of mutational effects on phenotypic traits, it provides a simple mechanism for the emergence of genotypic epistasis from the nonlinear mapping of phenotypes to fitness. Of particular interest is the occurrence of reciprocal sign epistasis, which is a necessary condition for multipeaked genotypic fitness landscapes. Here we compute the probability that a pair of randomly chosen mutations interacts sign-epistatically, which is found to decrease with increasing phenotypic dimension n, and varies non-monotonically with the distance from the phenotypic optimum. We then derive expressions for the mean number of fitness maxima in genotypic landscapes composed of all combinations of L random mutations. This number increases exponentially with L, and the corresponding growth rate is used as a measure of the complexity of the landscape. The dependence of the complexity on the model parameters is found to be surprisingly rich, and three distinct phases characterized by different landscape structures are identified. Our analysis shows that the phenotypic dimension, which is often referred to as phenotypic complexity, does not generally correlate with the complexity of fitness landscapes and that even organisms with a single phenotypic trait can have complex landscapes. Our results further inform the interpretation of experiments where the parameters of Fisher's model have been inferred from data, and help to elucidate which features of empirical fitness landscapes can be described by this model.

evolutionary biology