bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,765 records · Page 98Linked to original sources

BATL: Bayesian annotations for targeted lipidomics

MotivationBioinformatic tools capable of annotating, rapidly and reproducibly, large, targeted lipidomic datasets are limited. Specifically, few programs enable high-throughput peak assessment of liquid chromatography-electrospray ionization tandem mass spectrometry (LC-ESI-MS/MS) data acquired in either selected or multiple reaction monitoring (SRM and MRM) modes. ResultsWe present here Bayesian Annotations for Targeted Lipidomics (BATL), a Gaussian naive Bayes classifier for targeted lipidomics that annotates peak identities according to eight features related to retention time, intensity, and peak shape. Lipid identification is achieved by modelling distributions of these eight input features across biological conditions and maximizing the joint posterior probabilities of all peak identities at a given transition. When applied to sphingolipid and glycerophosphocholine SRM datasets, we demonstrate over 95% of all peaks are rapidly and correctly identified. Availability and implementationBATL software is freely accessible online at https://complimet.ca/batl/ and is compatible with Safari, Firefox, Chrome and Edge. Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

CONSULT: Accurate contamination removal using locality-sensitive hashing

A fundamental question appears in many bioinformatics applications: Does a sequencing read belong to a large dataset of genomes from some broad taxonomic group, even when the closest match in the set is evolutionarily divergent from the query? For example, low-coverage genome sequencing (skimming) projects either assemble the organelle genome or compute genomic distances directly from unassembled reads. Using unassembled reads needs contamination detection because samples often include reads from unintended groups of species. Similarly, assembling the organelle genome needs distinguishing organelle and nuclear reads. While k-mer-based methods have shown promise in read-matching, prior studies have shown that existing methods are insufficiently sensitive for contamination detection. Here, we introduce a new read-matching tool called CONSULT that tests whether k-mers from a query fall within a user-specified distance of the reference dataset using locality-sensitive hashing. Taking advantage of large memory machines available nowadays, CONSULT libraries accommodate tens of thousands of microbial species. Our results show that CONSULT has higher true-positive and lower false-positive rates of contamination detection than leading methods such as Kraken-II and improves distance calculation from genome skims. We also demonstrate that CONSULT can distinguish organelle reads from nuclear reads, leading to dramatic improvements in skims-based mitochondrial assemblies.

bioinformatics↗

Population-specific genome graphs improve high-throughput sequencing data analysis: A case study on the Pan-African genome

Graph-based genome reference representations have seen significant development, motivated by the inadequacy of the current human genome reference to represent the diverse genetic information from different human populations and its inability to maintain the same level of accuracy for non-European ancestries. While there have been many efforts to develop computationally efficient graph-based toolkits for NGS read alignment and variant calling, methods to curate genomic variants and subsequently construct genome graphs remains an understudied problem that inevitably determines the effectiveness of the overall bioinformatics pipeline. In this study, we discuss obstacles encountered during graph construction and propose methods for sample selection based on population diversity, graph augmentation with structural variants and resolution of graph reference ambiguity caused by information overload. Moreover, we present the case for iteratively augmenting tailored genome graphs for targeted populations and demonstrate this approach on the whole-genome samples of African ancestry. Our results show that population-specific graphs, as more representative alternatives to linear or generic graph references, can achieve significantly lower read mapping errors and enhanced variant calling sensitivity, in addition to providing the improvements of joint variant calling without the need of computationally intensive post-processing steps.

bioinformatics↗

Asc-Seurat - Analytical single-cell Seurat-based web application

SummarySingle-cell RNA sequencing (scRNA-seq) has become a popular approach for studying the transcriptome, providing a powerful tool for discovering and characterizing cell types and their developmental trajectories. However, scRNA-seq analysis is complex, requiring a continuous, iterative process to refine the data processing and uncover relevant biological information. We present Asc-Seurat, a feature rich workbench, providing a user-friendly and easy-to-install web application encapsulating the necessary tools for an all-encompassing and fluid scRNA-seq data analysis. Availability and implementationAsc-Seurat is available at https://github.com/KirstLab/asc_seurat/ and released under GNU 3 license. Contactmkirst@ufl.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

BubbleGun: Enumerating Bubbles and Superbubbles in Genome Graphs

MotivationWith the fast development of third generation sequencing machines, de novo genome assembly is becoming a routine even for larger genomes. Graph-based representations of genomes arise both as part of the assembly process, but also in the context of pangenomes representing a population. In both cases, polymorphic loci lead to bubble structures in such graphs. Detecting bubbles is hence an important task when working with genomic variants in the context of genome graphs. ResultsHere, we present a fast general-purpose tool, called BubbleGun, for detecting bubbles and superbubbles in genome graphs. Furthermore, BubbleGun detects and outputs runs of linearly connected bubbles and superbubbles, which we call bubble chains. We showcase its utility on de Bruijn graphs and compare our results to vgs snarl detection. We show that BubbleGun is considerably faster than vg especially in bigger graphs, where it reports all bubbles in less than 30 minutes on a human sample de Bruijn graph of around 2 million nodes. AvailabilityBubbleGun is available and documented at https://github.com/fawaz-dabbaghieh/bubble_gun under MIT license. Contactfawaz@hhu.de or tobias.marschall@hhu.de Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Factors Associated with Emerging and Re-emerging of SARS-CoV-2 Variants

Global spread of Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2) has triggered unprecedented scientific efforts, as well as containment and treatment measures. Despite these efforts, SARS-CoV-2 infections remain unmanageable in some parts of the world. Due to inherent mutability of RNA viruses, it is not surprising that the SARS-CoV-2 genome has been continuously evolving since its emergence. Recently, four functionally distinct variants, B.1.1.7, B.1.351, P.1 and CAL.20C, have been identified, and they appear to more infectious and transmissible than the original (Wuhan-Hu-1) virus. Here we provide evidence based upon a combination of bioinformatics and structural approaches that can explain the higher infectivity of the new variants. Our results show that the greater infectivity of SARS-CoV-2 than SARS-CoV can be attributed to a combination of several factors, including alternate receptors. Additionally, we show that new SARS-CoV-2 variants emerged in the background of D614G in Spike protein and P323L in RNA polymerase. The correlation analyses showed that all mutations in specific variants did not evolve simultaneously. Instead, some mutations evolved most likely to compensate for the viral fitness.

bioinformatics↗

ModPhred: an integrative toolkit for the analysis and storage of nanopore sequencing DNA and RNA modification data

MotivationDNA and RNA modifications can now be identified using Nanopore sequencing. However, we currently lack a flexible software to efficiently encode, store, analyze and visualize DNA and RNA modification data. ResultsHere we present ModPhred, a versatile toolkit that facilitates DNA and RNA modification analysis from nanopore sequencing reads in a user-friendly manner. ModPhred integrates probabilistic DNA and RNA modification information within the FASTQ and BAM file formats, can be used to encode multiple types of modifications simultaneously, and its output can be easily coupled to genomic track viewers, facilitating the visualization and analysis of DNA and RNA modification information in individual reads in a simple and computationally efficient manner. Availability and ImplementationModPhred is available at https://github.com/novoalab/modPhred, is implemented in Python3, and is released under an MIT license. Supplementary DataSupplementary Data are available at Bioinformatics online.

bioinformatics↗

Haplotype-aware pantranscriptome analyses using spliced pangenome graphs

Pangenomics is emerging as a powerful computational paradigm in bioinformatics. This field uses population-level genome reference structures, typically consisting of a sequence graph, to mitigate reference bias and facilitate analyses that were challenging with previous reference-based methods. In this work, we extend these methods into transcriptomics to analyze sequencing data using the pantranscriptome: a population-level transcriptomic reference. Our novel toolchain can construct spliced pangenome graphs, map RNA-seq data to these graphs, and perform haplotype-aware expression quantification of transcripts in a pantranscriptome. This workflow improves accuracy over state-of-the-art RNA-seq mapping methods, and it can efficiently quantify haplotype-specific transcript expression without needing to characterize a samples haplotypes beforehand.

bioinformatics↗

KinaFrag explores the kinase-ligand fragment interaction space for selective kinase inhibitor discovery

Protein kinases play a crucial role in many cellular signaling processes, making them one of the most important families of drug targets. But selectivity put a barrier at the design of kinase inhibitors. Fragment-based drug design strategies have been successfully applied to develop novel selective kinase inhibitors. However, the complicate kinase-fragment interaction and fragment-to-lead process pose challenges to fragment-based kinase discovery. Here, we developed a web source KinaFrag to investigate kinase-fragment interaction space and perform fragment-to-lead optimization. KinaFrag contained 31464 fragments from reported kinase inhibitors, which involved 3244 crystal fragment structures and 7783 crystal kinase-fragment complexes. These crystal fragments were classified by their binding cleft and subpockets, and their 3D structure and interactions were displayed in KinaFrag. In addition, the structural information, physicochemical information, similarity information, and substructure relationship information were contained in KinaFrag. Moreover, a computational fragment growing strategy obviously developed by our group was implemented in the KinaFrag. We test this fragment growing strategy using our fragment libraries, and obtained a lead compound of c-Met with ~1000-fold in vitro activity improvement compared with the hit compound. We hope KinaFrag could become a powerful tool for the fragment-based kinase inhibitor design. KinaFrag is freely available at http://chemyang.ccnu.edu.cn/ccb/database/KinaFrag/. Biographical noteZhi-Zheng Wang is a PhD student at College of Chemistry, Central China Normal University (CCNU), and the direction of his thesis is computational molecular simulation. Xing-Xing Shi is PhD student at College of Chemistry, of CCNU, and the direction of her thesis is computational molecular simulation. Fan Wang is a lecturer at College of Chemistry of CCNU. He received the PhD degree in Computational Chemistry from University of Amiens, France. Ge-Fei Hao is Professor in Bioinformatics in College of Chemistry of CCNU. He received his PhD in Pesticide Science from CCNU. Guang-Fu Yang is Professor in Chemical Biology. He is the group leader and has been the Dean at College of Chemistry of CCNU. He received the PhD degree in Pesticide Science from Nankai University, Tianjin, China.

bioinformatics↗

ALOHA: Aggregated local extrema splines for high-throughput dose-response analysis

Computational methods for genomic dose-response integrate dose-response modeling with bioinformatics tools to evaluate changes in molecular and cellular functions related to pathogenic processes. These methods use parametric models to describe each genes dose-response, but such models may not adequately capture expression changes. Additionally, current approaches do not consider gene co-expression networks. When assessing co-expression networks, one typically does not consider the dose-response relationship, resulting in co-regulated gene sets containing genes having different dose-response patterns. To avoid these limitations, we develop an analysis pipeline called Aggregated Local Extrema Splines for High-throughput Analysis (ALOHA), which computes individual genomic dose-response functions using a flexible class Bayesian shape constrained splines and clusters gene co-regulation based upon these fits. Using splines, we reduce information loss due to parametric lack-of-fit issues, and because we cluster on dose-response relationships, we better identify co-regulation clusters for genes that have co-expressed dose-response patterns from chemical exposure. The clustered pathways can then be used to estimate a dose associated with a pre-specified biological response, i.e., the benchmark dose (BMD), and approximate a point of departure dose corresponding to minimal adverse response in the whole tissue/organism. We compare our approach to current parametric methods and our biologically enriched gene sets to cluster on normalized expression data. Using this methodology, we can more effectively extract the underlying structure leading to more cohesive estimates of gene set potency.

bioinformatics↗

OGUs enable effective, phylogeny-aware analysis of even shallow metagenome community structures

We introduce Operational Genomic Unit (OGU), a metagenome analysis strategy that directly exploits sequence alignment hits to individual reference genomes as the minimum unit for assessing the diversity of microbial communities and their relevance to environmental factors. This approach is independent from taxonomic classification, granting the possibility of maximal resolution of community composition, and organizes features into an accurate hierarchy using a phylogenomic tree. The outputs are suitable for contemporary analytical protocols for community ecology, differential abundance and supervised learning while supporting phylogenetic methods, such as UniFrac and phylofactorization, that are seldomly applied to shotgun metagenomics despite being prevalent in 16S rRNA gene amplicon studies. As demonstrated in one synthetic and two real-world case studies, the OGU method produces biologically meaningful patterns from microbiome datasets. Such patterns further remain detectable at very low metagenomic sequencing depths. Compared with taxonomic unit-based analyses implemented in currently adopted metagenomics tools, and the analysis of 16S rRNA gene amplicon sequence variants, this method shows superiority in informing biologically relevant insights, including stronger correlation with body environment and host sex on the Human Microbiome Project dataset, and more accurate prediction of human age by the gut microbiomes in the Finnish population. We provide Woltka, a bioinformatics tool to implement this method, with full integration with the QIIME 2 package and the Qiita web platform, to facilitate OGU adoption in future metagenomics studies. ImportanceShotgun metagenomics is a powerful, yet computationally challenging, technique compared to 16S rRNA gene amplicon sequencing for decoding the composition and structure of microbial communities. However, current analyses of metagenomic data are primarily based on taxonomic classification, which is limited in feature resolution compared to 16S rRNA amplicon sequence variant analysis. To solve these challenges, we introduce Operational Genomic Units (OGUs), which are the individual reference genomes derived from sequence alignment results, without further assigning them taxonomy. The OGU method advances current read-based metagenomics in two dimensions: (i) providing maximal resolution of community composition while (ii) permitting use of phylogeny-aware tools. Our analysis of real-world datasets shows several advantages over currently adopted metagenomic analysis methods and the finest-grained 16S rRNA analysis methods in predicting biological traits. We thus propose the adoption of OGU as standard practice in metagenomic studies.

bioinformatics↗

ECNano: A Cost-Effective Workflow for Target Enrichment Sequencing and Accurate Variant Calling on 4,800 Clinically Significant Genes Using a Single MinION Flowcell

BackgroundThe application of long-read sequencing using the Oxford Nanopore Technologies (ONT) MinION sequencer is getting more diverse in the medical field. Having a high sequencing error of ONT and limited throughput from a single MinION flowcell, however, limits its applicability for accurate variant detection. Medical exome sequencing (MES) targets clinically significant exon regions, allowing rapid and comprehensive screening of pathogenic variants. By applying MES with MinION sequencing, the technology can achieve a more uniform capture of the target regions, shorter turnaround time, and lower sequencing cost per sample. MethodWe introduced a cost-effective optimized workflow, ECNano, comprising a wet-lab protocol and bioinformatics analysis, for accurate variant detection at 4,800 clinically important genes and regions using a single MinION flowcell. The ECNano wet-lab protocol was optimized to perform long-read target enrichment and ONT library preparation to stably generate high-quality MES data with adequate coverage. The subsequent variant-calling workflow, Clair-ensemble, adopted a fast RNN-based variant caller, Clair, and was optimized for target enrichment data. To evaluate its performance and practicality, ECNano was tested on both reference DNA samples and patient samples. ResultsECNano achieved deep on-target depth of coverage (DoC) at average >100x and >98% uniformity using one MinION flowcell. For accurate ONT variant calling, the generated reads sufficiently covered 98.9% of pathogenic positions listed in ClinVar, with 98.96% having at least 30x DoC. ECNano obtained an average read length of 1,000 bp. The long reads of ECNano also covered the adjacent splice sites well, with 98.5% of positions having [≥] 30x DoC. Clair-ensemble achieved >99% recall and accuracy for SNV calling. The whole workflow from wet-lab protocol to variant detection was completed within three days. ConclusionWe presented ECNano, an out-of-the-box workflow comprising (1) a wet-lab protocol for ONT target enrichment sequencing and (2) a downstream variant detection workflow, Clair-ensemble. The workflow is cost-effective, with a short turnaround time for high accuracy variant calling in 4,800 clinically significant genes and regions using a single MinION flowcell. The long-read exon captured data has potential for further development, promoting the application of long-read sequencing in personalized disease treatment and risk prediction.

bioinformatics↗

On the Analysis of Transcriptional Noise From RNA-sequencing Data

RNA-sequencing (RNA-seq) has revolutionized our understanding of molecular and cellular biology. A central cornerstone in the analysis of RNA-seq is the bioinformatic tools that quantify the data. To evaluate the efficacy of these tools, scientists rely heavily on simulation of RNA-seq. Recently Varabyou et al. took simulation of RNA-seq data to the next level by providing simulated data, that includes simulation of transcriptional noise. While this represents a significant step forward in our ability to perform realistic benchmarks of RNA-seq tools, the data provided by Varabyou et al. need refinement. In the following, I suggest a few improvements with a specific focus on splicing noise. PrefaceI wrote this paper intending to submit it as a Commentary on the Varabyou et al. 2020 Genome Research paper1, but apparently, Genome Research does not publish correspondence-type articles. That is why it is currently on BioRxiv. If you have suggestions about where this paper could potentially be published do not hesitate to contact me.

bioinformatics↗

Clustering single-cell RNA-seq data by rank constrained similarity learning

MotivationRecent breakthroughs of single-cell RNA sequencing (scRNA-seq) technologies offer an exciting opportunity to identify heterogeneous cell types in complex tissues. However, the unavoidable biological noise and technical artifacts in scRNA-seq data as well as the high dimensionality of expression vectors make the problem highly challenging. Consequently, although numerous tools have been developed, their accuracy remains to be improved. ResultsHere, we introduce a novel clustering algorithm and tool RCSL (Rank Constrained Similarity Learning) to accurately identify various cell types using scRNA-seq data from a complex tissue. RCSL considers both local similarity and global similarity among the cells to discern the subtle differences among cells of the same type as well as larger differences among cells of different types. RCSL uses Spearmans rank correlations of a cells expression vector with those of other cells to measure its global similarity, and adaptively learns neighbour representation of a cell as its local similarity. The overall similarity of a cell to other cells is a linear combination of its global similarity and local similarity. RCSL automatically estimates the number of cell types defined in the similarity matrix, and identifies them by constructing a block-diagonal matrix, such that its distance to the similarity matrix is minimized. Each block-diagonal submatrix is a cell cluster/type, corresponding to a connected component in the cognate similarity graph. When tested on 16 benchmark scRNA-seq datasets in which the cell types are well-annotated, RCSL substantially outperformed six state-of-the-art methods in accuracy and robustness as measured by three metrics. AvailabilityThe RCSL algorithm is implemented in R and can be freely downloaded at https://github.com/QinglinMei/RCSL. Contactguojunsdu@gmail.com, zcsu@uncc.edu Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

Germline testing data validate inferences of mutational status for variants detected from tumor-only sequencing

Structured AbstractO_ST_ABSBackgroundC_ST_ABSPathogenic germline variants (PGV) in cancer susceptibility genes are usually identified in cancer patients through germline testing of DNA from blood or saliva: their detection can impact patient treatment options and potential risk reduction strategies for relatives. PGV can also be identified, in tumor sequencing assays, often performed without matched normal specimens. It is then critical to determine whether detected variants are somatic or germline. Here, we evaluate the clinical utility of computational inference of mutational status in tumor-only sequencing compared to germline testing results. Patients and MethodsTumor-only sequencing data from 1,608 patients were retrospectively analyzed to infer germline-versus-somatic status of variants using an information-theoretic, gene-independent approach. Loss of heterozygosity (LOH) was also determined. The predicted mutational models were compared to clinical germline testing results. Statistical measures were computed to evaluate performance. ResultsTumor-only sequencing detected 3,988 variants across 70 cancer susceptibility genes for which germline testing data were available. Our analysis imputed germline-versus-somatic status for >75% of all detected variants, with a sensitivity of 65%, specificity of 88%, and overall accuracy of 86% for pathogenic variants. False omission rate was 3%, signifying minimal error in misclassifying true PGV. A higher portion of PGV in known hereditary tumor suppressors were found to be retained with LOH in the tumor specimens (72%) compared to variants of uncertain significance (58%). ConclusionsTumor-only sequencing provides sufficient power to distinguish germline and somatic variants and infer LOH. Although accurate detection of PGV from tumor-only data is possible, analyzing sequencing data in the context of specimens tumor cell content allows systematic exclusion of somatic variants, and suggests a balance between type 1 and 2 errors for identification of patients with candidate PGV for standard germline testing. Our approach, implemented in a user-friendly bioinformatics application, facilities objective analysis of tumor-only data in clinical settings. HighlightsO_LIMost pathogenic germline variants in cancer predisposition genes can be identified by analyzing tumor-only sequencing data. C_LIO_LIInformation-theoretic gene-independent analysis of common sequencing data accurately infers germline vs. somatic status. C_LIO_LIA reasonable statistical balance can be established between sensitivity and specificity demonstrating clinical utility. C_LIO_LIPathogenic germline variants are more often detected with loss of heterozygosity vs. germline variants of uncertain significance. C_LI

bioinformatics↗

Design and application of a knowledge network for automatic prioritization of drug mechanisms

MotivationDrug repositioning is an attractive alternative to de novo drug discovery due to reduced time and costs to bring drugs to market. Computational repositioning methods, particularly non-black-box methods that can account for and predict a drugs mechanism, may provide great benefit for directing future development. By tuning both data and algorithm to utilize relationships important to drug mechanisms, a computational repositioning algorithm can be trained to both predict and explain mechanistically novel indications. ResultsIn this work, we examined the 123 curated drug mechanism paths found in the drug mechanism database (DrugMechDB) and after identifying the most important relationships, we integrated 18 data sources to produce a heterogeneous knowledge graph, MechRepoNet, capable of capturing the information in these paths. We applied the Rephetio repurposing algorithm to MechRepoNet using only a subset of relationships known to be mechanistic in nature and found adequate predictive ability on an evaluation set with AUROC value of 0.83. The resulting repurposing model allowed us to prioritize paths in our knowledge graph to produce a predicted treatment mechanism. We found that DrugMechDB paths, when present in the network were rated highly among predicted mechanisms. We then demonstrated MechRepoNets ability to use mechanistic insight to identify a drugs mechanistic target, with a mean reciprocal rank of .525 on a test set of known drug-target interactions. Finally, we walked through a repurposing example of the anti-cancer drug imantinib for use in the treatment of asthma, to demonstrate this methods utility in providing mechanistic insight into repurposing predictions it provides. Availability and implementationThe Python code to reproduce the entirety of this analysis is available at: https://github.com/SuLab/MechRepoNet Contactasu@scripps.edu Supplementary informationSupplemental information is available at Bioinformatics online.

bioinformatics↗

stPlus: a reference-based method for the accurate enhancement of spatial transcriptomics

MotivationSingle-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of tran-scriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the tech-nologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and in-sufficient ability of cell-population identification still impede the applications of these methods. ResultsWe propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial gene expression via a weighted k-NN. stPlus outperforms baseline meth-ods with higher gene-wise and cell-wise Spearman correlation coefficients. We also introduce a clustering-based approach to assess the enhancement performance systematically. Using the data enhanced by stPlus, cell populations can be better identified than using the measured data. The predicted expression of genes unique to scRNA-seq data can also well characterize spatial cell heterogeneity. Besides, stPlus is robust and scalable to datasets of diverse gene detection sensitivity levels, sample sizes, and number of spatially meas-ured genes. We anticipate stPlus will facilitate the analysis of spatial transcriptomics. AvailabilitystPlus with detailed documents is freely accessible at http://health.tsinghua.edu.cn/software/stPlus/ and the source code is openly available on https://github.com/xy-chen16/stPlus. Contactruijiang@tsinghua.edu.cn Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

ChIP-AP, An Integrated ChIP-Seq Analysis Pipeline

ChIP-Seq is a technique used to analyse protein-DNA interactions. The protein-DNA complex is pulled down using a protein antibody, after which sequencing and analysis of the bound DNA fragments is performed. A key bioinformatics analysis step is "peak" calling - identifying regions of enrichment. Benchmarking studies have consistently shown that no optimal peak caller exists. Peak callers have distinct selectivity and specificity characteristics which are often not additive and seldom completely overlap in many scenarios. In the absence of a universal peak caller, we rationalized one ought to utilize multiple peak-callers to 1) gauge peak confidence as determined through detection by multiple algorithms, and 2) more thoroughly survey the protein-bound landscape by capturing peaks not detected by individual peak callers owing to algorithmic limitations and biases. We therefore developed an integrated ChIP-Seq Analysis Pipeline (ChIP-AP) which performs all analysis steps from raw fastq files to final result, and utilizes four commonly used peak callers to more thoroughly and comprehensively analyse datasets. Results are integrated and presented in a single file enabling users to apply selectivity and sensitivity thresholds to select the consensus peak set, the union peak set, or any sub-set in-between to more confidently and comprehensively explore the protein-bound landscape. (https://github.com/JSuryatenggara/ChIP-AP).

bioinformatics↗