bioRxiv ScienceSearch

Biology subjects

Ngo, V.

Publications and source records attributed to Ngo, V..

3 recordsLinked to original sources

Motto: Representing motifs in consensus sequences with minimum information loss

Sequence analysis frequently requires intuitive understanding and convenient representation of motifs. Typically, motifs are represented as position weight matrices (PWMs) and visualized using sequence logos. However, in many scenarios, representing motifs by wildcard-style consensus sequences is compact and sufficient for interpreting the motif information and search for motif match. Based on mutual information theory and Jenson-Shannon Divergence, we propose a mathematical framework to minimize the information loss in converting PWMs to consensus sequences. We name this representation as sequence Motto and have implemented an efficient algorithm with flexible options for converting motif PWMs into Motto from nucleotides, amino acids, and customized alphabets. Here we show that this representation provides a simple and efficient way to identify the binding sites of 1156 common TFs in the human genome. The effectiveness of the method was benchmarked by comparing sequence matches found by Motto with PWM scanning results found by FIMO. On average, our method achieves 0.81 area under the precision-recall curve, significantly (p-value < 0.01) outperforming all existing methods, including maximal positional weight, Douglas and minimal mean square error. We believe this representation provides a distilled summary of a motif, as well as the statistical justification.\n\nAVAILABILITYMotto is freely available at http://wanglab.ucsd.edu/star/motto.

bioinformatics

Identification of DNA motifs that regulate DNA methylation

DNA methylation is an important epigenetic mark but how its locus-specificity is decided in relation to DNA sequence is not fully understood. Here, we have analyzed 34 diverse whole-genome bisulfite sequencing datasets in human and identified 313 motifs, including 92 and 221 associated with methylation (methylation motifs, MMs) and unmethylation (unmethylation motifs, UMs), respectively. The functionality of these motifs is supported by multiple lines of evidences. First, the methylation levels at the MM and UM motifs are respectively higher and lower than the genomic background. Second, these motifs are enriched at the binding sites of methylation modifying enzymes including DNMT3A and TET1, indicating their possible roles of recruiting these enzymes. Third, these motifs significantly overlap with SNPs associated with gene expression and those with DNA methylation. Fourth, disruption of these motifs by SNPs is associated with significantly altered methylation level of the CpGs in the neighbor regions. Furthermore, these motifs together with somatic SNPs are predictive of cancer subtypes and patient survival. We revealed some of these motifs were also associated with histone modifications, suggesting possible interplay between the two types of epigenetic modifications. We also found some motifs form feed forward loops to contribute to DNA methylation dynamics.

bioinformatics

Scalable and Accurate Drug-target Prediction Based on Heterogeneous Bio-linked Network Mining

MotivationDespite the existing classification- and inference-based machine learning methods that show promising results in drug-target prediction, these methods possess inevitable limitations, where: 1) results are often biased as it lacks negative samples in the classification-based methods, and 2) novel drug-target associations with new (or isolated) drugs/targets cannot be explored by inference-based methods. As big data continues to boom, there is a need to study a scalable, robust, and accurate solution that can process large heterogeneous datasets and yield valuable predictions. ResultsWe introduce a drug-target prediction method that improved our previously proposed method from the three aspects: 1) we constructed a heterogeneous network which incorporates 12 repositories and includes 7 types of biomedical entities (#20,119 entities, # 194,296 associations), 2) we enhanced the feature learning method with Node2Vec, a scalable state-of-art feature learning method, 3) we integrate the originally proposed inference-based model with a classification model, which is further fine-tuned by a negative sample selection algorithm. The proposed method shows a better result for drug-target association prediction: 95.3% AUC ROC score compared to the existing methods in the 10-fold cross-validation tests. We studied the biased learning/testing in the network-based pairwise prediction, and conclude a best training strategy. Finally, we conducted a disease specific prediction task based on 20 diseases. New drug-target associations were successfully predicted with AUC ROC in average, 97.2% (validated based on the DrugBank 5.1.0). The experiments showed the reliability of the proposed method in predicting novel drug-target associations for the disease treatment.

bioinformatics