bioRxiv ScienceSearch

Biology subjects

Dutta, A.

Publications and source records attributed to Dutta, A..

6 recordsLinked to original sources

Inference of splicing motifs through visualization of recurrent networks

Neural models have been able to obtain state-of-the-art performances on several genome sequence-based prediction tasks. Such models take only nucleotide sequences as input and learn relevant features on its own. However, extracting the interpretable motifs from the model remains a challenge. This work explores various existing visualization techniques in their ability to infer relevant sequence information learned by a recurrent neural network (RNN) on the task of splice junction identification. The visualization techniques have been modulated to suit the genome sequences as input. The visualizations inspect genomic regions at the level of a single nucleotide as well as a span of consecutive nucleotides. This inspection is performed based on modification of input sequences (perturbation-based) or the embedding space (back-propagation based). We infer features pertaining to both canonical and non-canonical splicing from a single neural model. Results indicate that the visualization techniques produce comparable performance for branchpoint detection. However, in case of canonical donor and acceptor junction motifs, perturbation based visualizations perform better than back-propagation based visualizations and vice-versa for non-canonical motifs.

bioinformatics

A prognostic signature for lower-grade gliomas based on expression of long noncoding RNAs

Diffuse low-grade and intermediate-grade gliomas (together known as lower-grade gliomas, WHO grade II and III) develop in the supporting glial cells of brain and are the most common types of primary brain tumor. Despite a better prognosis for lower-grade gliomas, 70% of patients undergo high-grade transformation within 10 years, stressing the importance of better prognosis. Long non-coding RNAs (lncRNAs) are gaining attention as potential biomarkers for cancer diagnosis and prognosis. We have developed a computational model, UVA8, for prognosis of lower-grade gliomas by combining lncRNA expression, Cox regression and L1-LASSO penalization. The model was trained on a subset of patients in TCGA. Patients in TCGA, as well as a completely independent validation set (CGGA) could be dichotomized based on their risk score, a linear combination of the level of each prognostic lncRNA weighted by its multivariable cox regression coefficient. UVA8 is an independent predictor of survival and outperforms standard epidemiological approaches and previous published lncRNA-based predictors as a survival model. Guilt-by-association studies of the lncRNAs in UVA8, all of which predict good outcome, suggest they have a role in suppressing interferon stimulated response and epithelial to mesenchymal transition. The expression levels of 8 lncRNAs can be combined to produce a prognostic tool applicable to diverse populations of glioma patients. The 8 lncRNA (UVA8) based score can identify grade II and grade III glioma patients with poor outcome and thus identify patients who should receive more aggressive therapy at the outset.

genomics

Topological Scoring of Protein Interaction Networks

It remains a significant challenge to define individual protein associations within networks where an individual protein can directly interact with other proteins and/or be part of large complexes, which contain functional modules. Here we demonstrate the topological scoring (TopS) algorithm for the analysis of quantitative proteomic analyses of affinity purifications. Data is analyzed in a parallel fashion where a bait protein is scored in an individual affinity purification by aggregating information from the entire dataset. A broad range of scores is obtained which indicate the enrichment of an individual protein in every bait protein analyzed. TopS was applied to interaction networks derived from human DNA repair proteins and yeast chromatin remodeling complexes. TopS captured direct protein interactions and modules within complexes. TopS is a rapid method for the efficient and informative computational analysis of datasets, is complementary to existing analysis pipelines, and provides new insights into protein interaction networks.

systems biology

SpliceVec: distributed feature representations for splice junction prediction

Identification of intron boundaries, called splice junctions, is an important part of delineating gene structure and functions. This also provides valuable in-sights into the role of alternative splicing in increasing functional diversity of genes. Identification of splice junctions through RNA-seq is by mapping short reads to the reference genome which is prone to errors due to random sequence matches. This encourages identification of splicing junctions through computa-tional methods based on machine learning. Existing models are dependent on feature extraction and selection for capturing splicing signals lying in vicinity of splice junctions. But such manually extracted features are not exhaustive. We introduce distributed feature representation, SpliceVec, to avoid explicit and biased feature extraction generally adopted for such tasks. SpliceVec is based on two widely used distributed representation models in natural language processing. Learned feature representation in form of SpliceVec is fed to multi-layer perceptron for splice junction classification task. An intrinsic evaluation of SpliceVec indicates that it is able to group true and false sites distinctly. Our study on optimal context to be considered for feature extraction indicates inclusion of entire intronic sequence to be better than flanking upstream and downstream region around splice junctions. Further, SpliceVec is invariant to canonical and non-canonical splice junction detection. The proposed model is consistent in its performance even with reduced dataset and class-imbalanced dataset. SpliceVec is computationally efficient and can be trained with user defined data as well.

bioinformatics

Global Gene Repression By Dicer-Independent tRNA Fragments

tRNA derived RNA fragments (tRFs) is an emerging group of small RNAs as abundant as miRNAs, and yet their roles are not well understood. Here, we focus on endogenous tRFs (18-22 bases) derived from 3 end of human mature tRNAs (tRF-3) and their functions in gene repression. tRF-3 levels increase upon parental tRNA over-expression or tRNA induction by c-Myc oncogene activation. Elevated tRF-3 levels lead to repression of target genes with a sequence complementary to the tRF-3 in the 3 UTR. The tRF-3-mediated repression is Dicer-independent, Argonaute-dependent and the targets are recognized by 5 seed sequence rules similar to miRNAs. Furthermore, tRF-3s associate with GW proteins in P-bodies. RNA-seq identifies the endogenous target genes of tRF-3s that are specifically repressed upon tRF-3 induction. Overall, our analysis shows Dicer-independent tRF-3s, generated upon tRNA upregulation such as c-Myc overexpression, regulate gene expression globally through Argounate via seed sequence matches.

molecular biology

mlh3 separation of function and endonuclease defective mutants display an unexpected effect on meiotic recombination outcomes

Mlh1-Mlh3 is an endonuclease hypothesized to act in meiosis to resolve double Holliday junctions into crossovers. It also plays a minor role in eukaryotic DNA mismatch repair (MMR). To understand how Mlh1-Mlh3 functions in both meiosis and MMR, we analyzed in bakers yeast 60 new mlh3 alleles. Five alleles specifically disrupted MMR, whereas one (mlh3-32) specifically disrupted meiotic crossing over. Mlh1-mlh3 representatives for each separation of function class were purified and characterized. Both Mlh1-mlh3-32 (MMR+, crossover-) and Mlh1-mlh3-45 (MMR-, crossover+) displayed wild-type endonuclease activities in vitro. Msh2-Msh3, an MSH complex that acts with Mlh1-Mlh3 in MMR, stimulated the endonuclease activity of Mlh1-mlh3-32 but not Mlh1-mlh3-45, suggesting that Mlh1-mlh3-45 is defective in MSH interactions. Whole genome recombination maps were constructed for two mlh3 mutants with opposite separation of function phenotypes, and an endonuclease defective mutant. Unexpectedly, all three showed increases in the number of non-crossover events that were not observed in mlh3{Delta}. Our observations provide a structure-function map for Mlh3 that reveals the importance of protein-protein interactions in regulating Mlh1-Mlh3s enzymatic activity. They also illustrate how defective meiotic components can alter the fate of meiotic recombination intermediates, providing new insights for how meiotic recombination pathways are regulated.\n\nAuthor SummaryDuring meiosis, diploid germ cells that become eggs or sperm undergo a single round of DNA replication followed by two consecutive chromosomal divisions. The segregation of chromosomes at the first meiotic division is dependent in most organisms on at least one genetic exchange, or crossover event, between chromosome homologs. Homologs that do not receive a crossover frequently undergo non-disjunction at the first meiotic division, yielding gametes that lack chromosomes or contain additional copies. Such events have been linked to human disease and infertility. Recent studies suggest that the Mlh1-Mlh3 complex is an endonuclease that resolves recombination intermediates into crossovers. Interestingly, this complex also acts as a matchmaker in DNA mismatch repair (MMR) to remove DNA replication errors. How does one complex act in two different processes? We investigated this question by performing a mutational analysis of the bakers yeast Mlh3 protein. Five mutations were identified that disrupted MMR but not crossing over, and one mutation disrupted crossing over while maintaining MMR. Using a combination of biochemical and genetic analyses to further characterize these mutants we illustrate the importance of protein-protein interactions for Mlh1-Mlh3s activity. Importantly, we illustrate how defective meiotic components can alter the outcome of meiotic recombination events. They also provide new insights in our understanding of the basis of infertility syndromes.

genetics