bioRxiv ScienceSearch

Biology subjects

Anand, A.

Publications and source records attributed to Anand, A..

6 recordsLinked to original sources

Inference of splicing motifs through visualization of recurrent networks

Neural models have been able to obtain state-of-the-art performances on several genome sequence-based prediction tasks. Such models take only nucleotide sequences as input and learn relevant features on its own. However, extracting the interpretable motifs from the model remains a challenge. This work explores various existing visualization techniques in their ability to infer relevant sequence information learned by a recurrent neural network (RNN) on the task of splice junction identification. The visualization techniques have been modulated to suit the genome sequences as input. The visualizations inspect genomic regions at the level of a single nucleotide as well as a span of consecutive nucleotides. This inspection is performed based on modification of input sequences (perturbation-based) or the embedding space (back-propagation based). We infer features pertaining to both canonical and non-canonical splicing from a single neural model. Results indicate that the visualization techniques produce comparable performance for branchpoint detection. However, in case of canonical donor and acceptor junction motifs, perturbation based visualizations perform better than back-propagation based visualizations and vice-versa for non-canonical motifs.

bioinformatics

Global Adoption of High-Sensitivity Cardiac Troponins and the Universal Definition of Myocardial Infarction

ImportanceThe third Universal Definition of Myocardial Infarction aimed to standardize the approach to the diagnosis and management of myocardial infarction. High-sensitivity cardiac troponin testing was recommended, as these assays have improved precision at low concentrations, but concerns over specificity may have limited implementation.\n\nObjectiveTo determine the global adoption of high-sensitivity cardiac troponin assays and key recommendations from the Universal Definition.\n\nDesign, Setting and ParticipantsGlobal survey of 1,902 medical centers across 23 countries evenly distributed across all five continents. Included respondents were involved in the diagnosis and management of patients with suspected acute coronary syndrome at their institutions.\n\nMain Outcomes and MeasuresStructured questionnaire detailing the primary biomarker used for myocardial infarction, diagnostic thresholds and critical elements of clinical pathways for comparison to the third Universal Definition recommendations.\n\nResultsCardiac troponin was the primary diagnostic biomarker for myocardial infarction at 96% of all sites surveyed. Only 41% of centers had adopted high-sensitivity cardiac troponin assays, with wide variation from 7% in North America to 60% in Europe. Sites using high-sensitivity assays more frequently employed serial sampling pathways (91% vs. 78%) and the 99th percentile diagnostic threshold (74% vs. 66%) when compared to sites using the previous generation of troponin assays. Furthermore, sites using high-sensitivity assays more often used earlier serial sampling ([≤]3 hours) and accelerated diagnostic pathways. However, fewer than 1 in 5 sites using high-sensitivity assays had adopted sex-specific thresholds (18%).\n\nConclusions and RelevanceProgress has been made in adopting the recommendations of the Universal Definition of Myocardial Infarction, particularly in the use of the 99th percentile diagnostic threshold and serial sampling. However, high-sensitivity assays are used in a minority of sites and sex-specific thresholds in even fewer. These findings highlight regions where additional efforts are required to improve the risk stratification and diagnosis of patients with myocardial infarction.

biochemistry

Sequence variation of rare outer membrane protein β-barrel domains in clinical strains provides insights into the evolution of Treponema pallidum subsp. pallidum, the syphilis spirochete.

In recent years, considerable progress has been made in topologically and functionally characterizing integral outer membrane proteins (OMPs) of Treponema pallidum subspecies pallidum (TPA), the syphilis spirochete, and identifying its surface-exposed {beta}-barrel domains. Extracellular loops in OMPs of Gram-negative bacteria are known to be highly variable. We examined the sequence diversity of {beta}-barrel-encoding regions of tprC, tprD, and bamA, in 31 specimens from Cali, Colombia; San Francisco, California; and the Czech Republic and compared them to allelic variants in the 41 reference genomes in the NCBI database. To establish a phylogenetic framework, we used tp0548 genotyping and tp0558 sequences to assign strains to the Nichols or SS14 clades. We found that (i) {beta}-barrels in clinical strains could be grouped according to allelic variants in TPA reference genomes; (ii) for all three OMP loci, clinical strains within the Nichols or SS14 clades often harbored {beta}-barrel variants that differed from the Nichols and SS14 reference strains; and (iii) OMP variable regions often reside in predicted extracellular loops containing B-cell epitopes. Based upon structural models, non-conservative amino acid substitutions in predicted transmembrane {beta}-strands of TprC and TprD2 could give rise to functional differences in their porin channels. OMP profiles of some clinical strains were mosaics of different reference strains and did not correlate with results from enhanced molecular typing. Our observations suggest that human host selection pressures drive TPA OMP diversity and that genetic exchange contributes to the evolutionary biology of TPA. They also set the stage for topology-based analysis of antibody responses against OMPs and help frame strategies for syphilis vaccine development.\n\nIMPORTANCEDespite recent progress characterizing outer membrane proteins (OMPs) of Treponema pallidum (TPA), little is known about how their surface-exposed, {beta}-barrel-forming domains vary among strains circulating within high-risk populations. In this study, sequences for the {beta}-barrel-encoding regions of three OMP loci, tprC, tprD, and bamA, in TPA from a large number of patient specimens from geographically disparate sites were examined. Structural models predict that sequence variation within {beta}-barrel domains occurred predominantly within predicted extracellular loops. Amino acid substitutions in predicted transmembrane strands that could potentially affect porin channel function also were noted. Our findings suggest that selection pressures exerted by human populations drive TPA OMP diversity and that recombination at OMP loci contributes to the evolutionary biology of syphilis spirochetes. These results also set the stage for topology-based analysis of antibody responses that promote clearance of TPA and frame strategies for vaccine development based upon conserved OMP extracellular loops.

microbiology

Rapid Reconstruction of Time-varying Gene Regulatory Networks

--Rapid advancements in high-throughput technologies has resulted in genome-scale time series datasets. Uncovering the temporal sequence of gene regulatory events, in the form of time-varying gene regulatory networks (GRNs), demands computationally fast, accurate and scalable algorithms. The existing algorithms can be divided into two categories: ones that are time-intensive and hence unscalable; others that impose structural constraints to become scalable. In this paper, a novel algorithm, namely an algorithm for reconstructing Time-varying Gene regulatory networks with Shortlisted candidate regulators (TGS), is proposed. TGS is time-efficient and does not impose any structural constraints. Moreover, it provides such flexibility and time-efficiency, without losing its accuracy. TGS consistently outperforms the state-of-the-art algorithms in true positive detection, on three benchmark synthetic datasets. However, TGS does not perform as well in false positive rejection. To mitigate this issue, TGS+ is proposed. TGS+ demonstrates competitive false positive rejection power, while maintaining the superior speed and true positive detection power of TGS. Nevertheless, main memory requirements of both TGS variants grow exponentially with the number of genes, which they tackle by restricting the maximum number of regulators for each gene. Relaxing this restriction remains a challenge as the actual number of regulators is not known a priori.\n\nReproducibilityThe datasets and results can be found at: https://github.com/aaiitg-grp/TGS. This manuscript is currently under review. As soon as it is accepted, the source code will be made available at the same link. There are mentions of a supplementary document throughout the text. The supplementary document will also be made available after acceptance of the manuscript. If you wish to be notified when the supplementary document and source code are available, kindly send an email to saptarshipyne01@gmail.com with subject line TGS Source Code: Request for Notification. The email body can be kept blank.

bioinformatics

Heat shock factor 5 is conserved in vertebrates and essential forspermatogenesis in zebrafish

Heat shock factors (Hsfs) are transcription factors that regulate response to heat shock and to variety of other environmental and physiological stimuli. Four HSFs (HSF1-4) known in vertebrates till date, perform a wide variety of functions from mediating heat shock response to development and gametogenesis. Here, we describe a new yet conserved member of HSF family, Hsf5, which likely exclusively functions for spermatogenesis. The hsf5 is predominantly expressed in developing testicular tissues, in comparison to wider expression reported for other HSFs. HSF5 loss causes male sterility due to drastically reduced sperm count, and severe abnormalities in remaining few spermatozoa. While hsf5 mutant female did not show any abnormality. We show that Hsf5 is required for progression through meiotic prophase 1 during spermatogenesis. The hsf5 mutants indeed show misregulation of a substantial number of genes regulating cell cycle, DNA-damage repair, apoptosis and cytoskeleton proteins. We also show that Hsf5 physically binds to majority of these differentially expressed genes, suggesting its direct role in regulating the expression of many genes important for spermatogenesis.

developmental biology

SpliceVec: distributed feature representations for splice junction prediction

Identification of intron boundaries, called splice junctions, is an important part of delineating gene structure and functions. This also provides valuable in-sights into the role of alternative splicing in increasing functional diversity of genes. Identification of splice junctions through RNA-seq is by mapping short reads to the reference genome which is prone to errors due to random sequence matches. This encourages identification of splicing junctions through computa-tional methods based on machine learning. Existing models are dependent on feature extraction and selection for capturing splicing signals lying in vicinity of splice junctions. But such manually extracted features are not exhaustive. We introduce distributed feature representation, SpliceVec, to avoid explicit and biased feature extraction generally adopted for such tasks. SpliceVec is based on two widely used distributed representation models in natural language processing. Learned feature representation in form of SpliceVec is fed to multi-layer perceptron for splice junction classification task. An intrinsic evaluation of SpliceVec indicates that it is able to group true and false sites distinctly. Our study on optimal context to be considered for feature extraction indicates inclusion of entire intronic sequence to be better than flanking upstream and downstream region around splice junctions. Further, SpliceVec is invariant to canonical and non-canonical splice junction detection. The proposed model is consistent in its performance even with reduced dataset and class-imbalanced dataset. SpliceVec is computationally efficient and can be trained with user defined data as well.

bioinformatics