bioRxiv ScienceSearch

Biology subjects

Jiang, T.

Publications and source records attributed to Jiang, T..

At least 19 recordsLinked to original sources

rMETL: sensitive and fast mobile element insertion detection with long read realignment

SummaryMobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing provides the opportunity to sensitively discover MEIs. However, the signals of MEIs implied by noisy long reads are highly complex, due to the repetitiveness of mobile elements as well as the serious sequencing errors. Herein, we propose Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). rMETL takes advantage of its novel chimeric read re-alignment approach to well handle complex MEI signals. Benchmarking results on simulated and real datasets demonstrated that rMETL has the ability to more sensitivity discover MEIs as well as prevent false positives. It is suited to produce high quality MEI callsets in many genomics studies.\n\nAvailability and Implementation: rMETL is available from https://github.com/hitbc/rMETL.\n\nContact: ydwang@hit.edu.cn\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

Identification of genome-wide nucleotide sites associated with mammalian virulence in influenza A viruses

MotivationThe virulence of influenza viruses is a complex multigenic trait. Previous studies about the virulence determinants of influenza viruses mainly focused on amino acid sites, ignoring the influence of nucleotide mutations.\n\nResultsWe collected more than 200 viral strains from 21 subtypes of influenza A viruses with virulence in mammals and obtained over 100 mammalian virulence-related nucleotide sites across the genome by computational analysis. Interestingly, 50 of these nucleotide sites only experienced synonymous mutations. Further experiments showed that synonymous mutations in the top two of these nucleotide sites, i.e., PB1-2031 and PB1-633, enhanced the pathogenicity of the viruses in mice. Finally, machine-learning models with accepted accuracy for predicting mammalian virulence of influenza A viruses were built. Overall, this study highlighted the importance of nucleotide mutations, especially synonymous mutations in viral virulence, and provided rapid methods for evaluating the virulence of influenza A viruses. It could be helpful for early warning of newly emerging influenza A viruses.

microbiology

Integrative analysis of Zika virus genome RNA structure reveals critical determinants of viral infectivity

Since its outbreak in 2007, Zika virus (ZIKV) has become a global health threat that causes severe neurological conditions. Here we perform a comparative in vivo structural analysis of the RNA genomes of two ZIKV strains to decipher the regulation of their infection at the RNA level. Our analysis identified both known and novel functional RNA structural elements. We discovered a functional long-range intramolecular interaction specific for the Asian epidemic strains, which contributes to their infectivity. Our findings illuminate the structural basis of ZIKV regulation and provide a rich resource for the discovery of RNA structural elements that are important for ZIKV infection.

molecular biology

Bats adjust temporal features of echolocation calls but not those of communication calls in response to traffic noise

Summary statementThis study reveals the impact of anthropogenic noise on spectrally distinct vocalizations and the limitations of the acoustic masking hypothesis to explain the vocal response of bats to chronic noise.\n\nAbstractThe acoustic masking hypothesis states that auditory masking may occur if the target sound and interfering sounds overlap spectrally, and it suggests that animals exposed to noise will modify their acoustic signals to increase signal detectability. However, it is unclear if animals will put more effort into changing their signals that spectrally overlap more with the interfering sounds than when the signals overlap less. We examined the dynamic changes in the temporal features of echolocation and communication vocalizations of the Asian particolored bat (Vespertilio sinensis) when exposed to traffic noise. We hypothesized that traffic noise has a greater impact on communication vocalizations than on echolocation vocalizations and predicted that communication vocalization change would be greater than echolocation. The bats started to adjust echolocation vocalizations on the fourth day of noise exposure, including an increased number of call sequences, decreased number of calls, and vocal rate within a call sequence. However, there was little change in the duration of the call sequence. In contrast, these communication vocalization features were not significantly adjusted under noise conditions. These findings suggest that the degree of spectral overlap between noise and animal acoustic signals does not predict the level of temporal vocal response to the noise.

animal behavior and cognition

Membrane proteins with high N-glycosylation, high expression, and multiple interaction partners were preferred by mammalian viruses as receptors

Receptor mediated entry is the first step for viral infection. However, the relationship between viruses and receptors is still obscure. Here, by manually curating a high-quality database of 268 pairs of mammalian virus-host receptor interaction, which included 128 unique viral species or sub-species and 119 virus receptors, we found the viral receptors were structurally and functionally diverse, yet they had several common features when compared to other cell membrane proteins: more protein domains, higher level of N-glycosylation, higher ratio of self-interaction and more interaction partners, and higher expression in most tissues of the host. Additionally, the receptors used by the same virus tended to co-evolve. Further correlation analysis between viral receptors and the tissue and host specificity of the virus shows that the virus receptor similarity was a significant predictor for mammalian virus cross-species. This work could deepen our understanding towards the viral receptor selection and help evaluate the risk of viral zoonotic diseases.

microbiology

A Highly Efficient and Faithful MDS Patient-Derived Xenotransplantation Model for Pre-Clinical Studies

Comprehensive preclinical studies of Myelodysplastic Syndromes (MDS) have been elusive due to limited ability of MDS stem cells to engraft current immunodeficient murine hosts. We developed a novel MDS patient-derived xenotransplantation model in cytokine-humanized immunodeficient \"MISTRG\" mice that for the first time provides efficient and faithful disease representation across all MDS subtypes. MISTRG MDS patient-derived xenografts (PDX) reproduce patients' dysplastic morphology with multi-lineage representation, including erythro- and megakaryopoiesis. MISTRG MDS-PDX replicate the original sample's genetic complexity and can be propagated via serial transplantation. MISTRG MDS-PDX demonstrate the cytotoxic and differentiation potential of targeted therapeutics providing superior readouts of drug mechanism of action and therapeutic efficacy. Physiologic humanization of the hematopoietic stem cell niche proves critical to MDS stem cell propagation and function in vivo. The MISTRG MDS-PDX model opens novel avenues of research and long-awaited opportunities in MDS research.

cancer biology

NeoDTI: Neural integration of neighborinformation from a heterogeneous network fordiscovering new drug-target interactions

MotivationAccurately predicting drug-target interactions (DTIs) in silico can guide the drug discovery process and thus facilitate drug development. Computational approaches for DTI prediction that adopt the systems biology perspective generally exploit the rationale that the properties of drugs and targets can be characterized by their functional roles in biological networks.\n\nResultsInspired by recent advance of information passing and aggregation techniques that generalize the convolution neural networks (CNNs) to mine large-scale graph data and greatly improve the performance of many network-related prediction tasks, we develop a new nonlinear end-to-end learning model, called NeoDTI, that integrates diverse information from heterogeneous network data and automatically learns topology-preserving representations of drugs and targets to facilitate DTI prediction. The substantial prediction performance improvement over other state-of-the-art DTI prediction methods as well as several novel predicted DTIs with evidence supports from previous studies have demonstrated the superior predictive power of NeoDTI. In addition, NeoDTI is robust against a wide range of choices of hyperparameters and is ready to integrate more drug and target related information (e.g., compound-protein binding affinity data). All these results suggest that NeoDTI can offer a powerful and robust tool for drug development and drug repositioning.\n\nAvailability and implementationThe source code and data used in NeoDTI are available at: https://github.com/FangpingWan/NeoDTI.\n\nContactzengjy321@tsinghua.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

DeepHINT: Understanding HIV-1 integration via deep learning with attention

MotivationHuman immunodeficiency virus type 1 (HIV-1) genome integration is closely related to clinical latency and viral rebound. In addition to human DNA sequences that directly interact with the integration machinery, the selection of HIV integration sites has also been shown to depend on the heterogeneous genomic context around a large region, which greatly hinders the prediction and mechanistic studies of HIV integration.\n\nResultsWe have developed an attention-based deep learning framework, named DeepHINT, to simultaneously provide accurate prediction of HIV integration sites and mechanistic explanations of the detected sites. Extensive tests on a high-density HIV integration site dataset showed that DeepHINT can outperform conventional modeling strategies by automatically learning the genomic context of HIV integration solely from primary DNA sequence information. Systematic analyses on diverse known factors of HIV integration further validated the biological relevance of the prediction result. More importantly, in-depth analyses of the attention values output by DeepHINT revealed intriguing mechanistic implications in the selection of HIV integration sites, including potential roles of several basic helix-loop-helix (bHLH) transcription factors and zinc-finger proteins. These results established DeepHINT as an effective and explainable deep learning framework for the prediction and mechanistic study of HIV integration.\n\nAvailabilityDeepHINT is available as an open-source software and can be downloaded from https://github.com/nonnerdling/DeepHINT\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

Intrinsic functional reorganization of the attention network in the blind

Attention can bias visual perception by modulating the neuronal activity of visual areas. However, little is known if blindness can reshape the intrinsic functional organisation within the attention networks and between the attention and visual networks. A voxel-wise network-based functional connectivity strengthen mapping analysis was proposed to thirty congenitally, thirty early and thirty late blind subjects, and thirty sighted controls. Both the blind and sighted subjects exhibited similar spatial distributions of the intrinsic dorsal (DAN) and ventral (VAN) attention networks. Moreover, compared to the sighted controls, the blind subjects showed increased functional coupling within the DAN, and between the DAN and VAN, and between the attention sub-networks and visual areas, suggesting an increased information communication by visual deprivation. However, the onset age of blindness had little impact on the functional coupling of the attention network, indicating that non-visual sensory experience is enough for driving the development of intrinsic functional organization of the attention network. Finally, a positive correlation was identified between the duration of blindness and the functional coupling of the posterior inferior frontal gyrus with the visual network, representing an experience-dependent reorganisation after visual deprivation.

neuroscience

Reliability of Whole-Exome Sequencing for Assessing Intratumor Genetic Heterogeneity

Multi-region sequencing is used to detect intratumor genetic heterogeneity (ITGH) in tumors. To assess whether genuine ITGH can be distinguished from sequencing artifacts, we whole-exome sequenced (WES) three anatomically distinct regions of the same tumor with technical replicates to estimate technical noise. Somatic variants were detected with three different WES pipelines and subsequently validated by high-depth amplicon sequencing. The cancer-only pipeline was unreliable, with about 69% of the identified somatic variants being false positive. Even with matched normal DNA where 82% of the somatic variants were detected reliably, only 36%-78% were found consistently in technical replicate pairs. Overall 34%-80% of the discordant somatic variants, which could be interpreted as ITGH, were found to constitute technical noise. Excluding mutations affecting low mappability regions or occurring in certain mutational contexts was found to reduce artifacts, yet detection of subclonal mutations by WES in the absence of orthogonal validation remains unreliable.

genomics

Consistency in predicting functions from anatomical and functional connectivity profiles across the cortical cortex

More and more studies had used connectivity profiles to predict functions of the brain. However, whether anatomical connectivity can predict functions consistently with functional connectivity in various functional domains and whether the connectivity-function relationship is universal across the whole cortex are unknown. Using a linear model, we discovered that anatomical connectivity was comparative to functional connectivity in explaining the variance of functions in most cortical regions, with the exception that anatomical connectivity had poor explaining abilities in brain areas which had high individual task variations. In addition, anatomical connectivity were not that good at capturing individual functional differences and had less inter-subject variation than functional connectivity, however anatomical connectivity could be regarded as more stable in the perspective of parcellation. The current results provided the first comprehensive picture of the relationships between functions and connectivity in the whole human cortex at a fine-grained brain atlas.

neuroscience

Very low depth whole genome sequencing in complex trait association studies

MotivationVery low depth sequencing has been proposed as a cost-effective approach to capture low-frequency and rare variation in complex trait association studies. However, a full characterisation of the genotype quality and association power for very low depth sequencing designs is still lacking.\n\nResultsWe perform cohort-wide whole genome sequencing (WGS) at low depth in 1,239 individuals (990 at 1x depth and 249 at 4x depth) from an isolated population, and establish a robust pipeline for calling and imputing very low depth WGS genotypes from standard bioinformatics tools. Using genotyping chip, whole-exome sequencing (WES, 75x depth) and high-depth (22x) WGS data in the same samples, we examine in detail the sensitivity of this approach, and show that imputed 1x WGS recapitulates 95.2% of variants found by imputed GWAS with an average minor allele concordance of 97% for common and low-frequency variants. In our study, 1x further allowed the discovery of 140,844 true low-frequency variants with 73% genotype concordance when compared to high-depth WGS data. Finally, using association results for 57 quantitative traits, we show that very low depth WGS is an efficient alternative to imputed GWAS chip designs, allowing the discovery of up to twice as many true association signals than the classical imputed GWAS design.\n\nSupplementary DataSupplementary Data are appended to this manuscript.

genetics

Consequences Of Natural Perturbations In The Human Plasma Proteome

Proteins are the primary functional units of biology and the direct targets of most drugs, yet there is limited knowledge of the genetic factors determining inter-individual variation in protein levels. Here we reveal the genetic architecture of the human plasma proteome, testing 10.6 million DNA variants against levels of 2,994 proteins in 3,301 individuals. We identify 1,927 genetic associations with 1,478 proteins, a 4-fold increase on existing knowledge, including trans associations for 1,104 proteins. To understand consequences of perturbations in plasma protein levels, we introduce an approach that links naturally occurring genetic variation with biological, disease, and drug databases. We provide insights into pathogenesis by uncovering the molecular effects of disease-associated variants. We identify causal roles for protein biomarkers in disease through Mendelian randomization analysis. Our results reveal new drug targets, opportunities for matching existing drugs with new disease indications, and potential safety concerns for drugs under development.

genomics

Characterizing RNA Pseudouridylation By Convolutional Neural Networks

The most prevalent post-transcriptional RNA modification, pseudouridine ({Psi}), also known as the fifth ribonucleoside, is widespread in rRNAs, tRNAs, snRNAs, snoRNAs and mRNAs. Pseudouridines in RNAs are implicated in many aspects of post-transcriptional regulation, such as the maintenance of translation fidelity, control of RNA stability and stabilization of RNA structure. However, our understanding of the functions, mechanisms as well as precise distribution of pseudourdines (especially in mRNAs) still remains largely unclear. Though thousands of RNA pseudouridylation sites have been identified by high-throughput experimental techniques recently, the landscape of pseudouridines across the whole transcriptome has not yet been fully delineated. In this study, we present a highly effective model, called PULSE (PseudoUridyLation Sites Estimator), to predict novel {Psi} sites from large-scale profiling data of pseudouridines and characterize the contextual sequence features of pseudouridylation. PULSE employs a deep learning framework, called convolutional neural network (CNN), which has been successfully and widely used for sequence pattern discovery in the literature. Our extensive validation tests demonstrated that PULSE can outperform conventional learning models and achieve high prediction accuracy, thus enabling us to further characterize the transcriptome-wide landscape of pseudouridine sites. Overall, PULSE can provide a useful tool to further investigate the functional roles of pseudouridylation in post-transcriptional regulation.

bioinformatics

A powerful approach to estimating annotation-stratified genetic covariance using GWAS summary statistics

Despite the success of large-scale genome-wide association studies (GWASs) on complex traits, our understanding of their genetic architecture is far from complete. Jointly modeling multiple traits genetic profiles has provided insights into the shared genetic basis of many complex traits. However, large-scale inference sets a high bar for both statistical power and biological interpretability. Here we introduce a principled framework to estimate annotation-stratified genetic covariance between traits using GWAS summary statistics. Through theoretical and numerical analyses we demonstrate that our method provides accurate covariance estimates, thus enabling researchers to dissect both the shared and distinct genetic architecture across traits to better understand their etiologies. Among 50 complex traits with publicly accessible GWAS summary statistics (Ntotal {approx} 4.5 million), we identified more than 170 pairs with statistically significant genetic covariance. In particular, we found strong genetic covariance between late-onset Alzheimers disease (LOAD) and amyotrophic lateral sclerosis (ALS), two major neurodegenerative diseases, in single-nucleotide polymorphisms (SNPs) with high minor allele frequencies and in SNPs located in the predicted functional genome. Joint analysis of LOAD, ALS, and other traits highlights LOADs correlation with cognitive traits and hints at an autoimmune component for ALS.

genetics

TIDE: predicting translation initiation sites by deep learning

MotivationTranslation initiation is a key step in the regulation of gene expression. In addition to the annotated translation initiation sites (TISs), the translation process may also start at multiple alternative TISs (including both AUG and non-AUG codons), which makes it challenging to predict TISs and study the underlying regulatory mechanisms. Meanwhile, the advent of several high-throughput sequencing techniques for profiling initiating ribosomes at single-nucleotide resolution, e.g., GTI-seq and QTI-seq, provides abundant data for systematically studying the general principles of translation initiation and the development of computational method for TIS identification.\n\nMethodsWe have developed a deep learning based framework, named TITER, for accurately predicting TISs on a genome-wide scale based on QTI-seq data. TITER extracts the sequence features of translation initiation from the surrounding sequence contexts of TISs using a hybrid neural network and further integrates the prior preference of TIS codon composition into a unified prediction framework.\n\nResultsExtensive tests demonstrated that TITER can greatly outperform the state-of-the-art prediction methods in identifying TISs. In addition, TITER was able to identify important sequence signatures for individual types of TIS codons, including a Kozak-sequence-like motif for AUG start codon. Furthermore, the TITER prediction score can be related to the strength of translation initiation in various biological scenarios, including the repressive effect of the upstream open reading frames (uORFs) on gene expression and the mutational effects influencing translation initiation efficiency.\n\nAvailabilityTITER is available as an open-source software and can be downloaded from https://github.com/zhangsaithu/titer\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

An Adaptive Geometric Search Algorithm for Macromolecular Scaffold Selection

A wide variety of protein and peptidomimetic design tasks require matching functional three-dimensional motifs to potential oligomeric scaffolds. Enzyme design, for example, aims to graft active-site patterns typically consisting of 3 to 15 residues onto new protein surfaces. Identifying suitable proteins capable of scaffolding such active-site engraftment requires costly searches to identify protein folds that can provide the correct positioning of side chains to host the desired active site. Other examples of biodesign tasks that require simpler fast exact geometric searches of potential side chain positioning include mimicking binding hotspots, design of metal binding clusters and the design of modular hydrogen binding networks for specificity. In these applications the speed and scaling of geometric search limits downstream design to small patterns. Here we present an adaptive algorithm to searching for side chain take-off angles compatible with an arbitrarily specified functional pattern that enjoys substantive performance improvements over previous methods. We demonstrate this method in both genetically encoded (protein) and synthetic (peptidomimetic) design scenarios. Examples of using this method with the Rosetta framework for protein design are provided but our implementation is compatible with multiple protein design frameworks and is freely available as a set of python scripts (https://github.com/JiangTian/adaptive-geometric-search-for-protein-design).

bioengineering

A Deep Boosting Based Approach for Capturing the Sequence Binding Preferences of RNA-Binding Proteins from High-Throughput CLIP-Seq Data

Characterizing the binding behaviors of RNA-binding proteins (RBPs) is important for understanding their functional roles in gene expression regulation. However, current high-throughput experimental methods for identifying RBP targets, such as CLIP-seq and RNAcompete, usually suffer from the false positive and false negative issues. Here, we develop a deep boosting based machine learning approach, called DeBooster, to accurately model the binding sequence preferences and identify the corresponding binding targets of RBPs from CLIP-seq data. Comprehensive validation tests have shown that DeBooster can outperform other state-of-the-art approaches in predicting RBP targets and recover false negatives that are common in current CLIP-seq data. In addition, we have demonstrated several new potential applications of DeBooster in understanding the regulatory functions of RBPs, including the binding effects of the RNA helicase MOV10 on mRNA degradation, the influence of different binding behaviors of the ADAR proteins on RNA editing, as well as the antagonizing effect of RBP binding on miRNA repression. Moreover, DeBooster may provide an effective index to investigate the effect of pathogenic mutations in RBP binding sites, especially those related to splicing events. We expect that DeBooster will be widely applied to analyze large-scale CLIP-seq experimental data and can provide a practically useful tool for novel biological discoveries in understanding the regulatory mechanisms of RBPs. The scource code of DeBooster can be downloaded from http://github.com/dongfanghong/deepboost.

bioinformatics