bioRxiv Science⌕ Search

Biology subjects

WANG, R.

Publications and source records attributed to WANG, R..

6 recordsLinked to original sources

Probabilistic tensor decomposition extracts better latent embeddings from single-cell multiomic data

Single-cell sequencing technology enables the simultaneous capture of multiomic data from multiple cells. The captured data can be represented by tensors, i.e., the higher-rank matrices. However, the proposed analysis tools often take the data as a collection of two-order matrices, renouncing the correspondences among the features. Consequently, we propose a probabilistic tensor decomposition framework, SCOIT, to extract embeddings from single-cell multiomic data. To deal with sparse, noisy, and heterogeneous single-cell data, we incorporate various distributions in SCOIT, including Gaussian, Poisson, and negative binomial distributions. Our framework can decompose a multiomic tensor into a cell embedding matrix, a gene embedding matrix, and an omic embedding matrix, allowing for various downstream analyses. We applied SCOIT to seven single-cell multiomic datasets from different sequencing protocols. With cell embeddings, SCOIT achieves superior performance for cell clustering compared to seven state-of-the-art tools under various metrics, demonstrating its ability to dissect cellular heterogeneity. With the gene embeddings, SCOIT enables cross-omics gene expression analysis and integrative gene regulatory network study. Furthermore, the embeddings allow cross-omics imputation simultaneously, outperforming conventional imputation methods with the Pearson correlation coefficient increased by 0.03-0.28.

bioinformatics↗

A graph representation of gapped patterns in phage sequences for graph convolutional network

Genome sequencing technologies reveal a huge amount of genomic sequences. Neural network-based methods can be prime candidates for retrieving insights from these sequences because of their applicability to large and diverse datasets.However, the highly variable lengths of nucleic acid sequences severely impair the presentation of sequences as input to the neural network. Genetic variations further complicate tasks that involve sequence comparison or alignment. Here, we propose a graph representation of nucleic acid sequences called gapped pattern graphs. These graphs can be transformed through a Graph Convolutional Network to form lower-dimensional embeddings for downstream tasks. On the basis of the gapped pattern graphs, we implemented a neural network model and demonstrated its performance in studying phage sequences. We compared our model with equivalent models based on other forms of input in performing four tasks related to nucleic acid sequences--phage and ICE discrimination, phage integration site prediction, lifestyle prediction, and host prediction. Other state-of-the-art tools were also compared, where available. Our method consistently outperformed all the other methods in various metrics on all four tasks. In addition, our model was able to identify distinct gapped pattern signatures from the sequences.

bioinformatics↗

Resolving single-cell copy number profiling for large datasets

The advances of single-cell DNA sequencing (scDNA-seq) enable us to characterize the genetic heterogeneity of cancer cells. However, the high noise and low coverage of scDNA-seq impede the estimation of copy number variations (CNVs). In addition, existing tools suffer from intensive execution time and often fail on large datasets. Here, we propose SeCNV, a novel method that leverages structural entropy, to profile the copy numbers. SeCNV adopts a local Gaussian kernel to construct a matrix, depth congruent map, capturing the similarities between any two bins along the genome. Then SeCNV partitions the genome into segments by minimizing the structural entropy from the depth congruent map. With the partition, SeCNV estimates the copy numbers within each segment for cells. We simulate nine datasets with various breakpoint distributions and amplitudes of noise to benchmark SeCNV. SeCNV achieves a robust performance, i.e., the F1-scores are higher than 0.95 for breakpoint detections, significantly outperforming state-of-the-art methods. SeCNV successfully processes large datasets (>50,000 cells) within four minutes while other tools failed to finish within the time limit, i.e., 120 hours. We apply SeCNV to single-nucleus sequencing (SNS) datasets from two breast cancer patients and acoustic cell tagmentation (ACT) sequencing datasets from eight breast cancer patients. SeCNV successfully reproduces the distinct subclones and infers tumor heterogeneity. SeCNV is available at https://github.com/deepomicslab/SeCNV.

bioinformatics↗

A distinctive neural nexus in blind individuals supports Braille reading

Natural Braille reading, a demanding cognitive skill, poses a huge challenge for the brain network of the blind. Here, with behavioral measurement and functional MRI imaging data, we pinpointed the neural pathway and investigated the neural mechanisms of individual differences in Braille reading in late blindness. Using resting state fMRI, we identified a distinct neural link between the higher-tier visual cortex--the lateral occipital cortex (LOC), and the inferior frontal cortex (IFC) in the late blind brain, which is significantly stronger than sighted controls. Individual Braille reading proficiency positively correlated with the left-lateralized LOC-IFC functional connectivity. In a natural Braille reading task, we found an enhanced bidirectional information flow with a stronger top-down modulation of the IFC-to-LOC effective connectivity. Greater top-down modulation contributed to higher Braille reading proficiency via a broader area of task-engaged LOC. Together, we established a model to predict Braille reading proficiency, considering both functional and effective connectivity of the LOC-IFC pathway. This two-tale model suggests that developing the underpinning neural circuit and the top-down cognitive strategy contributes uniquely to superior Braille reading performance. SIGNIFICANCE STATEMENTFor late blind humans, one of the most challenging cognitive skills is natural Braille reading. However, little is known about the neural mechanisms of significant differences in individual Braille reading performance. Using functional imaging data, we identified a distinct neural link between the left lateral occipital cortex (LOC) and the left inferior frontal cortex (IFC) for natural Braille reading in the late blind brain. To better predict individual Braille reading proficiency, we proposed a linear model with two variables of the LOC-IFC link: the resting-state functional connectivity and the task-engaged top-down effective connectivity. These findings suggest that developing the underpinning neural pathway and the top-down cognitive strategy contributes uniquely to superior Braille reading performance.

neuroscience↗

Adipocyte-specific deletion of the oxygen-sensor PHD2 sustains elevated energy expenditure at thermoneutrality

Enhancing brown adipose tissue (BAT) function to combat metabolic disease is a promising therapeutic strategy. A major obstacle to this strategy is that a thermoneutral environment, relevant to most modern human living conditions, deactivates functional BAT. We showed that we can overcome the dormancy of BAT at thermoneutrality by inhibiting the main oxygen sensor HIF-prolyl hydroxylase, PHD2, specifically in adipocytes. Mice lacking adipocyte PHD2 (P2KOad) and housed at thermoneutrality maintained greater BAT mass, had detectable UCP1 protein expression in BAT and higher energy expenditure. Mouse brown adipocytes treated with the pan-PHD inhibitor, FG2216, exhibited higher Ucp1 mRNA and protein levels, effects that were abolished by antagonising the canonical PHD2 substrate, HIF-2a. Induction of UCP1 mRNA expression by FG2216, was also confirmed in human adipocytes isolated from obese individuals. Human serum proteomics analysis of 5457 participants in the deeply phenotyped Age, Gene and Environment Study revealed that serum PHD2 (aka EGLN1) associates with increased risk of metabolic disease. Our data suggest adipose-selective PHD2 inhibition as a novel therapeutic strategy for metabolic disease and identify serum PHD2 as a potential biomarker.

physiology↗

First Detection of Pathogenic Escherichia coli Isolates Associated With Donkey Foals' Diarrhea in Northern China

ObjectiveThe aim of this study was to identify the biological features, influence factor and Genome-wide properties of pathogenic donkey Escherichia coli (DEC) isolates associated with severe diarrhea in Northern China. MethodsThe isolation and identification of DEC isolates were carried out by the conventional isolationautomatic biochemical analysis systemserotype identification16S rRNA testanimal challenge and antibiotics sensitivity examination. The main virulence factors were identified by PCR. The complete genomic re-sequence and frame-sequence were analyzed. Results216 strains of DEC were isolated from diarrhea samples, conforming to the bacterial morphology and biochemical characteristics of E.coli. The average size of the pure culture was 329.4 nmx223.5 nm. Agglutination test showed that O78 (117/179, 65.4%) was the dominant serotype and ETEC(130/216, 60.1%) was the dominant pathogenic type. Noticeable pathogenic were observed in 9 of 10 (90%) randomly selected DEC isolates caused the death of test mice (100%, 5/5) within 6h[~]48h, 1 of 10 (10%) isolates caused the death of test mice (40%, 2/5) within 72h. Our data confirmed that DEC plays an etiology role in dirarrea/death case of donkey foal. Antibiotics sensitivity test showed significant susceptibility to DEC isolates were concentrated in NorEFTENRCIP and AMK,while the isolates with severe antibiotic resistance was AMTEAPRFFCRL and CN. Multi-drug resistance was also observed. A total of 15 virulence gene fragments were determined from DEC(n=30) including OMPA (73%), safD (77%), traTa (73%), STa(67%), EAST1 (67%), astA (63%), kspII (60%), irp2 (73%), iucD (57%), eaeA (57%), VAT (47%), iss (33%), cva (27%), ETT2 (73%) and K88 (60%) respectively. More than 10 virulence genes from 9 of 30(30%) DEC strains were detected, while 6 of 30(20%) DEC strains detected 6 virulence factors. phylogenetic evolutionary tree of 16S rRNA gene from different isolates shows some variability. The original data volume obtained from the genome re-sequencing of DEC La18 was 2.55G and Genome framework sequencing was carried out to demonstrate the predicted functions and evolutionary direction and genetic relationships with other animal E.coli. ConclusionsThese findings provide firstly fundamental data that might be useful in further study of the role of DEC and provide a new understanding of the hazards of traditional colibacillosis due to the appear of new production models.

pathology↗