bioRxiv Science⌕ Search

Biology subjects

YAN, Y.

Publications and source records attributed to YAN, Y..

3 recordsLinked to original sources

Predicting microbial transcriptome using genome sequence

We present TXpredict, a transformer-based framework for predicting microbial transcriptomes using annotated genome sequences. By leveraging information learned from a large protein language model, TXpredict achieves an average Spearman correlation of 0.53 and 0.62 in predicting gene expression for new bacterial and fungal genomes. We further extend this framework to predict transcriptomes for 2, 685 additional microbial genomes spanning 1, 744 genera, 82% of which remain uncharacterized at the transcriptional level. Our analysis highlights conserved and divergent transcriptional programs across understudied genera, providing a powerful resource for uncovering microbial adaptation strategies and metabolic potential across the tree of life. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=103 SRC="FIGDIR/small/630741v4_ufig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@d1cad1org.highwire.dtl.DTLVardef@15a9206org.highwire.dtl.DTLVardef@128e65borg.highwire.dtl.DTLVardef@2b7ca7_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A generative deep learning approach for global species distribution prediction

Anthropogenic pressures on biodiversity necessitate efficient and scalable methods to predict global species distributions. Current species distribution models (SDMs) face limitations with large-scale datasets, complex interspecies interactions, and data quality. Here, we introduce EcoVAE, an autoencoder-based generative model that integrates bioclimatic variables with georeferenced occurrences. The model is trained separately for plants, butterflies, and mammals to predict global distributions at both genus and species levels. EcoVAE achieves high precision and speed, outperforming traditional SDMs in spatial block cross-validation. Through unsupervised learning, it captures underlying distribution patterns and reveals species associations that align with known prey-predator relationships. Additionally, it evaluates global sampling efforts and interpolates distributions in data-limited regions, offering new applications for biodiversity exploration and monitoring.

ecology↗

Scaling Logical Density of DNA storage with Enzymatically-Ligated Composite Motifs

DNA is a promising candidate for long-term data storage due to its high density and endurance. The key challenge in DNA storage today is the cost of synthesis. In this work, we propose composite motifs, a frame-work that uses a mixture of prefabricated motifs as building blocks to reduce synthesis cost by scaling logical density. To write data, we introduce Bridge Oligonucleotide Assembly, an enzymatic ligation technique for synthesizing oligos based on composite motifs. To sequence data, we introduce Direct Oligonucleotide Sequencing, a nanopore-based technique to sequence oligos without assembly and amplification. To decode data, we introduce Motif-Search, a novel consensus caller that provides accurate reconstruction despite synthesis and sequencing errors. Using the proposed methods, we present an end-to-end experiment where we store the text "HelloWorld" at a logical density of 84 bits/cycle (14-42x improvement over state-of-the-art.)

synthetic biology↗