bioRxiv Science⌕ Search

Biology subjects

Yau, S. S.-T.

Publications and source records attributed to Yau, S. S.-T..

4 recordsLinked to original sources

Primitive GLMY Homology: An Algebraic Topology Approach for the Quantitative Characterization of Graph Pangenomes toward Population Genetic Analysis

A central task in population genetics is to identify genetic diversity in a population containing a number of individuals. In recent years, with the development of the third generation sequencing (TGS) technology, pan-genome research has become a hot topic. Although graphical representation has been a popular way to represent the pangenome, few works have attempted to describe it in a more mathematical way. In this paper, we used 79 high-quality assembly data of third-generation sequencing in yeast (includingSaccharomyces cerevisiae and Saccharomyces paradoxus) to construct the graph pangenome, and introduced the Primitive GLMY (Grigoryan-Lin-Muranov-Yau) Homology in algebraic topology to quantitatively represent the pan-genome. We further made an intriguing attempt to conduct a population genetic analysis of this resulting dataset from the topological features of the graph pangenome. We found that there was good agreement between the obtained results and the biological context. We believe this study has developed a method for population genetic analysis of the genetic diversity of genome structural variation.

genetics↗

Using Natural Vector Method for Population Genomic Analysis on Human Mitochondrial Genome Data

The natural vector method is an important method for the analysis of biological sequences. In this study, we applied this method to population genetic analysis, with the core purpose of using it to evaluate the characteristics of a set of sequences rather than just pairwise comparison. We used the mitochondrial genome dataset from the human 1000 Genomes Project as a dataset to verify the feasibility of this improved natural vector method. The results showed that the modified natural vector method could be used for various population genetic approaches at least in the sense of population average, including the calculation of principal component analysis, population structure analysis and genetic diversity parameters. The results were in good agreement with those based on traditional molecular genetic markers such as SNP. The new method validates the feasibility of natural vector method for population genetic analysis and provides a framework for the application of matchless pair method to population genomic analysis on a wider scale.

genetics↗

Evolution is All You Need in Promoter Design and Optimization

Predicting the strength of promoters and guiding their directed evolution is a crucial task in synthetic biology. This approach significantly reduces the experimental costs in conventional promoter engineering. Previous studies employing machine learning or deep learning methods have shown some success in this task, but their outcomes were not satisfactory enough, primarily due to the neglect of evolutionary information. In this paper, we introduce the Chaos-Attention net for Promoter Evolution (CAPE) to address the limitations of existing methods. We comprehensively extract evolutionary information within promoters using chaos game representation and process the overall information with DenseNet and Transformer. Our model achieves state-of-the-art results on two kinds of distinct tasks. The incorporation of evolutionary information enhances the models accuracy, with transfer learning further extending its adaptability. Furthermore, experimental results confirm CAPEs efficacy in simulating in silico directed evolution of promoters, marking a significant advancement in predictive modeling for prokaryotic promoter strength. Our paper also presents a user-friendly website for the practical implementation of in silico directed evolution on promoters.

synthetic biology↗

Grand Biological Universe: Genome space geometry unravels looking for a single metric is likely to be futile in evolution

Understanding the differences between genomic sequences of different lives is crucial for biological classification and phylogeny. Here, we downloaded all the reliable sequences of the seven kingdoms and determined the dimensions of the genome space embedded in the Euclidean space, along with the corresponding Natural Metrics. The concept of the Grand Biological Universe is further proposed. In the grand universe, the convex hulls formed by the universes of seven kingdoms are mutually disjoint, and the convex hulls formed by different biological groups within each kingdom are mutually disjoint. This study provides a novel geometric perspective for studying molecular biology and also offers an accurate way for large-scale sequence comparison in a real-time manner. Most importantly, this study shows that, due to the space-time distortion in the biological genome space similar to Einsteins theory, it is futile to look for a single metric to measure different biological universes, as previous studies have done.

evolutionary biology↗