bioRxiv ScienceSearch

Biology subjects

Yang-Yu Liu

Publications and source records attributed to Yang-Yu Liu.

3 recordsLinked to original sources

Revealing complex ecological dynamics via symbolic regression

Complex ecosystems, from food webs to our gut microbiota, are essential to human life. Understanding the dynamics of those ecosystems can help us better maintain or control them. Yet, reverse-engineering complex ecosystems (i.e., extracting their dynamic models) directly from measured temporal data has not been very successful so far. Here we propose to close this gap via symbolic regression. We validate our method using both synthetic and real data. We firstly show this method allows reverse engineering two-species ecosystems, inferring both the structure and the parameters of ordinary differential equation models that reveal the mechanisms behind the system dynamics. We find that as the size of the ecosystem increases or the complexity of the inter-species interactions grow, using a dictionary of known functional responses (either previously reported or reverse-engineered from small ecosystems using symbolic regression) opens the door to correctly reverse-engineer large ecosystems.

Bioinformatics

Pitfalls in Inferring Human Microbial Dynamics from Temporal Metagenomics Data

Human gut microbiota is a very complex and dynamic ecosystem that plays a crucial role in our health and well-being. Inferring microbial community structure and dynamics directly from time-resolved metagenomics data is key to understanding the community ecology and predicting its temporal behavior. Many methods have been proposed to perform the inference. Yet, we point out that there are several pitfalls along the way, from uninformative temporal measurements to the compositional nature of the relative abundance data, focusing on highly abundant species by ignoring or grouping low-abundance species, and implicit assumptions in various regularization methods. These issues have to be seriously considered in ecological modeling of human gut microbiota.

Bioinformatics

A Novel Clustering Method for Patient Stratification

Patient stratification or disease subtyping is crucial for precision medicine and personalized treatment of complex diseases. The increasing availability of high-throughput molecular data provides a great opportunity for patient stratification. In particular, many clustering methods have been employed to tackle this problem in a purely data-driven manner. Yet, existing methods leveraging high-throughput molecular data often suffers from various limitations, e.g., noise, data heterogeneity, high dimensionality or poor interpretability. Here we introduced an Entropy-based Consensus Clustering (ECC) method that overcomes those limitations all together. Our ECC method employs an entropy-based utility function to fuse many basic partitions to a consensus one that agrees with the basic ones as much as possible. Maximizing the utility function in ECC has a much more meaningful interpretation than any other consensus clustering methods. Moreover, we exactly map the complex utility maximization problem to the classic K-means clustering problem with a modified distance function, which can then be efficiently solved with linear time and space complexity. Our ECC method can also naturally integrate multiple molecular data types measured from the same set of subjects, and easily handle missing values without any imputation. We applied ECC to both synthetic and real data, including 35 cancer gene expression benchmark datasets and 13 cancer types with four molecular data types from The Cancer Genome Atlas. We found that ECC shows superior performance against existing clustering methods. Our results clearly demonstrate the power of ECC in clinically relevant patient stratification.

Bioinformatics