bioRxiv ScienceSearch

Biology subjects

Pei, J.

Publications and source records attributed to Pei, J..

4 recordsLinked to original sources

Inferring the ancestry of parents and grandparents from genetic data

Inference of admixture proportions is a classical statistical problem in population genetics. Standard methods implicitly assume that both parents of an individual have the same admixture fraction. However, this is rarely the case in real data. In this paper we show that the distribution of admixture tract lengths in a genome contains information about the admixture proportions of the ancestors of an individual. We develop a Hidden Markov Model (HMM) framework for estimating the admixture proportions of the immediate ancestors of an individual, i.e. a type of decomposition of an individuals admixture proportions into further subsets of ancestral proportions in the ancestors. Based on a genealogical model for admixture tracts, we develop an efficient algorithm for computing the sampling probability of the genome from a single individual, as a function of the admixture proportions of the ancestors of this individual. This allows us to perform probabilistic inference of admixture proportions of ancestors only using the genome of an extant individual. We perform extensive simulations to quantify the error in the estimation of ancestral admixture proportions under various conditions. To illustrate the utility of the method, we apply it to real genetic data. Author summaryAncestry inference is an important problem in genetics and is used commercially by a number of companies affecting millions of consumers of genetic ancestry tests. In this paper, we show that it is possible, not only to estimate the ancestry fractions of an individual, but also, with some uncertainty, to estimate the ancestry fractions of an individuals ancestors. For example, if an individual traces his/her ancestry 50% to Asia and 50% to Europe, it is possible to distinguish between the individual having two parents that each are 50:50 composites of Asian and European ancestry, or one parent from Asia and one from Europe. It is likewise also possible to make inferences about grandparents. We present a computationally efficient method for making such inferences called PedMix. PedMix is based on a probabilistic model for the descendant and the recent ancestors. PedMix infers admixture proportions of recent ancestors (parents, grandparents or even great grandparents) using whole-genome genetic variation data from a focal individual. Results on both simulated and real data show that PedMix performs reasonably well in most scenarios.

genetics

Recombinant expression of Proteorhodopsin and biofilm regulators in Escherichia coli for nanoparticle binding and removal in a wastewater treatment model

The small size of nanoparticles is both an advantage and a problem. Their high surface-area-to-volume ratio enables novel medical, industrial, and commercial applications. However, their small size also allows them to evade conventional filtration during water treatment, posing health risks to humans, plants, and aquatic life. This project aims to remove nanoparticles during wastewater treatment using genetically modified Escherichia coli in two ways: 1) binding citrate-capped nanoparticles with the membrane protein Proteorhodopsin, and 2) trapping nanoparticles using Escherichia coli biofilm produced by overexpressing two regulators: OmpR234 and CsgD. We demonstrate experimentally that Escherichia coli expressing Proteorhodopsin binds to 60 nm citrate-capped silver nanoparticles. We also successfully upregulate biofilm production and show that Escherichia coli biofilms are able to trap 30 nm gold particles. Finally, both Proteorhodopsin and biofilm approaches are able to bind and remove nanoparticles in simulated wastewater treatment tanks. We envision integrating our trapping system in both rural and urban wastewater treatment plants to efficiently capture all nanoparticles before treated water is released into the environment.\n\nFinancial DisclosureThis work was funded by the Taipei American School. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.\n\nCompeting InterestsThe authors have declared that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityYes - all data are fully available without restriction. Sequences for the plasmids used in this study are available through the Registry of Standard Biological Parts. Links to raw data are included in Supplementary Information.

synthetic biology

CLADES: A Classification-based Machine Learning Method for Species Delimitation from Population Genetic Data

Species are considered to be the basic unit of ecological and evolutionary studies. Since multi-locus genomic data are becoming increasingly available, there has been considerable interests in the use of DNA sequence data to delimit species. In this paper, we show that machine learning can be used for species delimitation. There exists no species delimitation methods that are based on machine learning. Our method treats the species delimitation problem as a classification problem. It is a problem of identifying the category of a new observation on the basis of training data. Extensive simulation is first conducted over a broad range of evolutionary parameters for training purpose. Each pair of known populations are combined to form training samples with a label of \"same species\" or \"different species\". We use Support Vector Machine (SVM) to train a classifier using a set of summary statistics computed from training samples as features. The trained classifier can classify a test sample to two outcomes: \"same species\" or \"different species\". Given multi-locus genomic data of multiple related organisms or populations, our method (called CLADES) performs species delimitation by first classifying pairs of populations. CLADES then delimits species by maximizing the likelihood of species assignment for multiple populations. CLADES is evaluated through extensive simulation and also tested on real genetic data. We show that CLADES is both accurate and efficient for species delimitation when compared with existing methods. CLADES can be useful especially when existing methods have difficulty in delimitation, e.g. with short species divergence time and gene flow.

evolutionary biology

CircMarker: A Fast and Accurate Algorithm for Circular RNA Detection

While RNA is often created from linear splicing during transcription, recent studies have found that non-canonical splicing sometimes occurs. Non-canonical splicing joins 3 and 5 and forms the socalled circular RNA. It is now believed that circular RNA plays important biological roles such as affecting susceptibility in some diseases. within these few years, several experimental methods have been developed to enrich circular RNA while degrade linear RNA. Although several useful software tools for circRNA detection have been developed as well, these tools may miss many circular RNA. Also, existing tools are slow for large data because those tools often depend on reads mapping. In this paper, we present a new computational approach, named CircMarker, based on k-mers rather than reads mapping for circular RNA detection. CircMarker takes advantage of transcriptome annotation files to create k-mer table for circular RNA detection. Empirical results show that CircMarker outperforms existing tools in circular RNA detection on accuracy and efficiency in many simulated and real datasets. CircMarker can be downloaded from https://github.com/lxwgcool/CircMarker.

bioinformatics