bioRxiv Science⌕ Search

Biology subjects

Ioannidis, A. G.

Publications and source records attributed to Ioannidis, A. G..

9 recordsLinked to original sources

Adversarial Attacks on Genotype Sequences

Adversarial attacks can drastically change the output of a method by performing a small change on its input. While they can be a useful framework to analyze worst-case robustness, they can also be used by malicious agents to perform damage in machine learning-based applications. The proliferation of platforms that allow users to share their DNA sequences and phenotype information to enable association studies has led to an increase in large databases. Such open platforms are, however, vulnerable to malicious users uploading corrupted genetic sequence files that could damage downstream studies. Such studies commonly include steps involving the analysis of the genomic sequences structure using dimensionality reduction techniques and ancestry inference methods. In this paper we show how white-box gradient-based adversarial attacks can be used to corrupt the output of genomic analyses, and we explore different machine learning techniques to detect such manipulations.

bioinformatics↗

Nationwide genomic biobank in Mexico unravels demographic history and complex trait architecture from 6,057 individuals

Latin America continues to be severely underrepresented in genomics research, and fine-scale genetic histories as well as complex trait architectures remain hidden due to the lack of Big Data. To fill this gap, the Mexican Biobank project genotyped 1.8 million markers in 6,057 individuals from 32 states and 898 sampling localities across Mexico with linked complex trait and disease information creating a valuable nationwide genotype-phenotype database. Through a suite of state-of-the-art methods for ancestry deconvolution and inference of identity-by-descent (IBD) segments, we inferred detailed ancestral histories for the last 200 generations in different Mesoamerican regions, unraveling native and colonial/post-colonial demographic dynamics. We observed large variations in runs of homozygosity (ROH) among genomic regions with different ancestral origins reflecting their demographic histories, which also affect the distribution of rare deleterious variants across Mexico. We analyzed a range of biomedical complex traits and identified significant genetic and environmental factors explaining their variation, such as ROH found to be significant predictors for trait variation in BMI and triglycerides.

genomics↗

Predicting Dog Phenotypes from Genotypes

We analyze dog genotypes (i.e., positions of dog DNA sequences that often vary between different dogs) in order to predict the corresponding phenotypes (i.e., unique observed characteristics). More specifically, given chromosome data from a dog, we aim to predict the breed, height, and weight. We explore a variety of linear and non-linear classification and regression techniques to accomplish these three tasks. We also investigate the use of a neural network (both in linear and non-linear modes) for breed classification and compare the performance to traditional statistical methods. We show that linear methods generally outperform or match the performance of non-linear methods for breed classification. However, we show that the reverse is true for height and weight regression. Finally, we evaluate the results of all of these methods based on the number of input features used in the analysis. We conduct experiments using different fractions of the full genomic sequences, resulting in input sequences ranging from 20 SNPs to [~]200k SNPs. In doing so, we explore the impact of using a very limited number of SNPs for prediction. Our experiments demonstrate that these phenotypes in dogs can be predicted with as few as 0.5% of randomly selected SNPs (i.e., 992 SNPs) and that dog breeds can be classified with 50% balanced accuracy with as few as 0.02% SNPs (i.e., 40 SNPs).

genomics↗

Generative Moment Matching Networks for Genotype Simulation

The generation of synthetic genomic sequences using neural networks has potential to ameliorate privacy and data sharing concerns and to mitigate potential bias within datasets due to under-representation of some population groups. However, there is not a consensus on which architectures, training procedures, and evaluation metrics should be used when simulating single nucleotide polymorphism (SNP) sequences with neural networks. In this paper, we explore the use of Generative Moment Matching Networks (GMMNs) for SNP simulation, we present some architectural and procedural changes to properly train the networks, and we introduce an evaluation scheme to qualitatively and quantitatively assess the quality of the simulated sequences.

bioinformatics↗

Hyperbolic Geometry-Based Deep Learning Methods to Produce Population Trees from Genotype Data

The production of population-level trees using the genomic data of individuals is a fundamental task in the field of population genetics. Typically, these trees are produced using methods like hierarchical clustering, neighbor joining, or maximum likelihood. However, such methods are non-parametric: they require all data to be present at the time of tree formation, and the addition of new data points necessitates the regeneration of the entire tree, a potentially expensive process. They also do not easily integrate with larger workflows. In this study, we aim to address these problems by introducing parametric deep learning methods for tree formation from genotype data. Our models specifically create continuous representations of population trees in hyperbolic space, which has previously proven highly effective in embedding hierarchically structured data. We present two different architectures - a multi-layer perceptron (MLP) and a variational autoencoder (VAE) - and we analyze their performance using a variety of metrics along with comparisons to established tree-building methods. Both models tested produce embedding spaces that reflect human evolutionary history. In addition, we demonstrate the generalizability of these models by verifying that addition of new samples to an existing tree occurs in a semantically meaningful manner. Finally, we use Dasguptas Cost to compare the quality of trees generated by our models to those produced by established methods. Despite the fact that the benchmark methods are directly fit on the evaluation data, our models are able to outperform some of these and achieve highly comparable performance overall. Author summaryTree production is a vital task in population genetics, but current approaches fall prey to several common shortfalls. Most notably, they lack the ability to add new data points after tree generation, and they are often difficult to use in larger pipelines. By leveraging cutting-edge advances pairing deep learning with hyperbolic geometry, we develop multiple models designed to rectify these issues. Through experiments on a dataset of humans from globally widespread ancestries, we demonstrate the generalizability of our models to new data, and we also show strong empirical performance with respect to currently used methods. In addition, we show that the data representations produced by our models are semantically meaningful and reflect known facts about human evolutionary history. Finally, we discuss the additional benefits our models could provide, including improved visualization, greater privacy preservation, and improved integration with downstream machine learning tasks. In conclusion, we present models that are accurate, flexible, and generalizable, with the potential to facilitate a variety of further applications.

evolutionary biology↗

Archetypal Analysis for Population Genetics

The estimation of genetic clusters using genomic data has application from genome-wide association studies (GWAS) to demographic history to polygenic risk scores (PRS) and is expected to play an important role in the analyses of increasingly diverse, large-scale cohorts. However, existing methods are computationally-intensive, prohibitively so in the case of nationwide biobanks. Here we explore Archetypal Analysis as an efficient, unsupervised approach for identifying genetic clusters and for associating individuals with them. Such unsupervised approaches help avoid conflating socially constructed ethnic labels with genetic clusters by eliminating the need for exogenous training labels. We show that Archetypal Analysis yields similar cluster structure to existing unsupervised methods such as ADMIXTURE and provides interpretative advantages. More importantly, we show that since Archetypal Analysis can be used with lower-dimensional representations of genetic data, significant reductions in computational time and memory requirements are possible. When Archetypal Analysis is run in this fashion, it takes several orders of magnitude less compute time than the current standard, ADMIXTURE. Finally, we demonstrate uses ranging across datasets from humans to canids. Author summaryThis work introduces a method that combines the singular value decomposition (SVD) with Archetypal Analysis to perform fast and accurate genetic clustering by first reducing the dimensionality of the space of genomic sequences. Each sequence is described as a convex combination (admixture) of archetypes (cluster representatives) in the reduced dimensional space. We compare this interpretable approach to the widely used genetic clustering algorithm, ADMIXTURE, and show that, without significant degradation in performance, Archetypal Analysis outperforms, offering shorter run times and representational advantages. We include theoretical, qualitative, and quantitative comparisons between both methods.

bioinformatics↗

Admixture dynamics in colonial Mexico and the genetic legacy of the Manila Galleon

Mexico has considerable population substructure due to pre-Columbian diversity and subsequent variation in admixture levels from trans-oceanic migrations, primarily from Europe and Africa, but also, to a lesser extent, from Asia. Detailed analyses exploring sub-continental structure remain limited and post-Columbian demographic dynamics within Mexico have not been inferred with genomic data. We analyze the distribution of ancestry tracts to infer the timing and number of pulses of admixture in ten regions across Mexico, observing older admixture timings in the first colonial cities and more recent timings moving outward into southern and southeastern Mexico. We characterize the specific origin of the heterogeneous Native American ancestry in Mexico: a widespread western-central Native Mesoamerican component in northern Aridoamerican states and a central-eastern Nahua contribution in Guerrero (southern Mexico) and Veracruz to its north. Yucatan shows lowland Mayan ancestry, while Sonora exhibits a unique northwestern native Mexican ancestry matching no sampled reference, each consistent with localized indigenous cultures. Finally, in Acapulco, Guerrero a notable proportion of East Asian ancestry was observed, an understudied heritage in Mexico. We identified the source of this ancestry within Southeast Asia--specifically western Indonesian and non-Negrito Filipino--and dated its arrival to approximately thirteen generations ago (1620 CE). This points to a genetic legacy from the 17th century Manila Galleon trade between the colonial Spanish Philippines and the Pacific port of Acapulco in Spanish Mexico. Although this piece of the colonial Spanish trade route from China to Europe appears in historical records, it has been largely ignored as a source of genetic ancestry in Mexico, neglected due to slavery, assimilation as "Indios" and incomplete historical records.

genetics↗

High Resolution Ancestry Deconvolution for Next Generation Genomic Data

As genome-wide association studies and genetic risk prediction models are extended to globally diverse and admixed cohorts, ancestry deconvolution has become an increasingly important tool. Also known as local ancestry inference (LAI), this technique identifies the ancestry of each region of an individuals genome, thus permitting downstream analyses to account for genetic effects that vary between ancestries. Since existing LAI methods were developed before the rise of massive, whole genome biobanks, they are computationally burdened by these large next generation datasets. Current LAI algorithms also fail to harness the potential of whole genome sequences, falling well short of the accuracy that such high variant densities can enable. Here we introduce Gnomix, a set of algorithms that address each of these points, achieving higher accuracy and swifter computational performance than any existing LAI method, while also enabling portable models that are particularly useful when training data are not shareable due to privacy or other restrictions. We demonstrate Gnomix (and its swift phase correction counterpart Gnofix) on worldwide whole-genome data from both humans and canids and utilize its high resolution accuracy to identify the location of ancient New World haplotypes in the Xoloitzcuintle, dating back over 100 generations. Code is available at https://github.com/AI-sandbox/gnomix.

genomics↗

Neural ADMIXTURE: rapid population clustering with autoencoders

Characterizing the genetic substructure of large cohorts has become increasingly important as genetic association and prediction studies are extended to massive, increasingly diverse, biobanks. ADMIXTURE and STRUCTURE are widely used unsupervised clustering algorithms for characterizing such ancestral genetic structure. These methods decompose individual genomes into fractional cluster assignments with each cluster representing a vector of DNA marker frequencies. The assignments, and clusters, provide an interpretable representation for geneticists to describe population substructure at the sample level. However, with the rapidly increasing size of population biobanks and the growing numbers of variants genotyped (or sequenced) per sample, such traditional methods become computationally intractable. Furthermore, multiple runs with different hyperparameters are required to properly depict the population clustering using these traditional methods, increasing the computational burden. This can lead to days of compute. In this work we present Neural ADMIXTURE, a neural network autoencoder that follows the same modeling assumptions as ADMIXTURE, providing similar (or better) clustering, while reducing the compute time by orders of magnitude. Indeed, the equivalent of one month of continuous compute can be reduced to hours. In addition, Neural ADMIXTURE can include multiple outputs, providing the equivalent results as running the original ADMIXTURE algorithm many times with different numbers of clusters. Our models can also be stored, allowing later cluster assignment to be performed with a linear computational time. The software implementation of Neural ADMIXTURE can be found at https://github.com/ai-sandbox/neural-admixture.

genomics↗