bioRxiv Science⌕ Search

Biology subjects

Sigurdsson, A. I.

Publications and source records attributed to Sigurdsson, A. I..

2 recordsLinked to original sources

Adversarial and variational autoencoders improve metagenomic binning

Assembly of reads from metagenomic samples is a hard problem, often resulting in highly fragmented genome assemblies. Metagenomic binning allows us to reconstruct genomes by regrouping the sequences by their organism of origin, thus representing a crucial processing step when exploring the biological diversity of metagenomic samples. Here we present Adversarial Autoencoders for Metagenomics Binning (AAMB), an ensemble deep learning approach that integrates sequence co-abundances and tetranucleotide frequencies into a common denoised space that enables precise clustering of sequences into microbial genomes. When benchmarked, AAMB presented similar or better results compared with the state-of-the-art reference-free binner VAMB, reconstructing [~]7% more near-complete (NC) genomes across simulated and real data. In addition, genomes reconstructed using AAMB had higher completeness and greater taxonomic diversity compared with VAMB. Finally, we implemented a pipeline integrating VAMB and AAMB that enabled improved binning, recovering 20% and 29% more simulated and real NC genomes, respectively, compared to VAMB with moderate additional runtime. AAMB is freely available at https://github.com/RasmussenLab/VAMB.

bioinformatics↗

Deep integrative models for large-scale human genomics

Polygenic risk scores (PRSs) are expected to play a critical role in achieving precision medicine. Currently, PRS predictors are generally based on linear models using summary statistics, and more recently individual-level data. However, these predictors mainly capture additive relationships and are limited in data modalities they can use. Here, we developed a deep learning framework (EIR) for PRS prediction which includes a model, genome-local-net (GLN), specifically designed for large scale genomics data. The framework supports multi-task (MT) learning, automatic integration of other clinical and biochemical data, and model explainability. When applied to individual level data in the UK Biobank, we found that GLN outperformed LASSO for a wide range of diseases and in particularly autoimmune diseases. Furthermore, we show that this was likely due to modelling epistasis, and we showcase this by identifying widespread epistasis for Type 1 Diabetes. Furthermore, we trained PRS by integrating genotype, blood, urine and anthropometrics and found that this improved performance for 93% of 290 diseases and disorders considered. Finally, we found that including genotype data provided better calibrated PRS models compared to using measurements alone. EIR is available at https://github.com/arnor-sigurdsson/EIR.

bioinformatics↗