bioRxiv Science⌕ Search

Biology subjects

Mowlaei, M. E.

Publications and source records attributed to Mowlaei, M. E..

2 recordsLinked to original sources

TRUHiC: A TRansformer-embedded U-2 Net to enhance Hi-C data for 3D chromatin structure characterization

High-throughput chromosome conformation capture sequencing (Hi-C) is a key technology for studying the three-dimensional (3D) structure of genomes and chromatin folding. Hi-C data reveals underlying patterns of genome organization, such as topologically associating domains (TADs) and chromatin loops, with critical roles in transcriptional regulation and disease etiology and progression. However, the sparsity of existing Hi-C data often hinders robust and reliable inference of 3D structures. Hence, we propose TRUHiC, a new computational method that leverages recent state-of-the-art deep generative modeling to augment low-resolution Hi-C data for the characterization of 3D chromatin structures. By applying TRUHiC to real low-resolution Hi-C data from the GM12329 cell line and across other publicly available Hi-C data for human and mice, we demonstrate that the augmented data significantly improve the characterization of TADs and loops across diverse cell lines and species. We further present a pre-trained TRUHiC on human lymphoblastoid cell lines that can be adaptable and transferable to improve chromatin characterization of various cell lines, tissues, and species.

bioinformatics↗

Split-Transformer Impute (STI): Genotype Imputation Using a Transformer-Based Model

MotivationDespite recent advances in sequencing technologies, genome-scale datasets continue to have missing bases and genomic segments. Such incomplete datasets can undermine downstream analyses, such as disease risk prediction and association studies. Consequently, the imputation of missing information is a common pre-processing step for which many methodologies have been developed. However, the imputation of genotypes of certain genomic regions and variants, including large structural variants, remains a challenging problem. ResultsHere, we present a transformer-based deep learning framework, called a split-transformer impute (STI) model, for accurate genome-scale genotype imputation. Empowered by the attention-based transformer model, STI can be trained for any collection of genomes automatically using self-supervision. STI handles multi-allelic genotypes naturally, unlike other models that need special treatments. STI models automatically learned genome-wide patterns of linkage disequilibrium (LD), evidenced by much higher imputation accuracy in high LD regions. Also, STI models trained through sporadic masking for self-supervision performed well in imputing systematically missing information. Our imputation results on the human 1000 Genomes Project show that STI can achieve high imputation accuracy, comparable to the state-of-the-art genotype imputation methods, with the additional capability to impute multi-allelic structural variants and other types of genetic variants. Moreover, STI showed excellent performance without needing any special presuppositions about the patterns in the underlying data when applied to a collection of yeast genomes, pointing to easy adaptability and application of STI to impute missing genotypes in any species.

genomics↗