bioRxiv Science⌕ Search

Biology subjects

Korsakova, A.

Publications and source records attributed to Korsakova, A..

3 recordsLinked to original sources

Borzoi-informed fine mapping improves causal variant prioritization in complex trait GWAS

1Genome-wide association studies (GWAS) have identified thousands of trait-associated loci. Prioritizing causal variants within these loci is critical for characterizing trait biology. Statistical fine mapping identifies causal variants at trait-associated loci, but linkage disequilibrium (LD) and limited GWAS sample sizes prevent the resolution of many associations. Functionally informed approaches augment fine mapping by estimating variant prior causal probabilities based on overlap with trait-relevant functional annotations. However, functional enrichment provides only an indirect proxy for variant functional impact. Sequence-to-function models directly estimate variant effects on molecular phenotypes from underlying sequence context. Borzoi is a long-context model that predicts sequence determinants of transcription, splicing, and polyadenylation across diverse tissues and cell types. Here we present Sniff, a Borzoi-informed fine-mapping approach that integrates broad genomic functional annotations with Borzoi-predicted variant effects via PolyFun to estimate variant prior causal probabilities. Applied to 15 UK Biobank traits, Sniff identifies 9.45% additional fine-mapped variants compared to PolyFun-Baseline at posterior inclusion probability (PIP) > 0.8. Sniff-prioritized variants exhibit allele-specific activity in reporter assays and are predicted to have tissue-specific activity in trait-relevant tissues. For most traits, genes nominated by Sniff receive higher scores from the orthogonal gene prioritization method PoPS compared to genes nominated using functional annotations alone. Because differentially prioritized variants are driven by Borzoi predictions, we leverage attribution techniques to characterize sequence features underlying fine mapping and generate mechanistic hypotheses for GWAS associations.

genetics↗

Shift augmentation improves DNA convolutional neural network indel effect predictions

Determining genetic variant effects on molecular phenotypes like gene expression is a task of paramount importance to medical genetics. DNA convolutional neural networks (CNNs) attain state-of-the-art performance at predicting variant effects on gene regulation. However, most applications of such models focus on single nucleotide polymorphisms (SNPs), as technical challenges limit their application to insertions and deletions (indels). Sequence shifts from indels introduce technical variance in deep CNNs through misalignment of pooling blocks and output boundaries, creating artificially inflated variant effect scores compared to SNPs and confounding their interpretation. In this work, we demonstrate this technical variance in model predictions and present two strategies based on data augmentation with sequence shifts that reduce it. Applied to the state-of-the-art Borzoi model, our stitching approach improves indel eQTL classification accuracy across GTEx tissues. Furthermore, we demonstrate these techniques and observe compelling eQTL concordance for larger structural variants and tandem repeats. We additionally introduce in silico deletion (ISD) as an interpretation technique and validate it using MPRA data, demonstrating concordance between predicted and experimental measurements for deletion effects. Our strategies expand the utility of regulatory sequence machine learning for studying the full spectrum of noncoding genetic variation in human development and disease.

genomics↗

Prediction of G4 formation in live cells with epigenetic data: a deep learning approach

G-quadruplexes (G4s) are secondary structures abundant in DNA that may play regulatory roles in cells. Despite the ubiquity of the putative G-quadruplex sequences (PQS) in the human genome, only a small fraction forms secondary structures in cells. Folded G4, histone methylation and chromatin accessibility are all parts of the complex cis regulatory landscape. We propose an approach for G4 formation prediction in cells that incorporates epigenetic and chromatin accessibility data. The novel approach termed epiG4NN efficiently predicts cell-specific G4 formation in live cells based on a local epigenomic snapshot. Our architecture confirms the close relationship between H3K4me3 histone methylation, chromatin accessibility and G4 structure formation. Trained on A549 cell data, epiG4NN was then able to predict G4x formation in HEK293T and K562 cell lines. We observe the dependency of model performance with different epigenetic features on the underlying experimental condition of G4 detection. We expect that this approach will contribute to the systematic understanding of correlations between structural and epigenomic feature landscape.

bioinformatics↗