bioRxiv Science⌕ Search

Biology subjects

Papastathopoulos-Katsaros, A.

Publications and source records attributed to Papastathopoulos-Katsaros, A..

3 recordsLinked to original sources

Vision-Based Genomic Model for Copy Number Variant Pathogenicity Prediction

Copy number variants (CNVs) are a major class of structural genomic alterations underlying rare disease, including neurodevelopmental delay and intellectual disability, yet predicting their pathogenicity remains challenging. Existing methods reduce CNVs to region-level numerical features, discarding the positional structure and cross-track patterns that expert clinical reviewers use to interpret genomic evidence. To address this, we introduce TO_SCPLOWESSERACTC_SCPLOW for CNV, a track-based spatial representation for CNV pathogenicity prediction, which represents each variant as a base-pair-resolution multi-track image and models spatial genomic patterns across annotation tracks while preserving positional structure and cross-track dependencies. Trained on a chromosome-level hold-out split of the ClinVar dataset, TO_SCPLOWESSERACTC_SCPLOW outperforms prior methods on held-out and curated noncoding benchmarks, improving AUROC by up to 0.10 over the state-of-the-art baseline. On the independent DECIPHER cohort, the model demonstrates generalizability by maintaining the highest AUROC and the highest F1 score across baselines. Furthermore, our model localizes pathogenic signals to clinically meaningful genomic subregions, providing track-annotated evidence that supports practical clinical interpretation.

bioinformatics↗

A massively parallel reporter assay of MECP2 cis-regulatory elements reveals genetic candidates for male-biased autism

Autism affects males four times more often than females, yet the basis of this sex bias remains unclear. One hypothesis is that hypomorphic variants in X-linked genes--genes where loss-of-function alleles cause syndromic neurodevelopmental disorders (NDDs) predominantly in females--produce milder, non-syndromic phenotypes in hemizygous males. We tested this by investigating cis-regulatory elements (CREs) of MECP2, a dosage-sensitive X-linked gene. Using a massively parallel reporter assay in human neurons, we mapped transcription factor binding sites within MECP2 CREs and tested autism-associated variants for functional impact. We identified two noncoding variants that change CRE activity, each with a male-biased phenotype. One of these, a promoter variant, disrupts NFY binding and reduces MECP2 expression by [~]30%, a magnitude that produces autism-like phenotypes in mice. These findings suggest noncoding MECP2 variants can cause non-syndromic, male-biased autism, and provide a framework for uncovering regulatory variants in other X-linked NDD genes that may contribute to autisms missing heritability.

genetics↗

Positional frequency chaos game representation for machine learning-based classification of crop lncRNAs

Alignment-based methods are fundamental for sequence comparison but are often computationally prohibitive for large-scale genomic analyses. This limitation has spurred the development of quicker, alignment-free alternatives, such as k-mer analysis, which are crucial for studying long noncoding ribonucleic acids (lncRNAs) in plants. These lncRNAs play critical roles in regulating gene expression at both the epigenetic and transcriptomic levels. However, existing alignmentfree approaches typically lose positional information, which can be vital for achieving accurate classification. We propose positional frequency chaos game representation (PFCGR), a novel encoding that improves the traditional frequency chaos game representation (FCGR) by incorporating four statistical moments of k-mer positions: mean, standard deviation, skewness, and kurtosis. This creates a multi-channel image representation of genomic sequences, enabling machine learning models such as Logistic Regression, Random Forests, and Convolutional Neural Networks to classify plant lncRNAs directly from raw genomic sequences. Tested on seven major crop species, our PFCGR-based classifiers achieve classification accuracies comparable to or exceeding those of the computationally intensive DNABERT-based model [1], while requiring 80% to 95% less computational time. These results demonstrate PFCGRs potential as an efficient and accurate tool for plant lncRNA identification, as well as its ability to facilitate large-scale computational studies in genomics.

bioinformatics↗