bioRxiv Science⌕ Search

Biology subjects

Majewski, M. F.

Publications and source records attributed to Majewski, M. F..

4 recordsLinked to original sources

Leveraging Human Pangenome for Improved Somatic Variant Detection

Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.

genomics↗

Maternal SETDB1 enables development beyond cleavage stages by extinguishing the MERVL-driven 2-cell totipotency transcriptional program in the mouse embryo

Loss of maternal SETDB1, a histone H3K9 methyltransferase, leads to developmental arrest prior to implantation, with very few mouse embryos advancing beyond the 8-cell stage, which is currently unexplained. We genetically investigate SETDB1s role in the epigenetic control of the transition from totipotency to pluripotency--a process demanding precise timing and forward directionality. Through single-embryo total RNA sequencing of 2-cell and 8-cell embryos, we find that Setdb1mat-/+ embryos fail to extinguish 1-cell and 2-cell transient genes--alongside persistent expression of MERVL retroelements and MERVL-driven chimeric transcripts that define the totipotent state in mouse 2-cell embryos. Comparative bioinformatics reveals that SETDB1 acts at MT2 LTRs and MERVL-driven chimeric transcripts, which normally acquire H3K9me3 during early development. The dysregulated targets substantially overlap with DUXBL-responsive genes, indicating a shared regulatory pathway for silencing the 2-cell transcriptional program. We establish maternal SETDB1 as a critical chromatin regulator required to extinguish retroelement-driven totipotency networks and ensure successful preimplantation development.

genetics↗

A Pangenomic Method for Establishing a Somatic Variant Detection Resource in HapMap Mixtures

Somatic mosaicism is essential in human biology and disease, yet robust benchmarks are scarce. The SMaHT Consortium mixed six HapMap cell lines to create artificial somatic variants spanning 0.25% to 16.5% variant allele fractions. We developed a technology-agnostic method that builds pangenome graphs from individual assemblies to create unified benchmarking sets: > 6M single-nucleotide variants, 1.8M small insertions/deletions, 49K structural variations, and 10K mobile element insertions across autosomes, X, and mitochondrial chromosomes. We validated the variants using ultra-deep simulated reads and developed a binomial-based model to estimate coverage requirements for variant detection. Evaluating multiple callers showed CHM13 alignment improves structural variant detection and offers advantages in difficult-to-map regions compared to GRCh38. Systematic characterization showed regions with low detection rate are enriched in centromeres, satellite sequences, tandem repeats, and falsely duplicated genes. This accurate, versatile resource enables systematic evaluation of somatic variant detection technologies.

genomics↗

Tranquillyzer: A Flexible Neural Network Framework for Structural Annotation and Demultiplexing of Long-Read Transcriptomes

Long-read single-cell RNA sequencing using platforms such as Oxford Nanopore Technologies (ONT) enables full-length transcriptome profiling at single-cell resolution. However, high sequencing error rates, diverse library architectures, and increasing dataset scale introduce major challenges for accurately identifying cell barcodes (CBCs) and unique molecular identifiers (UMIs) - key prerequisites for reliable demultiplexing and deduplication, respectively. Existing pipelines rely on hard-coded heuristics or local transition rules that cannot fully capture this broader structural context and often fail to robustly interpret reads with indel-induced shifts, truncated segments, or non-canonical element ordering. We introduce Tranquillyzer (TRANscript QUantification In Long reads-anaLYZER), a flexible, architecture-aware deep learning framework for processing long-read single-cell RNA-seq data. Tranquillyzer employs a hybrid neural network architecture and a global, context-aware design, and enables precise identification of structural elements - even when elements are shifted, partially degraded, or repeated due to sequencing noise or library construction variability. In addition to supporting established single-cell protocols, Tranquillyzer accommodates custom library formats through rapid, one-time model training on user-defined label schemas, typically completed within a few hours on standard GPUs. Additional features such as scalability across large datasets and comprehensive visualization capabilities further position Tranquillyzer as a flexible and scalable framework solution for processing long-read single-cell transcriptomic datasets.

bioinformatics↗