bioRxiv Science⌕ Search

Biology subjects

Kurz, N. S.

Publications and source records attributed to Kurz, N. S..

6 recordsLinked to original sources

AstraKit: Customizable, reproducible workflows for biomedical research and precision medicine

MotivationThe success of precision medicine and biomedical research depends on the availability of efficient software solutions for processing and interpreting genetic variants, interpreting multi-omics data, and integrating drug screen analyses. However, fragmented bioinformatics tools compel researchers and clinicians to resort to error-prone manual pipelines. ResultsWe present AstraKit, a unified KNIME workflow suite enabling end-to-end precision medicine analytics. AstraKit introduces three transformative innovations: 1) Dynamic variant interpretation with customizable annotation and filtering for disease-specific genomic contexts; 2) Multi-layered omics analyses integrating genomic, transcriptomic, and epigenetic data; and 3) Translational drug matching that correlates in vitro drug screens with clinical outcomes. Validated across oncology cohorts, AstraKit demonstrates concordance between experimental drug sensitivity and clinical outcomes, resolving discordances to uncover resistance mechanisms. By unifying variant analysis, multi-omics, and drug response modeling on a single customizable platform, AstraKit eliminates siloed workflows, accelerating biomarker validation and enabling clinicians to directly link molecular profiles to therapeutic decisions. As all AstraKit workflows are open-source and platform-independent, we provide a versatile comprehensive software suite for a multitude of tasks in bioinformatics and precision medicine. Availability and implementationThe KNIME workflows are available at KNIME Hub https://hub.knime.com/bioinf_goe/spaces/Public/AstraKit~lfVsGBY2HnPYc1h1/. The source code is available at https://gitlab.gwdg.de/MedBioinf/mtb/astrakit.

bioinformatics↗

PanCNV-Explorer: Deciphering copy number alterations across human cancers

Copy number variants (CNVs) are major drivers of cancer progression and genetic disorders, yet their interpretation, spanning biological mechanisms, clinical relevance, and therapeutic implications, remains fragmented across disparate resources. To bridge this gap, we present PanCNV-Explorer, a unified database integrating harmonized copy number variation data across 33 cancer types, cancer cell lines, and healthy cohorts. PanCNV-Explorer provides a genome-wide atlas of CNV frequency and functional impact, quantifying tissue-specific amplifications and deletions in both cancer and non-cancer contexts through rigorous cross-dataset normalization. The interactive web interface enables researchers to dynamically query CNVs by genomic coordinates or gene symbol, visualize cancer-type-specific frequencies with real-time comparative analysis, and explore integrated genomic features including transcripts, regulatory elements, and gene expression through a zoomable genome browser. Beyond exploration, the platform offers programmatic APIs for pan-cancer CNV analysis and visualization. A public web instance of PanCNV-Explorer is available at https://mtb.bioinf.med.uni-goettingen.de/pancnv-explorer/.

bioinformatics↗

FusionPath: Gene fusion pathogenicity prediction using protein structural data and contextual protein embeddings

Accurate prediction of gene fusion pathogenicity is critical for understanding oncogenic mechanisms and advancing precision oncology. While existing computational methods provide valuable insights, their performance remains limited by incomplete integration of multi-scale biological features and insufficient model interpretability for clinical translation. We present FusionPath, a novel deep learning framework for gene fusion pathogenicity prediction. FusionPath uniquely integrates embeddings from multiple pretrained protein language models, including FusON-pLM and ProtBERT, alongside retained protein domains and Gene Ontology (GO) functional annotations. The model was trained and validated on a large-scale dataset of annotated pathogenic and benign fusions. The model was trained and validated on a rigorously curated dataset of 100,433 gene fusions (78,115 benign, 22,318 pathogenic) derived from FusionPDB, ChimerDB4.0, and 27 RNA-seq datasets of normal tissues. FusionPath significantly outperformed state-of-the-art methods, achieving higher AUC scores of 0.95 and 0.87 on independent test sets. By synergistically leveraging sequence, structural, and functional information with explicit modeling of wild-type sequence context, FusionPath establishes a new standard for gene fusion pathogenicity prediction by effectively leveraging complementary sequence, structural, and functional information.

bioinformatics↗

AdaGenes: A streaming processor for high-throughput annotation and filtering of sequence variant data

MotivationThe amount of sequencing data resulting from whole exome and whole genome sequencing (WES / WGS) presents challenges for annotation, filtering, and analysis. These challenges are exacerbated by the need for efficient and scalable tools that can handle the vast amounts of data generated by modern sequencing technologies. ResultsWe introduce the Adaptive Genes processor (AdaGenes), a sequence variant streaming processor designed to efficiently annotate, filter, LiftOver and transform large-scale VCF files. AdaGenes provides a unified solution for researchers to streamline VCF processing workflows and address common challenges in genomic data processing, e.g. to filter out non-relevant variants to focus on further processing of the relevant positions. Ada-Genes integrates genomic, transcript and protein data annotations, while maintaining scalability and performance for high-throughput workflows. Leveraging a streaming architecture, AdaGenes processes variant data incrementally, enabling high-performance on large files due to low memory consumption. The interactive front end provides the user with the ability to dynamically filter variants based on user-defined criteria. It allows researchers and clinicians to efficiently analyze large genomic datasets, facilitating variant interpretation in diverse genomics applications, such as population studies, clinical diagnostics, and precision medicine. AdaGenes is able to parse and convert multiple file formats while preserving metadata, and provides a report of the changes made to the variant file. Availability and implementationA public instance of AdaGenes is available at https://mtb.bioinf.med.uni-goettingen.de/adagenes. The source code is available at https://gitlab.gwdg.de/MedBioinf/mtb/adagenes.

bioinformatics↗

TCRanalyzer: A user-friendly tool for comprehensive analysis of T-cell diversity, dynamics and potential antigen targets

T cells are critical for immune responses, recognizing antigens via their unique T-cell receptors (TCRs). Analyzing the diverse TCR repertoires, especially the hypervariable CDR3 region, is essential for understanding immune function in health and disease. Current TCR analysis tools often require specialized expertise, computational resources, or sacrifice biological information for efficiency. To address these limitations, we developed TCRanalyzer, a fast and comprehensive TCR analysis pipeline within a user-friendly graphical interface. TCRanalyzer covers all steps from data loading, aggregation and optional sequence clustering, to the analysis of TCR diversity metrics, clonal expansion and antigen specificity. Applied to datasets from patients with either benign or malignant tumors, TCRanalyzer identified changes in TCR clonality, clonal expansion and shifts in antigen specificity across different cohorts or following immunotherapy, thereby demonstrating its potential to dissect critical immunological processes. TCRanalyzer provides a robust and user-friendly tool for TCR sequence analysis, enhancing research in immunology and related fields. AvailabilityTCRanalyzer is available at https://hub.docker.com/r/tcranalyzer/application. Contactnicole.seifert@bioinf.med.uni-goettingen.de

bioinformatics↗

PriOmics: integration of high-throughput proteomic data with complementary omics layers using mixed graphical modeling with group priors

Mass spectrometry (MS)-based high-throughput proteomics data cover abundances of 1,000s of proteins and facilitate the study of co- and post-translational modifications (CTMs/PTMs) such as acetylation, ubiquitination, and phosphorylation. Yet, it remains an open question how to holistically explore such data and their relationship to complementary omics layers or phenotypical information. Network inference methods aim for a holistic analysis of data to reveal relationships between molecular variables and to resolve underlying regulatory mechanisms. Among those, graphical models have received increased attention as they can distinguish direct from indirect relationships, aside from their generalizability to diverse data types. We propose PriOmics as a graphical modeling approach to integrate proteomics data with complementary omics layers and pheno- and genotypical information. PriOmics models intensities of individual peptides and incorporates their protein affiliation as prior knowledge in order to resolve statistical relationships between proteins and CTMs/PTMs. We show in simulation studies that PriOmics improves the recovery of statistical associations compared to the state of the art and demonstrate that it can disentangle regulatory effects of protein modifications from those of respective protein abundances. These findings are substantiated in a dataset of Diffuse Large B-Cell Lymphomas (DLBCLs) where we integrate SWATH-MS-based proteomics data with transcriptomic and phenotypic information. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=140 SRC="FIGDIR/small/566517v1_ufig1.gif" ALT="Figure 1"> View larger version (41K): org.highwire.dtl.DTLVardef@19bcac0org.highwire.dtl.DTLVardef@11c126forg.highwire.dtl.DTLVardef@1fe6631org.highwire.dtl.DTLVardef@e73187_HPS_FORMAT_FIGEXP M_FIG C_FIG

systems biology↗