bioRxiv Science⌕ Search

Biology subjects

Mahabal, A.

Publications and source records attributed to Mahabal, A..

2 recordsLinked to original sources

Predicting Fungal Contaminants for Space Missions Using Proteome-Wide Screening for Protein Orthologs

Fungal contamination poses a growing threat to spacecraft integrity, crew health, and planetary protection efforts. We describe a scalable and interpretable pipeline for identifying fungi with adaptation potential to spaceflight-associated stress conditions such as extreme temperatures, radiation levels, etc., and pathogenicity risks. Starting with proteins known to confer stress resistance, we identify orthologs across over fifteen hundred fungal species and evaluate their contamination potential via comparative proteome analysis. Our pipeline integrates proteins with known functional inference, cross-database proteome matching, and identity-based scoring to generate a ranked list of fungal species of concern. We apply this approach to detections from spacecraft assembly facilities, highlighting species with combined stress-tolerance and pathogenic potential. This study establishes a foundation for future AI-based risk assessments that can scale to orders of magnitude more fungal species, thus laying the foundation for systematic identification and assessment of fungal contaminants with potential adaptation and pathogenicity risks in spaceflight environments, thereby supporting contamination control strategies for future space missions. We also present an interactive visual online tool for researchers to trivially check the contamination potential of species in their own samples.

microbiology↗

BiomarkerKB: FAIR and Integrated Biomarker Knowledge Connecting Biomolecular and Clinical Data Types

Biomarkers are essential tools for disease detection, risk assessment, therapeutic monitoring, and precision medicine. However, biomarker data are dispersed across heterogeneous resources, inconsistently reported in the literature, and rarely standardized for computational use. This fragmentation limits reproducibility, cross-study integration, and the discovery of novel biomarker and disease relationships. We developed BiomarkerKB, a knowledgebase designed to harmonize and integrate biomarker information under a standardized data model. The model follows the FDA-NIH BEST biomarker definition and captures both core fields (biomarker entity, disease/condition, exposure agent) and contextual metadata (specimen, biomarker role, evidence, provenance). Biomarker data and related annotations were either curated from publications or collected from public resources (e.g., OpenTargets, GWAS Catalog, ClinVar, CIViC, OncoMX) and were also contributed by the Common Fund Data Coordinating Centers and the Early Detection Research Network (EDRN). Standardization was achieved using ontologies and reference resources such as Disease Ontology, UBERON, UniProtKB, and HUGO Gene Nomenclature Committee (HGNC) gene symbols. BiomarkerKB data were ingested into a Neo4j-based knowledge graph and integrated with the Common Fund Data Ecosystem (CFDE) Knowledge Graph. The initial release of BiomarkerKB contains over 200,000 biomarker-disease associations spanning genes, proteins, metabolites, glycans, and chemical elements. The knowledge graph comprises more than 300,000 nodes and 1.2 million edges, enabling structured exploration of biomarker relationships within CFDE data as demonstrated through the knowledge graph query-based use cases presented in this study. A publicly accessible web portal (https://biomarkerkb.org) provides keyword search, filtering, data downloads, and access to graph visualization to support both researchers and computational analyses. BiomarkerKB addresses a critical gap in biomarker informatics by providing an integrated, FAIR (Findable, Accessible, Interoperable, and Reusable), and unified framework for biomarker knowledge exploration and discovery.

bioinformatics↗