bioRxiv Science⌕ Search

Biology subjects

Molchanova, M.

Publications and source records attributed to Molchanova, M..

2 recordsLinked to original sources

Evaluating Reference-Independent Pipelines for the Detection of Spreading Organisms in Metagenomic Datasets

The emergence of unidentified pathogens, or "Disease X," poses a significant threat to global health, necessitating the development of proactive surveillance strategies for the wildlife and human virosphere. Since novel viruses often lack universal genetic markers or known homologs, this study evaluates four reference-independent computational pipelines: coverage-based, k-mer-based, nucleotide clustering, and Large Language Model (LLM)-based designed to detect spreading organisms by comparing distinct metagenomic datasets. Using a real-world pandemic dataset of human nasopharyngeal RNA-seq runs and a semi-synthetic dataset enriched with divergent Egovirales sequences, we measured the sensitivity, selectivity, and computational efficiency of each approach. The coverage-based method proved most robust, consistently achieving 100% genome coverage of SARS-CoV-2 and maintaining high selectivity even at low viral concentrations, though it required extensive computational resources (20 days of CPU time for 2B reads). In contrast, the k-mer-based approach offered a tenfold reduction in execution time and high selectivity but was sensitive to data depletion, failing to detect targets at very low abundances. The clustering-based pipeline performed effectively at moderate concentrations but suffered from sequence fragmentation in sparse data, while the LLM-based method (using ViraLM), despite its efficiency, exhibited critically low selectivity due to current latent space partitioning limitations. These results demonstrate that while k-mer and LLM-based tools provide rapid screening capabilities, the coverage-based approach remains the most reliable for sensitive pathogen discovery. Ultimately, these reference-independent workflows are essential for illuminating metagenomic "dark matter" and establishing early warning systems for emerging infectious diseases

bioinformatics↗

AliMarko: A Novel Tool for Eukaryotic Virus Identification Using Expert-Guided Approach

Metagenomic sequencing is a valuable tool for studying viral diversity in biological samples. Analyzing this data is complex due to the high variability of viral genomes and their low representation in databases. We present the Alimarko pipeline, designed to streamline virus identification in metagenomic data. A key feature of our tool is the focus on the interpretability of findings: results are provided with tabular and visual information to help determine the confidence level in the identified viral sequences. The pipeline employs two approaches for identifying viral sequences: mapping to reference genomes and de novo assembly followed by the application of Hidden Markov Models (HMM). Additionally, it includes a step for phylogenetic analysis, which constructs a phylogenetic tree to determine the evolutionary relationships with reference sequences. We also emphasize reducing false-positive results. Reads related to cellular organisms are computationally depleted, and the identified viral sequences are checked against a list of potential contaminants. The output is an HTML document containing visualizations and tabular information designed to assist researchers in making informed decisions about the presence of viruses. Using our pipeline for total RNA sequencing of bat feces, we identified a range of viruses and rapidly determined the validity and phylogenetic relationships of the findings to known sequences with the aid of reports generated by AliMarko.

bioinformatics↗