bioRxiv Science⌕ Search

Biology subjects

Brusselmans, M.

Publications and source records attributed to Brusselmans, M..

5 recordsLinked to original sources

Better data, better trees: GenBank-GISAID deduplication and source-specific artifact masking in viral genomics

GenBank and GISAID are the primary repositories for viral genomic data, but integrating records across them remains a challenge. The same sequence could be made available in both databases without any cross-reference linking the two entries. Consequently, there is no systematic way to identify this redundancy, which compromises the compilation of representative, non-redundant large-scale datasets. In parallel, the growth of viral genomic data has increased the risk of systematic technical artifacts introduced during sequencing or assembly. These artifacts can inflate substitution rate estimates and degrade temporal signal, biasing evolutionary rate estimates. To address both challenges, here we present a formal, reproducible workflow integrating two newly developed complementary tools: G2G matcher for cross-repository harmonization and Lab-Specific Bias FILTer (LSBFILT) for masking of laboratory-specific artifacts. Using the Eastern/Central/South African (ECSA) chikungunya virus lineage as a proof-of-concept, we demonstrate that our integrated workflow restores temporal signal and provides a robust, curated dataset for downstream phylodynamic analyses. Critically, restricting masking of homoplastic sites to specific sequences reduces the substitution rate estimate from an inflated 8.517 x 10-4 to 5.078 x 10-4 substitutions/site/year and increases the coefficient of determination (R2) of the root-to-tip regression analysis from 0.353 to 0.677. By enabling systematic cross-repository harmonization and source-specific artifact masking, we provide the molecular epidemiological community with scalable tools to reconcile fragmented genomic data and reduce technical biases, fostering more accurate and reproducible phylogenetic analysis. G2G matcher is available at https://github.com/andrezaleite/G2G-Matcher, and LSBFILT at https://github.com/khourious/LSBFILT.

bioinformatics↗

An optimized RNA polymerase II minigenome system for Nipah virus

Nipah virus is a highly lethal, zoonotic paramyxovirus that has caused recurring outbreaks in several South and Southeast Asian countries since its discovery in Malaysia in 1998. Symptoms of infection include severe respiratory and neurological disease, often resulting in death. As no approved vaccines or antivirals are currently available to reduce the burden of this virus, it is classified as a biosafety level 4 pathogen. There is an urgent need for systems that enable research in a lower biocontainment setting, especially since the World Health Organization declared Nipah virus a priority pathogen for pandemic concern. In the past, several minigenome systems have already been developed as safe alternatives to working with infectious virus; however, these systems remain relatively inefficient and lack robustness and reliability for further applications. Therefore, we developed novel optimized RNA polymerase II-driven minigenomes with nanoluciferase or enhanced green fluorescent protein reporter genes. Both systems outperform previously designed Nipah virus minigenomes, are easily operable, and can be implemented for antiviral compound screenings.

microbiology↗

Dispersal, adaptation and persistence of H5N1 in the sub-Antarctic and Antarctica

High pathogenicity avian influenza virus (HPAIV) H5N1 reached the sub-Antarctic and Antarctica in 2023, subsequently spreading to remote locations within this region where it had devastating impacts on seal, penguin and albatross populations. The threat to marine wildlife over this broad area exemplifies the need to understand H5N1 long-distance dispersal and evolution. We obtained 104 novel viral genomic sequences from samples that we collected at South Georgia, Kerguelen, Crozet, Prince Edward, Falklands/Malvinas Islands and the Antarctic Peninsula in a region spanning 8,000 kilometers. Using recent phylogeographic modeling advances we show that H5N1 spread encompassed numerous transmission events between distant locations, accumulating mammalian-adaptive mutations in the process. Seals are the most affected species, but we reveal that the long-distance eastward virus dispersal better aligns with the long-distance movements of large petrels and albatrosses. The risk of H5N1 endemisation, dispersal to other locations and ongoing evolution are highly concerning.

microbiology↗

Biological causes and impacts of rugged tree landscapes in phylodynamic inference

Phylodynamic analysis has been instrumental in elucidating epidemiological and evolutionary dynamics of pathogens. Bayesian phylodynamics integrates out phylogenetic uncertainty, which is typically substantial in phylodynamic datasets due to limited genetic diversity. Phylodynamic inference does not, however, scale with modern datasets, partly due to difficulties in traversing tree space. Here, we characterize tree space and landscape in phylodynamic inference and assess its impacts on analysis difficulty and key biological estimates. By running extensive Bayesian analyses of 15 classic large phylodynamic datasets and carefully analyzing the posterior samples, we find that the posterior tree landscape is diffuse yet rugged, leading to widespread tree sampling problems that usually stem from sequences in a small part of the tree. We develop clade-specific diagnostics to show that a few sequences--including putative recombinants and recurrent mutants--frequently drive the ruggedness and sampling problems, although existing data-quality tests show limited power to detect them. The sampling problems can significantly impact phylodynamic inferences or distort major biological conclusions; the impact is usually stronger on "local" estimates (e.g., introduction history) associated with particular clades than on "global" parameters (e.g., demographic trajectory) governed by general tree shape. We evaluate existing and newly-developed MCMC diagnostics, and offer strategies for optimizing phylodynamic analysis settings and mitigating sampling problem impacts. Our findings highlight the need and directions to develop efficient traversal over rugged tree landscapes, ultimately advancing scalable and reliable phylodynamics. Significance StatementBayesian phylodynamics is central to epidemiological studies, but exploring the vast and complex tree space is computationally challenging. Phylodynamic datasets comprise many highly similar sequences, sampled through time, creating a uniquely structured landscape of optimal trees. Here, we show that phylodynamic tree landscapes are often highly rugged, with multiple peaks separated by difficult-to-cross valleys. These features lead to widespread sampling problems which are often driven by a few sequences. These problems can significantly impact phylodynamic estimates, especially those associated with particular clades, distorting biological conclusions. We develop diagnostics to identify problematic sequences and provide solutions to mitigate their impacts. We offer strategies to optimize phylodynamic analysis workflows and to develop algorithms for navigating rugged landscapes, thereby advancing infectious disease investigation.

evolutionary biology↗

HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator X

In Bayesian phylogenetic and phylodynamic studies it is common to summarise the posterior distribution of trees with a time-calibrated consensus phylogeny. While the maximum clade credibility (MCC) tree is often used for this purpose, we here show that a novel consensus tree method - the highest independent posterior subtree reconstruction, or HIPSTR - contains consistently higher supported clades over MCC. We also provide faster computational routines for estimating both consensus trees in an updated version of TreeAnnotator X, an open-source software program that summarizes the information from a sample of trees and returns many helpful statistics such as individual clade credibilities contained in the consensus tree. HIPSTR and MCC reconstructions on two Ebola virus and two SARS-CoV-2 data sets show that HIPSTR yields consensus trees that consistently contain clades with higher support compared to MCC trees. The MCC trees regularly fail to include several clades with very high posterior probability ([≥] 0.95) as well as a large number of clades with moderate to high posterior probability ([≥] 0.50), whereas HIPSTR achieves near-perfect performance in this respect. HIPSTR also exhibits favorable computational performance over MCC in TreeAnnotator X. Comparison to the recently developed CCD0-MAP algorithm yielded mixed results, and requires more in-depth exploration in follow-up studies. TreeAnnotator X - which is part of the BEAST X (v10.5.0) software package - is available at https://github.com/beast-dev/beast-mcmc/releases.

bioinformatics↗