bioRxiv Science⌕ Search

Biology subjects

MinSeo, K.

Publications and source records attributed to MinSeo, K..

3 recordsLinked to original sources

Relaxed selection and convergent loss erode the retained plant-cell-wall-degrading enzymes of ectomycorrhizal fungi

The contraction of the plant-cell-wall-degrading enzyme (PCWDE) repertoire in ectomycorrhizal (ECM) fungi is among the most reproducible results in fungal comparative genomics, and it is routinely read as a loss of decay capacity - a statement about the genes that persist, which counts cannot make. We addressed that question in 183 fungal genomes at two levels and reached opposite conclusions. Gene loss is measurable. On a 182-tip species tree, ECM genomes carry 0.367 times the PCWDE complement of free-living decomposers (phylogenetic generalized least squares; 95% CI 0.295-0.457; P = 1.3 x 10^-15; Pagel's {lambda} = 0.92); the ECM coefficient is not absorbed into the phylogenetic covariance term. The selective regime of the copies that remain is not measurable with the standard branch-partitioned test, and this version retracts the relaxed-selection claims we made for GH6, GH7 and GH28 in version 1. Four independent controls support the retraction. (i) Calibration: applying the identical design to single-copy orthologues, which have no reason to experience a change of selective regime specific to ECM lineages, the relaxation test rejected the null for 35.8% of loci at = 0.05 (29 of 81; 43.9% of the 66 returning an interpretable statistic), against a nominal 5%. (ii) Non-convergence: in 46 of 410 fits the likelihood-ratio statistic was negative, an outcome that cannot occur if the optimisation has converged. (iii) Non-reproducibility: across 80 pairs of fits sharing alignment, tree and branch labels, none recovered its own log-likelihood to within 0.01 units, while the nested model fits computed inside the same runs agreed to within 1.99 units. (iv) Resolution: with 81 to 82 null replicates the smallest attainable empirical P value is about 0.012, against a Bonferroni threshold of 0.0125 for the four families tested here - a margin we had not computed, and one that any expansion of the family set removes. No family-level statement about selective regime is supported by these data. We emphasise the symmetry: we do not claim that these genes are maintained either. The pattern of gene loss described by previous work stands, and is strengthened by phylogenetic correction; what does not stand is the inference from that pattern to the functional state of the genes that remain.

evolutionary biology↗

Cross-architecture ensembling of DNA foundation models improves the precision and stability of chimera detection in long-read metagenomic bins

MotivationChimeric metagenome-assembled genomes (MAGs) that pool DNA from multiple organisms contaminate downstream analyses. Marker-gene tools such as CheckM2 miss low-level chimerism, and DNA foundation models have been proposed as a sequence-composition alternative, but whether large autoregressive models (Evo2, 7B parameters) outperform smaller contrastive models (DNABERT-S, 117M) has not been rigorously tested. ResultsOn 131 MAGs from CAMI2 Nanopore (21 samples, 32 true chimeras), an Evo2 embedding-distance detector achieved high recall (0.84) but low precision (0.34, F1 0.49 [95% CI 0.37-0.60]), producing 52 false positives. Systematic diagnosis revealed convergent multi-factor bias: cross-sample same-species redundancy (50%), low coverage (76% FP at <11x coverage), AT-rich GC (95% FP at GC<0.62), and concentration in Pseudomonas_E/Rhizobium genera. DNABERT-S alone reached F1 0.61 [0.46-0.74] with only 11 false positives, showing no such bias despite 60x fewer parameters. Their errors were largely independent: 87% of Evo2s false positives were correctly rejected by DNABERT-S, while 36% of DNABERT-Ss false positives were rejected by Evo2. Their union at tuned thresholds (Evo2 > 43.96 OR DNABERT-S > 1.80) achieved in-sample F1 0.65 [0.49-0.79] with 7 false positives, and held-out test F1 0.57 with lowest cross-validation variance ({sigma}=0.05 vs 0.09-0.11 for single models). Ensemble improvement over single models is most robust in precision (0.73 vs 0.34, non-overlapping CIs), while F1 gains are marginal given sample size. On real Nanopore data (ZymoBIOMICS D6331), the tuned threshold transferred without false positives, though well-defined mock communities bin cleanly and reveal a strain-level detection floor for sequence-embedding methods. Training objective and cross-architecture complementarity matter more than parameter scale. Availabilityhttps://github.com/sunsungkim04-sys/evo2-mag. Contact2023024947@knu.ac.kr

bioinformatics↗

Context-dependent calibration of Evo2 likelihood with bacterial fitness: a quantitative characterization across five E. coli datasets

DNA foundation models are trained to predict the likelihood of natural sequences, but the calibration between such likelihood scores and laboratory fitness or directly measured molecular phenotypes depends strongly on gene context, sequence divergence from wild-type, and selection regime. We apply zero-shot variant scoring with Evo2 7B ({Delta}LLR, the change in pseudo-log-likelihood between mutant and reference windows) to five E. coli datasets and quantify this context-dependent calibration map. Calibration is strong in two settings. In the Firnberg 2014 deep mutational scan of TEM-1 {beta}-lactamase (13,027 nucleotide-level variants; plasmid-borne enzyme under band-pass ampicillin selection), Evo2 {Delta}LLR tracks measured fitness at Spearman {rho} = 0.545 (95% CI 0.532-0.557; SNV {rho} = 0.606, indel {rho} = 0.521). In the Tenaillon 2012 thermal-evolution dataset, type-stratified, window-tuned scoring reaches Insertion AUROC 0.882 (W = 2,048 bp) and Deletion AUROC 0.846 (W = 4,096 bp). Calibration is decisively absent in the same organism: the Ireland 2020 RegSeq promoter MPRA gives{rho} = 0.011 (95% CI 0.003-0.019; n = 64,665), flat even after -10/-35 mechanism stratification, and the Dewachter 2023 chromosomal-essentials scan (fabZ/lpxC/murA) gives{rho} = 0.041 (95% CI 0.025-0.058). The Papkou 2023 folA combinatorial landscape sits between, at{rho} = 0.237, with a sweep that falls monotonically from {rho} = 0.575 at two mutations from wild-type to {rho} = 0.065 at nine. Pooling per-gene and per-divergence correlations, we fit calibration as an explicit function{rho} = f(sequence divergence from WT, variant context): weighted regression gives a negative divergence coefficient and a negative regulatory-context coefficient (both in the predicted direction; R2 = 0.49) -- an explicit, if illustrative, fit rather than a metaphor. We further test -- and find unsupported -- the intuitive explanation for the residual TEM-1 vs. essentials gap: across five genes the chromosomal essentials are more represented than TEM-1 by raw public-database deposition count yet calibrate far worse (calibration does not track deposition count; if anything, inversely), so simple training over-representation does not explain the gap. Deposited variant diversity is a candidate but remains untested. We therefore reframe Evo2 not as a fitness predictor but as a likelihood predictor whose calibration with fitness is context-dependent. The deliverable is not a DMS pre-screen tool but a quantitative lookup table of when, where, and why the likelihood- fitness gap closes (training-rich plasmid CDS under stringent selection) or opens (chromosomal essentials, native promoter regulatory variants). Even within a single organism, plasmid vs. chromosomal context and strong vs. weak selection yield qualitatively different calibration regimes -- the central finding.

bioinformatics↗