bioRxiv Science⌕ Search

bioRxiv · 10.1101/2025.09.09.675119

Structure-based phylogenetic analysis reveals multiple events of convergent evolution of cysteine-rich antimicrobial peptides in legume-rhizobium symbiosis

Abstract

Nitrogen is essential for plant growth, yet its availability often limits agricultural productivity. Some legumes have evolved a unique ability to form symbiotic relationships with nitrogen-fixing soil bacteria called rhizobia, enabling them to thrive in nitrogen-deficient soils. In five legume clades, an exploitive strategy has evolved in which rhizobia undergo Terminal Bacteroid Differentiation (TBD), where the bacteria become larger, polyploid, and have a permeabilized membrane. Terminally differentiated bacteria are associated with higher N2-fixation and, thus, a higher return on investment to the plant. In several members of the IRLC (Inverted Repeat-Lacking Clade) and the Dalbergioid clades of legumes, this differentiation process is triggered by a set of apparently unrelated plant antimicrobial peptides with membrane-damaging activity, known as Nodule-specific Cysteine-Rich (NCR) peptides. However, whether NCR peptides are also implicated in symbiotic TBD in other legume clades and whether they are evolutionarily related remains unknown. Here, to address the molecular identity of NCR peptides and their evolution in different legume clades, we performed inter- and intra-clade comparisons of NCR peptides in representative species of four TBD-inducing legume clades. First, we collected genomic and proteomic data of species for which NCR peptides are known (1523 NCR peptides). We then used sequence similarity-based clustering to regroup the NCR peptides, resulting in over 400 different NCR clusters, each clade-specific. We obtained Hidden Markov Models for each cluster and used them to predict NCR peptides in 21 legume genomes (6 clades), including newly generated deep-sequenced root and nodule RNA-seq data of Indigofera argentea (Indigoferoid clade) and newly assembled high-quality transcriptomes of Lupinus luteus and Lupinus mariae-josephae (Genistoid clade), using tailored gene prediction pipeline and transcriptome matching. This resulted in 3710 NCR peptides in species that induce TBD. To date, the rapid diversification of NCR peptides that reduces the sequence similarities has masked the origin of NCR peptide evolution. We obtained high-confidence structural models for one sequence of each cluster. We performed structure-based clustering and phylogenetics, which resulted in 23 superclusters (14 inter-clade and nine clade-specific) that we represent in a structural distance-based tree. Our study revealed that the evolution of NCR peptides is a mix of divergent and convergent processes within each clade. We further chose nine independently evolved NCR peptides to test in vitro whether they are functional analogs in symbiosis. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=145 HEIGHT=200 SRC="FIGDIR/small/675119v1_ufig1.gif" ALT="Figure 1"> View larger version (50K): org.highwire.dtl.DTLVardef@8d1698org.highwire.dtl.DTLVardef@c65b98org.highwire.dtl.DTLVardef@a75994org.highwire.dtl.DTLVardef@ea1a73_HPS_FORMAT_FIGEXP M_FIG Overview of the experimental and computational workflow for NCR peptide detection, characterization, and structural analysis. Nodule and root samples from Indigofera argentea (8 weeks post-inoculation) were collected and subjected to RNA extraction, library preparation, and Illumina PE150 sequencing. Raw RNA-seq reads from two Lupinus species were also included (Lupinus luteus and Lupinus mariae-josephae). Bacteroid differentiation of I. argentea was assessed by flow cytometry and confocal microscopy. Transcriptomes were assembled de novo and analyzed for differential gene expression between root and nodule tissues. NCR peptides were identified from them and other legume genomes and transcriptomes using the SPADA pipeline and HMM profiles from NCR clusters of the known NCR peptides. The putative NCR peptides were filtered based on conserved cysteine motifs, length, and nodule expression to build an exhaustive NCR peptide database. 3D structural predictions of NCR clusters were performed using AlphaFold2 (pLDDT >70), followed by structural clustering (Foldseek) and phylogenetic analysis (Foldtree). Functional validation involved flow cytometry and antimicrobial assays (against Eschericha coli, Sinorhizobium meliloti, and Bacillus subtilis), enabling structural and evolutionary characterization of NCR peptides. The green box at the top represents the experimental analysis, the blue box represents the sequence-based computational pipeline, the red box represents the structure-based computational pipeline, and the grey box at the bottom left represents the functional validation and interpretation of the results. C_FIG

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Boukherissa, A., Sankari, S., Timchenko, T., Bourge, M., Mergaert, P., diCenzo, G. C., Shykoff, J. A., Alunni, B., Rodriguez de la Vega, R. C.. 2025-09-14. Structure-based phylogenetic analysis reveals multiple events of convergent evolution of cysteine-rich antimicrobial peptides in legume-rhizobium symbiosis. https://doi.org/10.1101/2025.09.09.675119

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

RELAX does not reproduce its own estimates at default settings, and its output does not show it

Selection-intensity estimates from RELAX are reported as a point value of K with a likelihood-ratio P. We report that, at default settings and on data of ordinary size, the program does not reproduce its own fits. Of 27 enzyme entries refitted under two optimiser configurations, none reproduced its log-likelihood to within 0.01 units; the median change was 103 units, the largest over 3,400, and four verdicts reversed. Eighty null orthologues reproduced none. A byte-identical command returned a distinct likelihood on every repetition, single-threaded, across three releases, and on alignments simulated under the fitted model, where 3.3 per cent of replicates reproduced. The documented random-number seed never reaches the generator when assigned on the command line, yet reads back as the value supplied. PAML localises the cause: its two-ratio model, without site classes, reproduced its log-likelihood for all 288 genes; its site-class models agreed for 27 to 67 per cent. The instability follows the mixture over sites, not the program. The output does not show it: 46 of 410 fits ended with a negative likelihood-ratio statistic, impossible under convergence, and 123 of 410 report a K re-estimated under a domain restriction rather than the unconstrained maximum. Of 234 published studies using RELAX, none reported a seed. Seeding while holding the thread count at one reproduced sixty of sixty runs on twenty genes under two releases; the seed alone reproduced none of five, and no documentation states the second condition. We recommend that fits be repeated and their dispersion published.

evolutionary biology↗

Sequential accumulation of adaptive alleles forms an inversion supergene in deer mice

Supergenes are clusters of co-inherited loci that affect multiple or complex phenotypes. Despite the growing number of chromosomal inversions identified as supergenes in natural populations, their molecular basis and evolutionary history often remain obscure. Here, we identified two candidate genes, Slc45a2 and Npr3, within a 41-Mb inversion supergene in the deer mouse (Peromyscus maniculatus) that respectively drive darker coats and longer tails - two traits associated with forest adaptation. Mice homozygous for the inversion (inv/inv) exhibit elevated Slc45a2 expression in melanocytes relative to the congenic standard genotype (std/std), disrupting pheomelanin production. In parallel, downregulation of Npr3 in inv/inv mouse growth plates prolongs postnatal growth of caudal vertebrae, resulting in tail elongation. Population-level analyses further implicate that this supergene arose through the subsequent accumulation of the Npr3 allele within the inversion, rather than by capturing all beneficial mutations at its origin.

evolutionary biology↗

Toxin structure shapes palatability in a chemically defended butterfly

The toxicity of chemical defences is well studied, but the potential contribution of compound structure to predator deterrence remains largely unexplored. Whether predation acts more strongly on toxicity or unpalatability remains largely untested, partly because few systems allow toxin structure to vary independently of quantity. Heliconius sara larvae provide such a system: those reared on Passiflora auriculata sequester cyclopentenyl cyanogenic glucosides (CGs), while those reared on P. biflora biosynthesise comparable quantities of aliphatic CGs. Using two invertebrate predators, Camponotus floridanus ants and Hierodula membranacea mantids, we tested whether this structural difference affects palatability independent of toxicity. Mantids rejected larvae with cyclopentenyl CGs more often than larvae with aliphatic CGs, despite no detectable difference in total CG content. This pattern was mirrored in extract-based assays with ants, independently of cyanide release: extracts with cyclopentenyl CGs remained deterrent, while extracts with aliphatic CGs did not differ in deterrence from water. Live larvae, by contrast, elicited similar responses from ants regardless of CG structure. These results show that variation in toxin structure can strongly affect palatability, with some compounds conferring greater protection than others. This demonstrates the importance of chemical structural diversity in the evolution of chemical defences.

evolutionary biology↗