bioRxiv Science⌕ Search

Biology subjects

Cloutier, S.

Publications and source records attributed to Cloutier, S..

6 recordsLinked to original sources

Integrated pangenome and population genomics reveal selection on standing genetic variation driving fiber flax-linseed divergence

Flax (Linum usitatissimum L.) has been domesticated for dual end uses as linseed and fiber flax, yet the genomic basis of morphotype divergence remains unclear. Here, we constructed a morphotype-resolved pangenome by integrating three newly generated near telomere-to-telomere genome assemblies with 14 previously published ones. Despite substantial variation in assembly size, driven primarily by DNA transposons, gene content was highly conserved, with little evidence for significant morphotype-specific gene presence-absence variation. Population genomic analyses of 407 accessions revealed that fiber flax had reduced nucleotide diversity, extended linkage disequilibrium, and a more compact population structure relative to linseed, consistent with stronger selection and a narrower genetic base. Genome-wide differentiation was heterogeneous and concentrated in discrete regions. Integration of FST, nucleotide diversity ratios, Tajimas D, and genome-wide association signals identified morphotype-enriched genomic blocks distributed across the genome. Many candidate regions are primarily supported by directional shifts in nucleotide diversity rather than extreme differentiation, indicating selection on standing genetic variation. Genome-wide association analyses identified 1,712 unique quantitative trait nucleotides (QTNs), with predominantly small effect sizes and strong enrichment in gene-proximal regions, consistent with a polygenic architecture. Overall, fiber flax traits tend to be controlled by fewer loci with moderate-to-large effects, whereas linseed traits exhibit a more diffuse genetic architecture. Patterns of Tajimas D further support non-classical selection dynamics, with predominantly positive values in linseed and localized negative values in fiber flax, consistent with selection on standing genetic variation. Together, our results suggest that flax morphotype divergence is driven primarily by selection on pre-existing allelic variation within a conserved gene repertoire. This study provides a comprehensive framework linking genome structure, population genomics, and trait architecture, and highlights the importance of standing genetic variation as a key resource for flax breeding and improvement.

genomics↗

Evolutionary dynamics of Aegilops revealed through comparative genome assembly of all 25 species

Aegilops species are the closest wild relatives of wheat and an important reservoir of genetic diversity for its improvement. Despite their potential, many Aegilops genomes remain poorly characterized. Here we present high-quality assemblies of 18 diploid, tetraploid, and hexaploid Aegilops genomes, which, along with the previously published genomes, complete the production of reference assemblies for all 25 genomes in this genus. Assembly sizes ranged from 5.24 Gb in diploids to 12.65 Gb in hexaploids, with scaffold N50 values up to 749.2 Mb. Gene annotation identified 53,035-156,779 protein-coding genes, of which 21,865-60,490 were classified as high-confidence. Orthogroup-based pangenome analysis across the 25 Aegilops genomes identified 80,521 orthogroups, including 15,809 core, 61,735 dispensable, and 2,977 species-specific orthogroups, highlighting substantial gene content variation among genomes. Phylogenetic analysis of 63 Triticum and Aegilops genomes/subgenomes based on near single-copy orthologs defines the phylogenetic relationships within the Triticum/Aegilops complex and confirms diploid progenitors of polyploid lineages. Ae. mutica (T) and Ae. speltoides (S) belong to the B lineage while the remaining Sitopsis grouped within the D lineage. Structural variation analyses using diploid progenitors as references revealed extensive large-scale rearrangements following polyploidization, emphasizing the dynamics of their evolution. Transposable element (TE) annotation further highlighted subgenome-specific TE expansions and contractions, providing insights into the mechanisms shaping genome structure after polyploidization. Collectively, these genomic resources provide a comprehensive framework for exploring Aegilops diversity, understanding polyploid evolution, and accelerating wheat improvement.

genomics↗

Near telomere-to-telomere Linum genomes reveal a lineage-specific DNA transposon associated with chromosome architecture remodeling

Chromosome number variation and structural reorganization are key drivers of plant evolution, yet their genomic basis remains unclear due to incomplete representation of repetitive regions in existing assemblies. The Linum genus exhibits exceptional karyotypic diversity (n = 7-43), providing a powerful system to investigate chromosome evolution. Here, we generated near telomere-to-telomere (T2T) genome assemblies for four species, including cultivated flax (L. usitatissimum cv. CDC Bethune; n = 15), its wild progenitor (L. bienne; n = 15), and two related species (L. decumbens and L. grandiflorum; n = 8). Together with published genomes of L. lewisii (n = 9) and L. tenue (n = 10), these enabled reconstruction of chromosome evolution across six lineages. Phylogenomic analyses revealed a shared ancestral whole-genome duplication (WGD) associated with the n = 9 karyotype, followed by lineage-specific WGDs and divergent diploidization. The transition from n = 8 to the derived n = 15 flax lineage not only occurred without chromosome length expansion, but also with genome size reduction, indicating extensive internal restructuring. Comparative analyses showed that this restructuring was associated with lineage-specific expansion of a single DNA transposon family (TE_00003234; hAT), which is highly enriched in expansive pericentromeric regions that are characterized by low gene density and nucleotide diversity, suppressed recombination, segregation distortion, and extensive synteny disruption, unlike the LTR retrotransposon-rich pericentromeres typical of most plant genomes. These findings support a model in which lineage-specific DNA transposon expansion is associated with remodeling of pericentromeric architecture and large-scale chromosome restructuring following polyploidization.

genomics↗

Genomic selection for seed yield enhances flax breeding efficiency

Genomic selection (GS) is a promising strategy to improve breeding efficiency for complex traits such as seed yield by enabling early selection and reducing reliance on extensive field testing. However, practical deployment of GS remains challenging due to limited training populations sizes and reduced predictive ability when models are applied to true breeding germplasm. In this study, we evaluated GS for flax (Linum usitatissimum L.) seed yield under realistic breeding scenarios, with a focus on across-population prediction (APP) and breeding decision support rather than model benchmarking. Using historical germplasm collections and a newly developed breeding-oriented population as training sets, GS performance was assessed across multiple independent test populations representing contemporary breeding lines evaluated in replicated yield trials. APP predictive abilities ranged from r = 0.67 to 0.84 depending on population relatedness when training and test populations were genetically aligned, supporting routine breeding deployment. Training population composition emerged as a key determinant of prediction success, with breeding-oriented populations consistently outperforming broad germplasm collections for predicting true breeding lines. Check-based selection analyses showed that GS reliably reproduced phenotypic advancement decisions while eliminating 61-91% of low-performing lines, resulting in 48-78% reduction in field evaluation costs for a typical cohort of 300 lines. Marker subsampling analyses further indicated that moderate-density genotyping-by-sequencing panels ([~]2,500-3,000 SNPs) are sufficient to achieve stable predictive abilities. Overall, these results demonstrate that GS for seed yield in flax is ready for routine integration into breeding programs, offering a practical pathway to reduce costs, accelerate breeding cycles, and enhance selection efficiency.

genomics↗

MultiGS: A comprehensive and user-friendly genomic prediction platform Integrating statistical, machine learning, and deep learning models for breeders

Genomic selection (GS) is a core strategy in modern breeding programs, yet the rapid expansion of statistical, machine-learning (ML), and deep-learning (DL) models has made systematic evaluation and practical deployment increasingly challenging. To address these issues, we developed MultiGS, a unified and user-friendly framework that integrates linear, ML, DL, hybrid, and ensemble GS models within a standardized and computationally efficient workflow. MultiGS is implemented through two complementary pipelines: MultiGS-R, a Java/R pipeline implementing 12 statistical and ML models, and MultiGS-P, a Python pipeline integrating 17 models including five linear models, three ML approaches, and nine recently developed DL architectures implemented within the framework. We benchmarked MultiGS using wheat, maize, and flax datasets representing contrasting prediction scenarios. Wheat and maize were evaluated using random training-test splits within the same population, reflecting suitable conditions for assessing model capacity and scalability. Under these scenarios, several DL, hybrid, and ensemble models achieved prediction accuracies comparable to RR-BLUP and consistently exceeded those of GBLUP. In contrast, the flax dataset represented a true across-population prediction scenario with limited training set size and strong population structure. In this challenging context, classical linear models provided stable baselines, while a subset of DL architectures--particularly graph-based models and BLUP-integrated hybrids--demonstrated comparatively improved generalization across populations. Comparisons with previously published DL tools showed that MultiGS models achieved comparable or improved prediction accuracies while requiring lower computational costs, enabling routine retraining and large-scale evaluation. Overall, MultiGS informs, scenario-specific model selection and provides a practical platform for deploying genomic prediction under realistic breeding conditions. The software is freely available on GitHub (https://github.com/AAFC-ORDC-Crop-Bioinfomatics/MultiGS).

bioinformatics↗

Phyllobacterium meliloti sp. nov. a novel non-symbiotic bacterium isolated from root nodules of Melilotus albus (white sweet clover) grown in Canada

Two novel bacterial strains isolated from root-nodules of white sweet clover (Melilotus albus) plants grown at a Canadian site were previously characterized and placed in the genus Phyllobacterium. Here we present phylogenomic and phenotypic data to support the description of strain T1293T as representative of a novel species and present the first complete closed genome sequence of a bacterial strain (T1018) representing the species P. pellucidum. Phylogenetic analysis of genome sequences as well as analysis of 53 core genes placed novel strain T1293T in a highly supported cluster of strains distinct from named Phyllobacterium species with P. myrsinacearum and P. calauticae as closest relatives. The highest average nucleotide identity (ANI) and digital DNA-DNA hybridization (dDDH) values of genome sequences of T1293T compared to closest species type strains (84.1% and 26.5%, respectively) are well below the threshold values for bacterial species circumscription. The genome of strain T1293T has a size of 5074034 bp with a DNA G+C content of 55 mol% and possesses three plasmids with sizes of 397619 bp, 476847 bp and 519835 bp. Detected in the genome were Type III and Type VI secretion system genes, implicated in plant-microbe and microbe-microbe interactions, but key nodulation, nitrogen-fixation and photosystem genes were not detected. Further analysis revealed that T1293T, like other Phyllobacterium species, possesses key genes encoding an enzyme complex implicated in the degradation of glyphosate, a widely used broad-spectrum herbicide that has negative consequences for many microorganisms including the human gut microbiome. A novel prophage (size [~] 41.5 kb) was also detected in the genome of T1293T. Data for multiple phenotypic tests complemented the sequence-based characterization of strain T1293T. The data presented support the description of a new species and the name Phyllobacterium meliloti sp. nov. is proposed with T1293T = LMG32641T = HAMBI 3765T as the species type strain.

microbiology↗