bioRxiv Science⌕ Search

SEARCH · bioRxiv Science

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

Genome dashboards: Framework and Examples

Genomics is a sequence-based informatics science and a 3D structure-based material science. Here we describe a framework for developing genome dashboards specifically designed to unify informatics with studies of chromatin structure and dynamics. The framework is based on the mathematical representation of geometrically exact rod models and the generalization of DNA base pair step parameters. A Model-View-Controller software design approach is proposed to implement genome dashboards as finite state machines either as desktop or web based applications. Two examples are demonstrated using our minimal genome dashboard called G-Dash-min. The data unification achieved with a genome dashboard supports the bi-directional exchange of data between informatics and structure. Thus any experimentally or theoretically determined sequence based informatics track can inform DNA, nucleosome or chromatin modeling (e.g. nucleosome positions) and structure features can be analyzed as informatics tracks in a genome browser (e.g. DNA base pair step parameters: Roll, Tilt, Twist). Here the framework is applied to chromatin, but genome dashboards are more broadly applicable. Genome dashboards are a novel means of investigating structure-function relationships for regions of the genome ranging from base pairs to entire chromosomes and for generating, validating, and testing mechanistic hypotheses.

genomics↗

What is in a lichen? A metagenomic approach to reconstruct the holo-genome of Umbilicaria pustulata

Lichens are valuable models in symbiosis research and promising sources of biosynthetic genes for biotechnological applications. Most lichenized fungi grow slowly, resist aposymbiotic cultivation, and are generally poor candidates for experimentation. Obtaining contiguous, high quality genomes for such symbiotic communities is technically challenging. Here we present the first assembly of a lichen holo-genome from metagenomic whole genome shotgun data comprising both PacBio long reads and Illumina short reads. The nuclear genomes of the two primary components of the lichen symbiosis - the fungus Umbilicaria pustulata (33 Mbp) and the green alga Trebouxia sp. (53 Mbp) - were assembled at contiguities comparable to single-species assemblies. The analysis of the read coverage pattern revealed a relative cellular abundance of approximately 20:1 (fungus:alga). Gap-free, circular sequences for all organellar genomes were obtained. The community of lichen-associated bacteria is dominated by Acidobacteriaceae, and the two largest bacterial contigs belong to the genus Acidobacterium. Gene set analyses showed no evidence of horizontal gene transfer from algae or bacteria into the fungal genome. Our data suggest a lineage-specific loss of a putative gibberellin-20-oxidase in the fungus, a gene fusion in the fungal mitochondrion, and a relocation of an algal chloroplast gene to the algal nucleus. Major technical obstacles during reconstruction of the holo-genome were coverage differences among individual genomes surpassing three orders of magnitude. Moreover, we show that G/C-rich inverted repeats paired with non-random sequencing error in PacBio data can result in missing gene predictions. This likely poses a general problem for genome assemblies based on long reads.

genomics↗

Investigating Bacillus anthracis genomic diversity and trait-specific lineages in an endemic area in northern Tanzania through a combination of traditional and culture-free sequencing approaches

Anthrax, caused by Bacillus anthracis (BA), is a prominent neglected zoonosis with major impacts on human, livestock, and wildlife health. Despite this, limited genomic investigation at the One Health interface constrains current understanding of BA transmission and of the ecological and host factors shaping its diversity and population structure. This includes the possibility of host-specific BA lineages, given that anthrax outbreaks often disproportionally affect individual species. This study characterises the genomic diversity of BA in an endemic area, the Ngorongoro Conservation Area (NCA), in northern Tanzania. We analysed 213 BA genomes from livestock, wildlife and humans from cultured isolates combined with a culture-free targeted capture (TC) approach. NCA sequences formed a distinct genetic cluster compared with those from surrounding areas, and we observed surprisingly high levels of strain diversity within apparent epidemiological clusters, as well as within single animals, though strain diversity was lowest at the within host scale. We found limited evidence for seasonal clustering of cases as well as for BA lineages clustering by host species. This indicates that disproportional impacts on certain species during outbreaks are more likely driven by host ecology factors or, hypothetically, by accessory parts of the bacterial genome not represented in our data. TC-derived data significantly expanded the range of host species and geographic locations for genomic analysis, demonstrating the value of this approach. Although TC data may contain artefactual variation, shared SNP profiles between isolate- and TC-derived genomes gave confidence in its use for genotyping. Our analysis demonstrates unexpectedly high BA strain diversity and limited population structure in this endemic area across a range of spatial scales, including within-host. It further highlights the need for high density sampling and adaptable sequencing strategies to generate adequate BA genomic datasets that can enable informative molecular epidemiological studies of anthrax at the One Health interface. Author SummaryAnthrax continues to threaten the health of people, livestock, and wildlife in many parts of the world, yet we still know surprisingly little about how this disease spreads in nature. One major gap is understanding why some species are affected more strongly during outbreaks, despite assumed equal susceptibility. To investigate this, we studied the genomic diversity of the anthrax-causing bacterium Bacillus anthracis in a large conservation area in northern Tanzania. We combined two ways of generating genetic data: traditional laboratory culture and a culture-free method that allowed us to recover bacterial DNA from a wider range of samples. By analysing a uniquely large and species-diverse dataset of over 200 bacterial genomes from people, livestock and wildlife, we found that bacteria from the study area formed a clearly defined group compared to those from surrounding regions. We also discovered unexpectedly high diversity of strains, not only across the landscape but even within single animals. Despite this diversity, we saw little evidence that certain bacterial lineages are tied to specific host species. Our results suggest that the behaviour and ecology of different animals, rather than host-adapted lineages of the bacterium, likely explain why some species are more affected than others.

genomics↗

Lineage-wide evolution of 3D genome organisation and centromeres in brown algae

Although 3D genome architecture has been described for an increasing number of plant and algal species, comparative analyses across closely related lineages remain scarce. Consequently, fundamental questions persist about how chromatin organization is maintained or reshaped over deep evolutionary time, and how such changes relate to life-history traits, genome size, and linear genome features. Here, we present a comprehensive analysis of 3D chromatin architecture across six brown algae species and one outgroup, spanning the phylogenetic breadth and biological complexity of this key photosynthetic lineage. We show that compact genomes lack chromatin domains whereas larger, transposable element-rich genomes of morphologically complex taxa tend to exhibit structured organization including TAD-like domains. We investigate chromatin folding patterns and gene expression over evolutionary time and uncover 3D chromatin features associated with transitions in sexual systems. Moreover, we reconstruct the ancestral brown algal karyotype, revealing deeply conserved macrosynteny and providing a new framework for interpreting chromosome-scale genome dynamics. Finally, we uncover an ancient and highly conserved association between centromeres and chromodomain-encoding retrotransposons, revealing a remarkable example of convergence in centromere-transposon co-evolution between brown algae and angio-sperms, and one of the most stable examples of centromere-linked transposable elements known in eukaryotes. Together, our findings elucidate the evolutionary history of 3D chromatin and linear genome architectures across an entire eukaryotic lineage and highlight extreme centromere stability in brown algae, providing a powerful point of comparison with land plants and deepening our understanding of genome evolution in independent multicellular lineages. One sentence summaryWe present a lineage-wide evolutionary analysis of 3D genome architecture and organization across brown algae, revealing conserved chromosomal features, lineage-specific rearrangements, and long-term co-evolution of centromeres and retrotransposons that together illuminate how nuclear architecture evolves over hundreds of millions of years.

genomics↗

Genomic selection accuracy and bias using imputed genotypes on growth, welfare and fitness traits in two Pekin duck lines

The current study investigated the genomic selection accuracies and biases estimates from two commercial Pekin duck lines reared under commercial breeding practices. A large dataset of 26K duck records comprising both phenotype and imputed genotype information (60K chip) were analysed for growth, welfare and primary feather length traits. First, we employed mixed linear models with relationship matrices computed from the pedigree (BLUP) or markers (GBLUP) to estimate the variance components and breeding values. Then, we estimated the selection accuracies and selection biases to assess the more appropriate models. Our results showed moderately high imputation accuracies of 0.93 and 0.92 for lines A and D respectively. In both lines, the heritability estimates obtained using the pedigree were generally higher than using genomic markers in all traits considered. These ranged for juvenile weight (JW) from 0.22{+/-}0.01 vs 0.25{+/-}0.01 in line A vs line D using marker information to 0.39{+/-}0.02 to 0.50{+/-}0.02 using the pedigree in line A vs line D for slaughter body weight (BW). We observed very low estimates of heritability for gait 0.07{+/-}0.01 using markers in both lines. Breast muscle depth (BD) also had lower estimates of 0.15-0.16 using markers. For line A, the genomic predictions were generally higher when using the G-matrix than the A-matrix with the highest prediction was for BW (r2=0.68-0.70) and JW with r2 of 0.49. The estimates for gait and foot pad dermatitis (FPD) were greatly improved by using the G-Matrix at 0.58 vs 0.24 and 0.68 vs 0.44 respectively for markers vs pedigree information. For line D, the same improvements for G-Matrix vs A-Matrix were observed with estimates for BD being similar in the two lines. However, for BD the G-Matrix greatly improved the estimates from 0.50 to 0.71 unlike in line A where they remained at 0.50. The bias in line A were minimal (0.01- 0.19) using the G-Matrix compared to 0.02- 0.41 when using A-Matrix. The highest observed bias was for JW followed by BD for the G-matrix whereas when using the A-matrix we observed higher biases in many traits (JW, BW, BD and gait). The biases for line D were generally lower for the G-matrix (0.02 - 0.17 vs 0.00 - 0.19) than those observed in line A using markers whereas higher biases were observed using the pedigree (0.01 - 0.37). Current findings pinpointed that all traits were heritable with higher prediction accuracies and lower biases when using GBLUP as opposed to traditional BLUP. The present study demonstrates the effectiveness of GBLUP for improving prediction accuracy and reducing bias in selection traits of Pekin ducks, particularly for traits with low heritability. Author SummaryThe study explored genomic selection in two commercial Pekin duck lines. Using a large dataset of 26,000 records, including phenotype and genotype data, researchers analyzed growth, welfare, and feather length traits. They applied statistical models to assess variance components and breeding values, comparing traditional pedigree-based methods (BLUP) with genomic marker-based methods (GBLUP). Results showed high imputation accuracies (93% for line A and 92% for line D). Heritability estimates varied, with genomic markers generally producing lower estimates than pedigrees, except for traits like gait and breast muscle depth where genomic predictions were superior. For example, line A showed higher accuracy using genomic data for body weight and juvenile weight. Overall, genomic predictions (GBLUP) provided higher accuracy and lower bias compared to traditional methods, especially for traits with low heritability. This highlights the effectiveness of GBLUP in improving selection processes in Pekin ducks.

genomics↗

Polymorphic 3D genome architecture mediated by transposable elements

The three-dimensional (3D) folding of the genome plays a crucial role in genome regulation. However, how 3D genome structure varies between individuals and consequently influences genome function and evolution remains poorly understood. One potential source of this variation is transposable elements (TEs), genomic parasites whose location and composition vary between species and individuals. Hosts typically silence TEs through enriching them with repressive epigenetic marks, turning euchromatic TEs into heterochromatin islands, which were shown to spatially interact with pericentromeric heterochromatin (PCH). Because most TE insertions are present in only a few individuals within Drosophila populations, we asked whether polymorphism in the presence/absence of TEs drives varying 3D structures through TE-PCH spatial interactions. We performed deep-coverage Hi-C of two wild-type Drosophila strains and developed a Hi-C analysis framework enabling allelic comparisons of spatial interactions with PCH. Supporting our hypothesis, nearly 40% of strain-specific euchromatic TEs cause their adjacent euchromatic regions to be spatially closer to PCH than TE-free homologous alleles. These interactions are not limited to specific TE families, and, surprisingly, telomere-proximal TEs show a similar propensity as centromere-proximal TEs to enhance PCH interactions. The most defining feature of TEs involved in PCH interactions is H3K9me3 enrichment, revealing a chromatin-based mechanism for TE-mediated 3D genome organization broadly applicable across TE families and genome locations. Importantly, TEs involved in PCH interactions reduce the expression of adjacent genes and are evolutionarily young, indicating stronger selection against them. Our study reveals a previously uncharacterized mechanism by which TEs influence the function and evolution of host genomes by generating polymorphic 3D genome organization.

genomics↗

Leveraging whole-genome re-sequencing for diversity, population structure, and a public mid-density genotyping enrichment panel in crimson clover (Trifolium incarnatum L.) for breeding purposes

AO_SCPLOWBSTRACTC_SCPLOWCrimson clover (Trifolium incarnatum L.) is an obligately outcrossing, cool-season annual legume valued for forage and cover cropping, yet genomic resources to support systematic improvement are limited. We performed the first and most comprehensive whole-genome resequencing (WGR) of global crimson clover germplasm to (i) characterize diversity and population structure and (ii) develop a public mid-density enrichment capture panel for breeding applications. A core set of 45 accessions sequenced at [~]50X generated 5.84 million variants, while 149 additional accessions sequenced at [~]2.54X yielded 17.05 million variants. After stringent filtering, we retained 542,790 high-confidence SNPs from the high-coverage dataset and [~]2.4 million from the low-pass cohort. Population analyses (PCA, ADMIXTURE) revealed compact clustering of cultivars, broader dispersion of wild and uncertain-status accessions, and low overall differentiation (FST = 0.0105) with excess heterozygosity (FIS = -0.0592), consistent with obligate outcrossing. Guided by these resources, we designed a 28,913-SNP TWIST hybrid-capture panel enriched for genic regions and evenly distributed across seven chromosomes. This panel is being deployed within Auburn Universitys crimson clover breeding program to support population improvement and cultivar development. The resulting genomic resources provide a reproducible, mid-density genotyping platform for trait discovery, predictive breeding, and diversity monitoring. Together, these advances bring crimson clover genomic resources on par with other legumes such as soybean (Glycine max (L.) Merr.) and alfalfa (Medicago sativa), establishing a robust foundation for genomics-assisted improvement of this key cover and forage crop in U.S. sustainable agriculture. CORE IDEASO_LIWhole-genome re-sequencing of 194 crimson clover accessions revealed >21 M variants. C_LIO_LIHigh-confidence SNP catalogs from 50X and 2X data enable cost-effective genotyping. C_LIO_LIGenetic diversity is weakly structured, with cultivars clustering narrowly by origin. C_LIO_LIA 28,913 SNP enrichment panel delivers uniform genome coverage and >75% genic content. C_LIO_LIThese genomic tools accelerate GWAS, genomic selection, and breeding innovation. C_LI

genomics↗

Alternaria atra from distinct ecological roles share functional genomic repertoires

Fungi, particularly ascomycetes, exhibit diverse ecological lifestyles, including endophytism, pathogenicity, and saprotrophy. Species of the genus Alternaria are taxonomically and ecologically diverse, yet the genomic determinants underlying different lifestyles remain poorly understood. Here, we investigate lifestyle-associated genomic variation in Alternaria atra using two newly collected isolates obtained as plant endophytes. We confirm their taxonomic identity and generate draft genome assemblies for both isolates. We assess their phenotypic behaviour under laboratory conditions and examine their genomic features alongside those of a previously published A. atra isolate described as pathogenic. Despite differing isolation histories, the endophytic and pathogenic isolates exhibit similar behaviour under laboratory conditions and possess highly comparable genomic repertoires, including predicted effector proteins, carbohydrate-active enzymes, and biosynthetic gene clusters. We detect no clear genomic signatures distinguishing endophytic and pathogenic origins or lifestyles. These findings suggest that A. atra harbours a shared genomic repertoire compatible with multiple ecological strategies, supporting a model of lifestyle plasticity rather than fixed genomic specialization. Our results add to growing evidence that genome content alone does not reliably predict ecological roles in ascomycete fungi.

genomics↗

Insights into goatpox virus and sheeppox virus genomes from pangenome graphs

The capripoxviruses comprise three species: goatpox virus (GTPV), sheeppox virus (SPPV) and lumpy skin disease virus (LSDV). They are large double-stranded DNA viruses with highly conserved core genomes and variable terminal regions. Previous studies have described variation in Capripoxvirus gene content, their broader population structure and the contribution of non-coding and structural variation remains opaque. This study investigated the genomic diversity and evolutionary history of GTPV and SPPV using phylogenetics, pangenome variation graphs (PVGs), and gene-specific analyses. We found clear differences in population structure between the two viruses. GTPV had three deeply divergent and genetically stable lineages with limited evidence of recent gene flow, whereas SPPV had weaker clade separation consistent with an ancestral bottleneck followed by recent population expansion. PVG-based analyses indicated that GTPV has a comparatively closed pangenome, while SPPVs was open, particularly at the inverted terminal repeats (ITRs). Structural and haplotype variation was concentrated at these ITRs, which moderate host immunity and specificity. In several lineages, extended putative ORFs spanning adjacent ITR genes were observed, indicating recurrent structural plasticity at these regions. Patterns of gene-specific conservation and divergence highlighted loci under strong constraint and lineage-specific structural changes that may contribute to host specificity. Together, these results demonstrated how graph-based genome models complement gene-based analyses in resolving poxvirus genome evolution and provide a resource for improved comparative and population genomic studies of large DNA viruses. SignificanceThe capripoxviruses are economically important livestock pathogens, yet the genomic mechanisms underlying their diversification and host specificity remain poorly resolved. By applying pangenome variation graphs alongside phylogenetic and gene-level analyses, this study reveals fundamental differences in how goatpox and sheeppox viruses have evolved. Goatpox virus had a deeper, more stable lineage structure, whereas sheeppox virus was more recent and diverse. Importantly, structural variation at the inverted terminal repeats emerged as a major driver of genomic diversity, including lineage-specific haplotypes and variable gene structures. These findings demonstrated the value of graph-based genome representations for resolving complex variation in large DNA viruses and provides approaches for improving genomic surveillance, comparative analyses, and future investigations into host range, virulence and tropism.

genomics↗

A High-Quality Genome Assembly of Chaetoceros muelleri Reveals Extensive Gene Duplication, Functional Diversification, and Unique Lineage-Specific Innovation

Diatoms are major contributors to marine primary production, yet high-quality nuclear genome resources remain scarce for ecologically dominant lineages such as Chaetoceros. Here, we present the first high-quality nuclear genome assembly of Chaetoceros muelleri, generated from living cells resurrected from resting spores preserved in Baltic Sea sediments and sequenced using PacBio HiFi long-read technology. The assembly is compact (43{square}Mb), highly contiguous (N50{square}={square}1.40{square}Mb), and highly complete (93% BUSCO). Comparative analyses across 14 diatom genomes revealed extensive lineage-specific and expanded gene families in C. muelleri, alongside a small, conserved core genome, reflecting rapid evolutionary turnover. Functional enrichment highlighted diversification of polysaccharide biosynthesis, vesicle-mediated trafficking, membrane remodelling, and transcriptional regulation, consistent with adaptations linked to frustule formation and environmental responsiveness. Transposable elements (TEs) strongly shape the genome, accounting for [~]18% of the assembly, with dominant LTR retrotransposons and a large fraction of unclassified repeats suggesting novel or highly diverged TE lineages. Enrichment of DNA replication, recombination, and repair functions further indicates compensatory genome maintenance associated with TE-driven structural dynamics. Direct comparison with C. tenuissimus revealed contrasting patterns of gene family expansion and regulatory innovation, underscoring divergent evolutionary strategies within Chaetoceros. By integrating resurrection ecology with long-read genomics, this study provides a foundational genomic resource for C. muelleri and highlights the role of TE-mediated genome plasticity in diatom evolution.

genomics↗

A Reference Genome for the Critically Endangered Philippine Eagle (Pithecophaga jefferyi), the National Bird of the Philippines

The worlds largest and rarest eagle, the Philippine Eagle (Pithecophaga jefferyi), also known as the monkey-eating eagle, is the national bird of the Philippines. This raptor species is endemic to the Philippine archipelago, with populations on the islands of Luzon, Leyte, Samar, and Mindanao. It is critically endangered, with an average estimated population of 392 potentially breeding pairs or 784 mature individuals. In this paper, we describe a reference genome of the Philippine Eagle (Pithecophaga jefferyi) from a female juvenile from the province of Nueva Ecija on the island of Luzon. We generated a de novo genome assembly with high contiguity and completeness, comprising 178 contigs totaling 1.345 Gbp. The genome was sequenced at a coverage of 75.2x, and Benchmarking Universal Single-Copy Orthologs (BUSCO)/Compleasm analysis yielded a BUSCO score of 99.92% (aves_odb12), corresponding to 99.7% single-copy, 0.21% duplicated, and 0.08% fragmented genes. A consensus mitogenome sequence of 19,377 bp was also generated. The genome assembly included 23,847 putative genes, and our annotation estimated that 15.78% of the genome consisted of repetitive elements. Genome heterozygosity (H) was estimated to be 0.020%, in comparison to other birds with genome heterozygosity values ranging from 0.0103% to 0.923%. Whole-genome comparisons with publicly available genomes suggest that the Philippine eagle belongs to the snake-eagle subfamily (Circaetinae) rather than the harpy-eagle subfamily (Harpiinae). Pairwise sequentially Markovian coalescent (PSMC) analysis suggests that the effective population size was around 4,000 individuals from about 100 KYA to about 1 KYA. Finally, we constructed a minimum spanning network, which revealed that our mitogenome from the northern island of Luzon occupies a peripheral position, separated from the dominant haplotype cluster found in the southern island of Mindanao by multiple mutational steps, indicating substantial mitochondrial divergence.

genomics↗

Single-Plant Genome-Wide Association Study Identifies Loci Controlling Multiple Vegetative Architecture Traits in Cultivated Northern Wild Rice (Zizania palustris L.)

Cultivated Northern Wild Rice (Zizania palustris L.) is an obligately outcrossing, self-incompatible cereal grown in aquatic paddies in the United States. Genetic improvement has relied primarily on phenotypic recurrent selection, and genomic approaches remain largely unexplored in this emerging crop. We applied a single-plant genome-wide association study (sp-GWAS) framework to dissect vegetative architecture traits in five open-pollinated cultivated populations evaluated across three years (n = 2,173 plants). Plant height (PH), basal stem width (BSW), primary stem width (PSW), flag leaf length (FLL), and flag leaf width (FLW) were analyzed using a mixed linear model accounting for population structure and kinship. Broad-sense heritability ranged from 0.03 to 0.34, and year effects explained up to 54% of phenotypic variance, indicating strong environmental influence. After filtering 73,363 SNPs, genome-wide linkage disequilibrium decayed rapidly (r{superscript 2} = 0.1 at [~]2.3 kb). A total of 124 significant SNPs (FDR < 0.01) were consolidated into 98 loci, of which 46 were associated with multiple traits and 11 were shared across four traits. Candidate genes near multi-trait loci included conserved regulatory classes implicated in grass architecture, including HLH/bHLH transcription factors. Diplotype analyses at candidate loci revealed both simple biallelic and complex multi-allelic haplotype structures, indicating that locus-level haplotype effects underlie several GWAS signals. Results demonstrate that sp-GWAS can detect statistically robust associations in a highly heterozygous, non-replicable crop system and suggest a polygenic, coordinated genetic architecture governing vegetative growth. These findings support genomic prediction and multi-trait selection strategies to accelerate improvement of cultivated Northern Wild Rice. PLAIN LANGUAGE SUMMARYCultivated Northern Wild Rice is an important specialty crop grown in flooded paddies in the United States. Unlike many major crops, it is naturally outcrossing and highly variable, which makes traditional breeding challenging and slow. Most improvement efforts have relied on selecting plants based only on how they look in the field, and genomic tools have rarely been used. In this study, we used DNA markers to better understand the genetics behind plant structure traits such as plant height, stem thickness, and leaf width. We evaluated more than 2,000 plants from five cultivated populations over three growing seasons. Because weather and growing conditions strongly influence these traits, we used statistical models to separate environmental effects from genetic effects. We identified 98 regions of the genome associated with variation in plant structure. Many of these regions influenced more than one trait, showing that plant height, stem strength, and leaf size are genetically connected. Several regions contained genes similar to those known to control plant growth and development in other grasses. We also found that, in some cases, combinations of nearby DNA variants (haplotypes) explained trait differences better than single genetic markers. Overall, this work shows that modern genomic tools can successfully identify useful genetic variation in cultivated Northern Wild Rice, even though it is highly outcrossing and genetically diverse. These results provide a foundation for using genomic selection to improve plant structure, lodging resistance, and overall performance in breeding programs. CORE IDEASO_LISingle-plant GWAS successfully detects genetic associations in obligately outcrossing cultivated Northern Wild Rice where conventional replicated mapping populations are impractical. C_LIO_LIVegetative architecture traits exhibit low heritability but retain recoverable polygenic signal, where nearly half of detected loci influence multiple architecture traits, indicating integrated developmental control. C_LIO_LIGenome-wide linkage disequilibrium decays rapidly ([~]2.3 kb), consistent with expectations for an obligately outcrossing species and supporting relatively localized association signals. C_LIO_LICandidate genes include conserved regulatory classes (TE1-like, HLH/bHLH, SPL). C_LIO_LIGiven extensive overlap between QTL and environmental effect, multi-trait, multi-environment genomic prediction provides a pragmatic breeding strategy to improve canopy efficiency, lodging resistance, and harvestability in aquatic production systems. C_LI

genomics↗

Highly contiguous reference genome assembly of the endangered Orces blue whiptail Holcosus orcesi

Holcosus orcesi, the Orces Blue Whiptail, is a Critically Endangered lizard endemic to the upper Jubones River basin in southern Ecuador. Restricted to a narrow elevational range within semi-arid Andean shrublands, it represents one of the few montane members of a predominantly lowland lineage. Here we present the first high-quality reference genome for H. orcesi, generated using Oxford Nanopore Technologies long-read sequencing. The assembly spans 1.68 Gb across only 91 contigs, with an N50 of 76.2 Mb and a BUSCO completeness of 96.8%, making it among the most contiguous and complete squamate genomes to date. Structural annotation predicted 25,682 genes, of which 85% showed homology to known proteins and 45% were assigned Gene Ontology terms. Repetitive elements accounted for 46.3% of the genome, with LINEs representing the predominant class. This genome provides a foundational resource for future evolutionary, comparative and conservation-genomic research of H. orcesi and other mountain reptiles, enabling studies of population genomics, local adaptation, and genomic erosion in isolated populations. By expanding the genomic representation of tropical montane reptiles, this work helps address longstanding phylogenetic and geographic gaps in global biodiversity genomics and provides a foundation for evidence-based conservation of H. orcesi and related taxa.

genomics↗

Genomic Architecture, Differentiation, and Adaptation in Three Large Falcons

Recent chromosomal rearrangements and divergence in large falcon species make them excellent foci for studies on evolution and genomic architecture. Here, we use high-coverage (44-74X) whole genome resequencing with 10X Genomics Linked-Reads to assess patterns of genomic divergence in peregrine, saker, and gyrfalcons and we link these to chromosomal type and chromosomal rearrangements. We first use admixture analysis and cross-coalescent MSMC2 to demonstrate distinct species boundaries between the large falcons and retrace their demography. We assessed genomic landscapes in terms of recombination rate, nucleotide diversity ({pi}), Tajimas D, autozygosity and Fst between saker and gyrfalcons: {pi} had higher values on smaller chromosomes and Fst had higher values on larger chromosomes. Recombination rate concealed other chromosome type effects on {pi} and Tajimas D but largely explained variation in Fst. We find 39 selective sweeps--some shared--across the falcons. However, only five candidate genes--mostly housekeeping genes--were implicated as targets of balancing selection across all falcons, with 4 of these shared between Hierofalco and three shared across all the falcons. Occurrence of selective sweeps and balancing selection were not enriched by chromosome type or in the context of chromosome fusions. Overall, our findings provide insights into divergence and adaptation in large falcons, and demonstrate an association of genomic architecture and chromosomal fusions with all population genomic indicators and metrics of differentiation between species. Significance StatementFalcons are culturally and commercially important birds that have undergone recent chromosomal rearrangement, providing a natural system for studies on chromosomal heterogeneities and evolution. By analyzing genomic variation across three large falcon species, we show that chromosome type and chromosomal fusions structure patterns of recombination, diversity, and divergence. Our findings highlight the importance of underlying genomic architecture to common forms of evolutionary inference and call attention to the role of chromosomal fusions in shaping falcon evolution.

genomics↗

Recombination and repetitive genomic landscapes are decoupled in a close relative of Caenorhabditis elegans

Genomes exhibit chromosomal heterogeneity. Distributions of genes, repetitive elements, and polymorphisms are not uniform along chromosomes in multiple species. One explanation for these patterns is recombination rate variation. As recombination interacts with selection to shape the evolutionary fates of alleles, recombination rate variation could promote differences in the chromosomal distribution of genomic features. Thus, clarifying the relationship between recombination rates and genomic organization is a major goal of genetics. In the nematode Caenorhabditis elegans, recombination rate correlates with multiple genomic features that are non-uniformly distributed along chromosomes; recombination rates and repeats are higher on chromosome ends compared to chromosome centers. Its closest known relative, C. inopinata, harbors a radically altered genome with nearly uniform chromosomal distributions of repetitive elements. Is this dramatic change in genomic organization connected to the evolution of recombination rates? Here, we describe a genetic map of C. inopinata constructed via whole-genome sequencing of 180 individual F2 recombinants. This reveals four chromosomes have a conserved recombination rate domain structure whereas two other chromosomes harbor divergent, more uniform recombination rate distributions. Comparisons of these intrachromosomal recombination rates with genomic features reveal little covariation between recombination rate, diversity, gene density, and repeat content in C. inopinata (in stark contrast to most Caenorhabditis species). This suggests that the evolution of recombination may not be entirely responsible for the atypical uniform distribution of repetitive elements across C. inopinata chromosomes. Taken together, these observations reveal that recombination rates can be decoupled from the genomic organization of repetitive elements.

genomics↗

Trustworthy agentic genomics through versioned skill libraries

Genomics is adopting autonomous AI agents that interpret genomes from natural-language instructions faster than it is building the means to trust them. We report the first large-scale controlled evaluation of where, in an agentic genomic pipeline, correctness must reside for the system to be trustworthy at clinical scale. Using pharmacogenomics, a domain where errors are measurable and sometimes lethal, we benchmarked nine frontier large language models across 44,550 scored evaluations on 110 pharmacogenomic cases, and tested model interpretation of real star-allele diplotypes from more than 7,000 individuals in three ancestrally diverse populations. Trustworthiness proved to be a property of pipeline architecture, not of the model. Letting the model reason was stochastic and unsafe, and grounding it in the correct guidelines by retrieval paradoxically increased lethal-class errors. Encoding the validated decision logic as a versioned skill and executing it as code made the pharmacogenomic mapping exact, auditable and identical across models, confining all residual error to a single input-interpretation step. On individual genomes, unguarded model interpretation degraded along an ancestry gradient; execution removes this gradient from the clinical mapping, relocating it to the auditable completeness of the input caller. This establishes a generalisable, auditable architecture for trustworthy agentic genome interpretation at scale. HighlightsO_LICorrectness must be executed, not reasoned or retrieved, to be trustworthy C_LIO_LIRetrieval raises phenotype accuracy yet increases lethal-class errors; skills do not C_LIO_LIExecution makes the clinical mapping exact and model-invariant; error stays at input C_LIO_LIA deterministic input caller is the predicted route to all-correct emitted answers C_LI In briefCorpas and colleagues show that trustworthy agentic genome interpretation comes not from making language models reason correctly about biology, but from confining them to interpreting input while versioned, validated skills do the reasoning as executed code. Across nine large language models and 110 pharmacogenomics cases, executing the skill makes the clinical mapping deterministic, auditable and model-invariant. SignificanceGenomics is adopting autonomous, language-model-mediated agents faster than it is building the standards needed to trust them. On a pharmacogenomic benchmark with lethal-class consequences, we show that an agents trustworthiness is not a property of the model but of how the agent is constrained: correctness must be moved out of the stochastic model into a versioned skill executed as code, with the model confined to interpreting heterogeneous input. This gives the field a transferable architecture for trustworthy agentic genome interpretation, a predicted route to deploying it so that every emitted answer is correct (execute the validated skill, call the input deterministically, and abstain on the irreducible residual), and a way to develop genomic skills as validated, executable, versioned units rather than prompts. Following a validation framework described elsewhere, we use clinical-grade to mean determinism, auditability, traceability to versioned components and population-invariant performance, all achieved under skill-constrained execution. We distinguish two senses of population performance: the executed clinical mapping is population-invariant by construction, verified across European, Latin American and East African origin individuals, whereas the models interpretation of real, ancestrally diverse diplotypes is not, degrading along an ancestry gradient, which is precisely why the mapping must be executed rather than reasoned. We do not claim full clinical validation, which would additionally require non-canonical inputs, real-world genomic and clinical data, human comparators and multi-site concordance.

genomics↗

Chromosome-scale assembly of the Cupressus sempervirens genome unravels new insights into the evolutionary history of conifers

Conifers, which comprise nearly two-thirds of extant gymnosperm species, are ecologically and economically important but remain genomically understudied because of their exceptionally large, repeat-rich genomes. Here, we report a chromosome-level assembly of the haploid genome of Cupressus sempervirens generated using PacBio HiFi reads and scaffolded with optical and genetic maps. The 10 Gb assembly shows exceptional contiguity for a conifer genome (contig N50 = 29.8 Mb) and was organized into 11 pseudomolecules. Iso-Seq-supported annotation identified 42,980 protein-coding genes. Repetitive elements account for over 80% of the genome, with LTR retrotransposons alone representing 52.5%. Transposable elements (TE) are pervasive in both intergenic and genic regions and have a major impact on gene architecture: TE insertions within introns generate ultra-long introns, often exceeding 100 kb, and drive gene size expansion. Analyses of LTR retrotransposon dynamics indicate that genome enlargement in C. sempervirens was driven not by recent transpositional bursts, but by the long-term accumulation and incomplete removal of ancient LTR retrotransposons. Consistent with this pattern, paleogenomic reconstruction across representative gymnosperms found no evidence of whole-genome duplication in the Cupressus lineage. This reference genome provides a valuable resource for studying conifer genome evolution, gene structure, and traits of agronomic and ecological interest, including cypress pollinosis.

genomics↗

A Telomere-to-telomere genome of the hibernating fat-tailed dwarf lemur (Cheirogaleus medius)

Madagascars dwarf lemurs (genus Cheirogaleus) are the only obligate-hibernating primates and closest relative to humans capable of hibernation. Endemic to the increasingly fragmented dry forests of Madagascar, the fat tailed dwarf lemur (Cheirogaleus medius) represents a unique model for understanding primate physiology and tropical hibernation. Here we present FatTail1, a highly contiguous diploid genome assembly generated from a male C. medius at the Duke Lemur Center using Oxford Nanopore Technologies PromethION sequencing. The assembly spans 2.3 Gb, with an N50 of 103Mb, L50 of 10, and a BUSCO completeness score of over 99%. In addition to a complete mitogenome, we generated allele-specific DNA methylation profiles and annotated 23,925 genes using NCBIs EGAPX. FatTail1 exceeds the gap-free contiguity of previously published strepsirrhine genomes, representing the first telomere-to-telomere genome of a Strepsirrhine primate, and providing a foundation for future studies of primate hibernation, epigenetic regulation, and conservation genomics. ARTICLE SUMMARYDwarf lemurs are the only primates and closest relative to humans capable of months-long hibernation, making them an important model for understanding metabolic adaptations with relevance to human physiology. Here we present FatTail1, the first telomere-to-telomere genome assembly of a strepsirrhine primate, generated from a fat-tailed dwarf lemur (Cheirogaleus medius) using Oxford Nanopore Technologies PromethION sequencing. In addition to a highly complete nuclear genome, we assembled the mitogenome, identified allele-specific DNA methylation profiles and annotated 23,925 genes. FatTail1 exceeds the gap-free contiguity of previously published strepsirrhine genomes and provides an improved genomic resource for studies of hibernation, comparative genomics, epigenetic regulation, evolutionary biology, and conservation of this threatened primate.

genomics↗