bioRxiv ScienceSearch

EXPLORE THE ARCHIVE

Genomics

Find preprints about genomes, sequencing and genetic variation.

37 recordsLinked to original sources

Dog-wise canine gut metagenome assemblies with reconstructed bacterial genomes and viral candidates

Long-read metagenomic sequencing can improve genome recovery from complex gut microbial communities, yet directly reusable canine gut genome resources remain limited. Here we describe DogMAG, a canine gut metagenome resource based on dog-wise long-read and hybrid assemblies generated by grouping sequencing libraries according to canonical dog identity before assembly. The final dataset comprises 41 assemblies linked to 277 FASTQ records, including 30 Flye long-read-only and 11 OPERA-MS hybrid assemblies. A single integrated BASALT workflow produced 11,276 selected bin/version records, followed by explicit quality-based re-selection of 3,418 medium-quality-or-better metagenome-assembled genome candidates. External dRep dereplication yielded 792 strain-like representatives at 99% average nucleotide identity and 135 species/SGB-like representatives at 95%. GTDB-Tk classified all 792 representatives as Bacteria. Viral screening identified 22,068 geNomad predictions, of which 3,374 Complete, High-quality or Medium-quality viral/proviral candidate rows passed CheckV filtering with contamination [≤]10%. DogMAG provides assemblies, genome and viral candidate sequences, metadata, provenance tables and workflow scripts for reuse, benchmarking and reanalysis.

microbiology

Comparative genomics of clinical isolates of Pseudomonas aeruginosa from cystic fibrosis patients in Mexico

Pseudomonas aeruginosa (P. aeruginosa) is the primary pathogen responsible for morbidity and mortality in patients with cystic fibrosis (CF). Its genomic plasticity and constant selective pressure from antimicrobial treatments have favored the emergence of multidrug-resistant clones. This study conducted a comparative genomic analysis of 41 P. aeruginosa isolated from pediatric patients with CF in Mexico from 2015 to 2024, with the aim of characterizing their evolutionary dynamics, resistome, and virulome. Whole-genome sequencing (MGI, Illumina, and PacBio platforms) was used, with de novo assemblies performed using Unicycler v0.4.8 on the BV-BRC platform. The databases used for the resistome were CARD and NDARO, and for the virulome, VFDB. Phylogenetic reconstruction was based on core-genome alignments generated with Roary v3.13.0, with maximum likelihood reconstruction performed in IQ-TREE v2.1.2. The statistical significance of the segregation of resistance and virulence patterns was evaluated using PERMANOVA analysis. The results revealed a significant clonal prevalence of sequence types (ST) 307 and ST 167. Phylogenomic analysis grouped the isolates into three main clades; Clade 1 stood out for having the highest resistance gene load (mean of 75 genes/genome), establishing itself as the main reservoir of multidrug-resistant profiles. Genotype-phenotype concordance reached 65.5% overall, with high accuracy for aminoglycosides (87.8%) and fluoroquinolones (82.9%). Furthermore, virulome analysis identified 67 distinct patterns that were significantly segregated among the clades (PERMANOVA: R2=0.31, p=0.001). These findings demonstrate that the evolution of P. aeruginosa lineages in the pediatric clinical setting involves parallel and coordinated adaptations in both their resistance potential and their virulence arsenal. This study underscores the need to adopt a multidisciplinary approach to the clinical management of chronic P. aeruginosa infections in pediatric patients. The persistence of extensively drug-resistant (XDR) strains calls for the integration of genomic surveillance and functional diagnostics, as well as the search for therapeutic alternatives for the clinical management of patients with cystic fibrosis.

microbiology

Using sequence-to-function models to interpret archaic hominin introgression

Understanding the functional impact of archaic hominin introgression remains challenging due to the poor representation of global introgression in publicly available genomics resources. Sequence-to-function models can predict the effects of any possible variant in the human genome and may fill this gap. Here, we used AlphaGenome to predict the effects of 144,139 introgressed SNPs segregating in present-day individuals of Papuan genetic ancestry. AlphaGenome's chromatin accessibility predictions recapitulate experimentally observed effects, but gene expression performs no better than chance. Predictions correlate more strongly with an independent reporter assay of single-variant activity than with the same variants' effects in live cells, indicating that AlphaGenome captures the regulatory potential of individual variants more reliably. Predictions carry tissue specificity, allowing us to predict specific tissues potentially impacted by introgressed haplotypes. We identify genes, including JAK1 and TAB2, that are associated with haplotypes that contain an excess of variants predicted by AlphaGenome to have large impacts on chromatin accessibility. Finally, we highlight the challenges and limitations associated with using sequence-to-function models for introgressed variant effect prediction, and show that while AlphaGenome's chromatin accessibility predictions can aid in prioritising candidate functional regions, expression predictions and the assignment of variants to target genes remain as open challenges.

genomics

HIF1A recruits primate-specific endogenous retroviruses into the human hypoxic and immune responses

Oxygen availability varies profoundly across the human body and changes further during inflammation, infection, tissue injury and disease. Immune cells must therefore continuously adapt their transcriptional and metabolic state based on the oxygen availability to them. Hypoxia-inducible factor 1 (HIF1A) is central to this adaptation and a marker of the cellular response to low oxygen, yet its genomic targets have been assembled from a non-repetitive fraction of the genome, leaving nearly half of the human genome largely unexplored. Here we define the gene and transposable-element (TE) landscape of the human hypoxic response across different human tissues, cell lines, and conditions. This directional TE response was reproduced in transformed cells and in primary immune cells isolated from blood and the physiologically oxygen-restricted tonsil. Single-cell profiling of peripheral blood mononuclear cells (PBMC) under hypoxia, pharmacological HIF stabilization, and interferon stimulation revealed a striking difference between the gene and retrotranscriptome responses. While gene responses were strongly cell-type dependent and in a bidirectional manner, TEs were overwhelmingly activated. This pattern extended to blood and tonsil immune cells, where ~70-90% of tested TE families were induced under hypoxia, with activated tonsil cells showing exclusively induced significant families, including THE1B, alongside increased LTR7 and HERVH. Integrating HIF1A ChIP-seq with transcriptional responses revealed that HIF1A does not engage repetitive DNA indiscriminately. Instead, its binding converged on LTR7, the promoter long terminal repeat of the HERVH endogenous retrovirus. Approximately 80% of HIF1A-bound LTR7 elements contained a canonical hypoxia-response element, and disruption of HIF1A DNA binding dramatically reduced the expression of occupied HERVH loci. CRISPR deletion of individual LTR7/HERVH loci altered the expression of distant and neighboring genes, demonstrating that hypoxia-responsive retroelements can participate directly in host gene regulation and contribute to overall physiology. Our findings reveal the repetitive genome as a previously underappreciated component of oxygen sensing. We propose that HIF1A recruits selected endogenous retroviral elements into the human hypoxic response, extending oxygen-dependent regulation beyond conventional gene promoters and providing an additional regulatory layer through which tissue oxygenation can shape immune-cell state and human physiology.

genomics

Critical Fragility Emerges from Chromosomal Instability in Cancer

Genomic instability is a major driver of tumor evolution, promoting diversification and adaptation while simultaneously increasing the accumulation of deleterious alterations. How tumor populations balance these opposing effects remains poorly understood. Here, we introduce a computational framework that explicitly represents diploid genomes, functional gene classes, point mutations, and chromosome-segregation errors in spatially constrained and well-mixed tumor populations. We identify a viability boundary separating sustained tumor expansion from instability-induced population collapse. Within the viable regime, mutation and selection generate a stable distribution of genomic-instability classes that is accurately captured by an analytical replicator--mutator description. Near the viability boundary, tumor dynamics exhibit prolonged extinction transients and strong sensitivity to stochastic fluctuations, with important differences between solid and liquid architectures. Chromosomal alterations further modify growth by creating transient benefits through increased gene dosage and genetic redundancy, while ultimately increasing genomic fragility. Finally, simulated interventions show that eliminating low-instability subpopulations or increasing the global mutational burden can displace tumors beyond their viability boundary and trigger irreversible collapse. These results identify genome instability as both an evolutionary advantage and an intrinsic vulnerability, providing a quantitative framework for developing therapies that exploit the limits of tumor evolution.

cancer biology

Hidden molecular states of bacterial replicons beyond the chromosome-plasmid dichotomy

Bacterial genomes are organized into autonomous replicons, traditionally classified as either chromosomes or plasmids-a binary framework that underpins genome annotation and evolution models. Yet whether this binary framework captures the full diversity of replicon organization remains unclear. Here we show that bacterial replicons occupy three recurrent organizational states rather than two canonical categories. By integrating quantitative measures of chromosome-plasmid sequence affinity (plasmidness) across more than 72,000 replicons from 21 bacterial genera, we identify a distinct class-intermediate replicons-that occupies a positional and functional middle ground. These replicons are plasmid-sized, harbor substantial chromosomal sequence ancestry, and lack canonical replication signatures typically associated with either class. Multiple complementary molecular properties converge on this same state. Comparative genomic analyses reveal their enrichment near recurrent chromosome remodeling regions and reveal close evolutionary ties to conjugative and antimicrobial resistance plasmids. Metagenomic data further corroborate their presence across natural ecosystems. Together, these findings reveal a previously unrecognized replicon state and redefine bacterial genome organization beyond the chromosome-plasmid dichotomy.

microbiology

Calibration-free compression brings Evo 2 to its full million-token context on a single GPU

Evo 2 is the largest openly available genomic foundation model, but its forty billion parameter configuration cannot be loaded onto a single 80 GB accelerator, placing genome-scale analysis beyond most laboratories. We present TurboQuant-Bio, an open toolkit that compresses Evo 2s weights and attention cache to four bits without calibration data, and serves both through fused kernels. Compression is near-lossless across perplexity spanning the tree of life, genomic classification, splice-site prediction, gene completion and clinically relevant variant-effect prediction. It brings Evo 2 40B onto one 80 GB GPU and Evo 2 7B to its full million-token context within a 40 GB memory budget, an eightfold gain in reachable context. We further show that the released chunked-prefill path is silently incorrect, returning plausible but uncorrelated likelihoods, and derive the block-wise continuation that repairs it: a complete 580-kilobase bacterial genome is now scored in one context in 22 minutes rather than 13.7 hours.

bioinformatics

Evolutionary origins of protein novelty across an entire yeast subphylum

Novel protein-coding sequences fuel molecular and cellular evolutionary innovations and frequently contribute to species-defining characteristics. They can originate either de novo from previously noncoding sequences or through extreme divergence of already coding ones. How frequently each mechanism occurs and how they shape the structural and functional potential of the resulting proteins remains unclear. Here, we conducted a broad computational investigation of genetic and protein novelty throughout the entire subphylum of Saccharomycotina yeasts. We detected more than 5,000 robust de novo genes across 332 species and compared them to more than 6,000 novel genes resulting from extreme sequence divergence, revealing two quantitatively similar but qualitatively distinct modes of evolution of novelty. A remarkable 40% of de novo proteins are predicted to localize to mitochondria compared to only 20% of divergent, with the latter also being substantially longer and more disordered. A detailed analysis of conservatively predicted tertiary structures of novel proteins shows that ''invention'' of new folds occurs more frequently through de novo emergence. We also illustrate cases of evolutionary ''re-invention'' of existing protein folds from noncoding sequences. Our work deepens our understanding of the origins and importance of novel proteins, opening new directions for further structural and functional characterization.

genomics

Single-Cell Analytics for Dose Response (SCADR) discriminates PTEN missense variants by lipid and protein phosphatase dysfunction

The proliferation of sequencing efforts has revealed a vast and expanding catalog of single nucleotide gene variants, many associated to, but with unclear roles in disease. Fully charactering variant impacts and linking specific protein dysfunctions to disease are challenging due to the multi-functional nature of many proteins and varying degree of variant effects on these functions. Lagging are sensitive approaches to empirically assess the impact of missense variant-induced single amino acid changes on a wide range of protein functions. To address these issues, we have developed an open-source computational analysis tool called SCADR (Single-Cell Analytics for Dose Response) for simultaneously measuring and comparing impacts of exogenously-expressed variants on multiple signaling pathways using multiplex phospho-antibody spectral flow cytometry in human cell lines. SCADR retains and correlates single-cell measures of signal protein activity states along with expression levels of exogenously-expressed variants, providing rich characterization of multiple protein functions, signaling protein interactions, and enhanced discrimination of variant impacts on different signaling pathways, highlighting each variants unique dysfunction profile. Here, we apply SCADR for analyses of the impact of 6 variants of the tumor-suppressor protein PTEN (P38H, C124S, G129E, Y138L, D268E, 4A) expressed in HEK293 cells on the phosphorylation states of the canonical and noncanonical downstream signaling proteins Akt, S6, CREB, ERK, and p38 detected with fluorophore-conjugated phospho-antibodies, along with an antibody detecting an N-terminal HA tag on PTEN variants allowing measures of dose-response effects of each variants expression on signaling cascades. Results identify variant-specific impacts on downstream signaling cascades.

genomics

Microsecond molecular dynamics of SOD1 variants suggest a structural basis for divergent ALS clinical outcomes

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by progressive motor neuron degeneration. Mutations in the SOD1 gene represent the second most common genetic cause of ALS (ALS), and distinct SOD1 missense variants present with markedly different clinical profiles. A4V leads to an aggressive form of the disease (median survival [~]1y), H46R confers a mild, slowly progressive course and I113T exhibits an intermediate phenotype. The molecular basis by which these mutations produce divergent clinical outcomes remains poorly understood. We performed extensive classical molecular dynamics simulations of wild-type SOD1 and the three ALS-associated variants in the apo monomeric state to attempt to investigate the mechanisms behind such phenotypic differences. Structural stability, global compactness, and conformational flexibility, as well as analysis of collective motions between residues and estimation of free energy, were assessed. The H46R, A4V, and I113T variants exhibited distinct dynamic behaviours, highlighting differences in structural stability, local flexibility, and intramolecular interactions. These findings suggest that specific structural regions may contribute differently to protein dysfunction and could represent key elements for understanding the relationship between molecular dynamic properties and the differing clinical severity associated with these variants. Most strikingly, H46R exhibited exceptional structural stability across every analytical level, the lowest global deviation, most attenuated local flexibility, strongest internal dynamic coordination, and the deepest, most confined free energy basins of any system examined. This convergent multi-layered evidence of structural restraint provides a compelling mechanistic basis for the mild and slowly progressive clinical course of H46R ALS, suggesting that enhanced conformational rigidity, rather than bulk destabilisation, is the defining biophysical feature of this variant, and that its pathogenic mechanism operates through a route fundamentally decoupled from the aggregation-driven toxicity that characterises the more aggressive SOD1-ALS mutations.

genomics

The nuclear actin cytoskeleton supports DNA double-strand break repair via VCP-mediated extraction of the KU70/80 complex from damaged chromatin

Double-strand breaks (DSBs) are critical lesions in genomic DNA, and their accurate repair is essential for maintaining genome stability. The nuclear actin cytoskeleton has been implicated in homology-directed repair (HDR) of DSBs. However, the underlying mechanism remains poorly understood. Here, we report that Myosin VI (Myo6), an actin-based motor protein, cooperates with F-actin in end resection and DSB mobilization. Our findings reveal that Myo6 directly interacts with both KU70 and the ubiquitin-dependent segregase VCP to facilitate the extraction of the KU70/80 complex from chromatin. This process is supported by F-actin, revealing an interplay between nuclear actin dynamics and the DSB repair machinery. By elucidating the function of Myo6 and its direct interactions with key repair factors, our study provides mechanistic insight into how repair mechanisms rely on nuclear actin to safeguard genome integrity.

cell biology

PGM3 inhibition rewires RUVBL2-dependent DNA repair and induces a BRCAness-like state in pancreatic cancer cells

Pancreatic ductal adenocarcinoma (PDAC) exhibits profound metabolic rewiring and strong resistance to DNA-damaging therapies, yet how metabolic pathways regulate genome maintenance remains poorly understood. The hexosamine biosynthetic pathway (HBP) integrates nutrient availability with protein glycosylation through production of UDP-GlcNAc, but its role in DNA damage response (DDR) regulation is unclear. Here we show that inhibition of the HBP enzyme phosphoglucomutase-3 (PGM3) reduces DNA repair capacity in pancreatic cancer cells. Transcriptomic and functional analyses reveal that the selective PGM3 inhibitor FR054 amplifies gemcitabine-induced replication stress, disrupts ATR-CHK1 and ATM-CHK2 checkpoint signaling, and selectively impairs homologous recombination. Glycoproteomic profiling identifies the AAA+ ATPase RUVBL2 as a key metabolic-DDR node. Gemcitabine increases RUVBL2 O-GlcNAcylation, with Thr81 identified as a modified residue within the Walker A nucleotide-binding motif. Structural modelling predicts that Thr81 O-GlcNAcylation stabilizes the RUVBL1-RUVBL2 complex without compromising ATP-Mg engagement. PGM3 inhibition and Thr81 mutation similarly reduced ATR and ATM abundance and promoted persistent DNA damage, supporting a role for RUVBL2 Thr81 O-GlcNAcylation in sustaining checkpoint signalling and genome stability. Consequently, PGM3 inhibition induces a BRCAness-like state that sensitizes pancreatic cancer cells to PARP inhibition, both in vitro and in vivo, as well as to ionizing radiation. These findings reveal a nutrient-sensitive mechanism linking protein glycosylation to genome maintenance and identify HBP-dependent DNA repair as a potentially actionable vulnerability in pancreatic cancer.

cancer biology

PhageTAILor leverages machine learning for phage tail-like elements detection and classification in plant-associated bacteria

Phage tail-like elements (PTEs) -- tailocins, bacterial type VI secretion systems (T6SS), and extracellular contractile injection systems (eCIS) -- are contractile nanomachines that bacteria use to kill their neighbors and compete within their micro-ecosystems. PTEs help shape microbial community composition. Most PTE detection tools only detect a single PTE class. Moreover, most tailocin detection methods are largely restricted to Pseudomonas, leaving a key part of tailocin diversity uncharacterized. In this work, we present PhageTAILor (https://github.com/hjcho-bio/PhageTAILor), an integrative and fully automated pipeline that detects and classifies prophages and 3 PTE classes from bacterial genomes. PhageTAILor combines a 6-detector homology-based candidate search (geNomad, tail-gene, PHROGs-tail, SecReT6, eCIStem, and a divergence-tolerant tail-HMM detector) with a LightGBM classifier comprising 1 multiclass and 3 binary heads, trained on 6,501 bacterial genomes carrying 13,082 prophages and PTEs. A phylogeny-free feature matrix used in our model keeps predictions reproducible between model construction and user inference. PhageTAILor performs strongly at the genome level and generalizes beyond its Pseudomonas-rich training set. On a 76-strain cross-clade benchmark, PhageTAILor detected tailocins at F1 = 0.955. Furthermore, it identified 12 of 13 experimentally validated tailocins spanning five genera versus 2 of 13 for a Pseudomonas-restricted tool TattleTail. PhageTAILor also demonstrated sensitivity equivalent to viral detection tool geNomad while avoiding its higher false-positive rate. Applied to 7,925 plant- and soil-associated bacterial isolates, PhageTAILor showed that prophages in the phyllosphere and tailocins in plant-associated bacteria, whereas eCIS are enriched in soil. PhageTAILor is distributed as an open-source, modular pipeline with a command-line interface.

microbiology

Absence of a spindle position checkpoint in the fungal pathogen Cryptococcus neoformans

To maintain genome stability, it is crucial that cells do not initiate cytokinesis until chromosomes have been properly segregated. In the model budding yeast Saccharomyces cerevisiae, a surveillance mechanism called the Spindle Position Checkpoint (SPoC) ensures this coordination by regulating the Mitotic Exit Network (MEN) to couple exit from mitosis and cytokinesis to spindle position. The MEN is conserved in Ascomycota where the orthologous pathway in the fission yeast Schizosaccharomyces pombe, the Septation Initiation Network (SIN), regulates cytokinesis in response to defects in spindle elongation. Here, we show that the MEN/SIN pathway is conserved in the basidiomycetous budding yeast and human pathogen, Cryptococcus neoformans, and controls cytokinesis. However, spindle position or elongation does not regulate pathway activation or cell cycle progression in C. neoformans. In essence, there appears to be no SPoC in this organism to delay cytokinesis upon defects in mitosis. We speculate that while increasing the risk of genome instability, the lack of a SPoC might facilitate C. neoformans's ability to change ploidy in the host.

cell biology

Parallel evolution under constraint shapes echinocandin resistance in Candida auris

Drug resistance emerges repeatedly in outbreaks of Candida fungal pathogens, but little is known about its origins or persistence. Here, we investigated the evolutionary processes shaping echinocandin resistance in Candida auris, a globally emerging and predominantly clonal fungal pathogen. Genome-wide association across over 600 isolates identified mutations in the {beta}-1,3-glucan synthase gene FKS1 as the most significant driver of resistance to an echinocandin drug. Ancestral reconstruction of this population traced shared resistance mutations among small groups typically consisting of 2-3 closely related isolates, but clusters could include up to 16 isolates. Nearly all resistant clusters consisted of isolates collected in the same year and region, consistent with local transmission. To further examine population-level selection, we measured adaptive signatures in FKS1 and the highly diverged paralog FKS2 across 22,000 genomes. This revealed excess nonsynonymous polymorphisms in FKS1, primarily due to independent, recurrent mutations at resistance hotspots, consistent with parallel evolution and incomplete fixation of adaptive alleles. In FKS2, there is no evidence of hotspots and little support for diversifying selection. Together, these results indicate that resistance mutations emerge under strong genetic constraint, with adaptation restricted to only one FKS homolog and predominantly at mutational hotspots.

genetics

Kaposi's sarcoma-associated herpesvirus forms and maintains R-loops at origins of lytic replication

GC-rich sequences are abundant in human herpesviruses genomes. GC-rich regions can form three-stranded RNA:DNA hybrid structures called R-loops. Though these hybrid structures serve important biological roles at telomeres or during cellular DNA synthesis, unscheduled or prolonged R-loop formation causes DNA damage and genome instability. For this reason, several mechanisms exist to resolve R-loops including endoribonucleases RNaseH1 (constitutively expressed) and RNaseH2A (cell cycle-regulated) which degrade the RNA portion of the R-loop. The Kaposi's sarcoma-associated herpesvirus (KSHV) origins of lytic replication (OriLyts) contain multiple cis-acting elements that are required for viral DNA replication including the production of GC-rich and repetitive transcripts, T1.4 (OriLyt-L) and kaposin (OriLyt-R). We previously showed that R-loops form at both OriLyts and that deleting kaposin repeats or decreasing their GC-rich content prevented R-loop formation at OriLyt-R, reduced genome amplification after primary infection and caused defects in latency establishment. To define the contribution that R-loops play in KSHV replication, we overexpressed RNaseH1, reasoning that excess RNaseH1 would resolve both OriLyt R-loops. However, RNaseH1 protein levels decreased following KSHV reactivation in both iSLK and BCBL-1 cell lines. Using co-transfection, we discovered that the KSHV viral replication and transcription activator protein, RTA, mediated RNaseH1 protein decreases in a E3 ligase domain-dependent manner without impacting levels of its cognate RNA transcript. We attempted to construct an RTA-resistant yet functional version of RNaseH1 by site-directed mutagenesis of lysine residues individually or in combination, yet these constructs remain susceptible to RTA-mediated protein decreases. An amino terminally tagged RNaseH1 displayed reduced susceptibility to RTA, suggesting that RTA may target the N-terminus of RNaseH1 for ubiquitination. However, overexpression of the cell-cycle regulated endonuclease, RNaseH2, exhibited RTA resistance, suggesting RNaseH2 may be a tool that will effectively resolve R-loops during KSHV infection. KSHV is not the only herpesvirus to encode a protein that reduces RNaseH1 levels, as co-expression of RTA homologs from the related gamma-herpesviruses EBV and MHV-68 likewise decreased steady-state levels of RNaseH1 protein. We propose that RTA-mediated RNaseH1 degradation is conserved strategy to ensure R-loop persistence during gamma-herpesvirus infection, underscoring the importance of these structures.

microbiology

AmPair: automating housekeeping-gene primer design for species-level metataxonomics

Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.

bioinformatics

A replicated patient-specific component of tumour telomere length across two pan-cancer cohorts

Bulk telomere length measured from tumour sequencing is routinely interpreted as a property of the cancer cells. However, a tumour specimen is a mixture, and the patient who supplies it has a telomere length of their own. Here I re-analyse published pan-cancer telomere estimates and ask how much of a tumour's telomere length is patient-specific. A calibration step comes first. Whole-genome and low-pass estimates recover the known cross-sectional attrition of leukocyte telomeres with age, at 26.6 bp per year in blood normals, whereas whole-exome estimates do not. After adjustment for cancer type, sequencing centre and sex, the exome slope is minus 0.6 bp per year. In 684 blood-normal aliquots sequenced by both assays, the whole-genome estimate declines at 38.9 bp per year, whereas the exome estimate from the same DNA shows no detectable decline. The difference between assays is 41.5 bp per year, with P = 3 x 10^-10. Because exome data constitute 78.6% of the original resource, downstream analyses use only whole-genome and low-pass libraries. Within those data, tumour telomere length tracks the patient's matched-normal telomere length. The Spearman correlation is 0.395 in TCGA, with positive associations in 22 of 23 cancer types. This finding replicates in PCAWG using a different telomere estimator, with a correlation of 0.472 and positive associations in all 24 histologies examined. Adjustment for cancer type, sequencing centre and library type leaves a regression coefficient of 0.385. The association is also stable after adjustment for age, sex, tumour purity, leukocyte fraction, ploidy, sequencing coverage and continental ancestry, with coefficients ranging from 0.406 to 0.429. Pure normal-cell admixture is rejected as the sole explanation. Under a two-compartment mixture model, the coefficient for host telomere length is expected to equal 1 and the host-by-purity interaction to equal minus 1. These restrictions are jointly rejected with P = 0.001. Tumour purity, leukocyte fraction and age each explain only about 1 to 3% of within-cohort variance and do not alter the cross-cancer ranking. By contrast, the between-cohort coefficient is not directly interpretable. Its apparent near one-to-one relationship with tissue-associated telomere length depends strongly on which tissue supplies the matched-normal reference and on the statistical spread of that predictor, falling to 0.44 when organ-matched solid tissue is used. Bulk tumour telomere length is therefore a composite phenotype containing a replicated patient-specific component. Telomere biomarker studies should include matched-normal telomere length as a covariate rather than treating tumour telomere length as exclusively tumour-intrinsic.

cancer biology