bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.08.05.743001

Near-infrared phenomic and genomic prediction for seed protein in winter legume white lupin (Lupinus albus L.): A utility comparison

Abstract

White lupin (Lupinus albus L.) is a cool-season grain legume with seed crude protein of 33-47%, competitive with soybean (Glycine max L.) meal. It also fixes nitrogen and mobilizes soil phosphorus. Because soybean is a summer crop, white lupin can occupy Southeastern winter fields as a complementary protein source. Breeding for seed protein is limited by the cost and throughput of reference phenotyping. To determine how each is best deployed, we compared the utility of near-infrared spectroscopy (NIRS)-based phenomic selection with genomic selection based on 246,847 SNPs from low-pass, whole genome sequencing in a panel of Auburn University breeding lines and USDA National Plant Germplasm System germplasm. A handheld NIR calibration against Dumas reference protein reached screening-grade accuracy (R2 = 0.81). Under common cross-validation, phenomic predictive ability was 0.93 and genomic was 0.12. The low genomic value was consistent with moderate heritability (H2 = 0.33) and strong genotype-by-year interaction. Beyond predictive ability, NIRS recovered superior accessions the strictest selection intensity, and 40 to 60 reference assays sufficed to calibrate the model. Handheld NIRS is a low-cost tool for protein calibration and early-generation screening, while genomic prediction remains suited to parental selection, together supporting a complementary strategy for legume breeding Plain Language SummarySoybean meal is the main protein source for livestock and fish farms in the United States. Because soybean is a summer crop, many Southeastern fields sit idle or grow low-value cover crops in winter. White lupin, a cool-season legume whose seeds are as protein-rich as soybean meal, makes a good complementary winter crop: it yields high-protein grain while serving as a cover crop that fixes nitrogen and frees up soil phosphorus for later crops. In our early-stage lupin breeding program, measuring seed protein by standard lab methods is slow and costly. We built a calibration that lets a handheld scanner estimate protein from light, and compared it with predicting protein from the plants DNA. The scanner gave accurate, low-cost protein screening from only about 40-60 lab tests, while DNA-based prediction remains suited to guiding parent selection. Used together, these tools offer breeders a practical path to develop high-protein white lupin. Core ideasO_LIHandheld NIRS provides screening-grade prediction of white lupin seed crude protein. C_LIO_LISpectra carried more usable protein signal than markers by measuring seed chemistry directly. C_LIO_LINIRS and genomic prediction serve different stages of a white lupin breeding program. C_LIO_LIAbout 40 to 60 reference assays sufficed to calibrate NIRS to near-full accuracy. C_LI

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Castillo, M. P., Oyebode, O. G., Lenahan, A., Orloski, A., Wolfe, M.. 2026-08-11. Near-infrared phenomic and genomic prediction for seed protein in winter legume white lupin (Lupinus albus L.): A utility comparison. https://doi.org/10.64898/2026.08.05.743001

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

PfPHAST: Plasmodium falciparum Public Health Amplicon Sequencing Tool, a Streamlined Panel for Malaria Genomic Surveillance

Genomic tools can support malaria control policy through surveillance of Plasmodium falciparum populations, tracking antimalarial drug resistance, pfhrp2/3 deletions that compromise rapid diagnostic tests, and selection at the circumsporozoite protein (PfCSP) vaccine target, as well as through molecular correction of therapeutic efficacy studies (TES). Multiplex Amplicons for Drug, Diagnostic, Diversity, and Differentiation Haplotypes using Targeted Resequencing (MAD4HatTeR), a comprehensive amplicon sequencing panel covering up to 276 targets, supports these applications but is tailored to research rather than routine programmatic use. We developed P. falciparum Public Health Amplicon Sequencing Tool (PfPHAST), a 56-target derivative of MAD4HatTeR spanning drug resistance loci, pfhrp2/3 deletion, PfCSP genotyping, non-falciparum species identification, and 20 high-heterozygosity microhaplotype loci for TES classification. We compared PfPHAST and MAD4HatTeR using laboratory strain controls, including two-strain dilution series and a five-strain mixture, across parasite densities of 100 to 10,000 parasites/L. At matched per-target depth, PfPHAST achieved a higher quality-control pass rate than MAD4HatTeR (94.4% versus 90.0%) and distributed reads more evenly across targets. The panels showed comparable recall and precision for drug resistance codons and microhaplotypes, reaching near-complete recall above 40% within-sample allele frequency (WSAF) at all densities, with reduced sensitivity for minor alleles below 10% WSAF at low parasite density in both panels. Observed and expected WSAF correlated strongly for both panels, and both resolved a five-strain polyclonal mixture, including a 5% minor strain. By concentrating sequencing capacity on targets of greatest programmatic relevance, PfPHAST offers a scalable, lower-cost alternative to comprehensive research panels without sacrificing performance on shared targets, complementing MAD4HatTeR for routine molecular malaria surveillance.

genomics↗

Structural variation in repeat elements is widespread in normal human tissues and in tumorigenesis

Somatic mosaicism contributes to genomic variation, yet postzygotic structural variants remain under-characterized. We performed long- and short-read WGS from multiple individuals (n=47 normal tissues; n=168 samples) and identified mosaic structural variants in all individuals and germ layers, impacting a median 285.2 kb/genome. Nearly half of breakpoints were independently validated, with tissue distributions reflecting both early and late developmental origins. Most mosaic variants were repeat-mediated and 8.3% overlapped functional elements, an enrichment compared to germline variants. To extend these analyses in samples where long-read sequencing is infeasible, we measured repeat alterations from short-read sequencing, recapitulating mosaic tissue-specific differences. We characterized tumor- and tissue- specific variation in repeats across 15 cancer types and found tumor-related repeat variation to be similar in scale to that of normal mosaic variation. Tracking repeat changes in cell-free DNA provided a noninvasive approach for tumor monitoring. Our analyses revealed widespread repeat-driven structural variation in health and disease.

genomics↗

RNA isoform-resolved multiplexed sequencing with bioorthogonal barcoding

RNA isoform dysregulation drives disease pathogenesis and is the target of FDA-approved splice-switching therapeutics. However, multiplexed sequencing methods discard splice junction information because only 3' termini are barcoded and counted. Here, we repurpose acylation and click chemistries to conjugate bioorthogonal barcodes (bobcodes) directly onto multiple internal positions along cellular RNAs. Bobcoded RNAs from multiple samples are pooled for multiplexed cDNA synthesis, during which reverse transcriptase switches from each RNA template onto its tethered bobcode with greater than 99% accuracy in species mixing experiments. Bobcode attachment intervals set cDNA insert sizes without a library fragmentation step, and priming with poly(dT) or random hexamers selects between 3'-end counting and full-length isoform capture. A bioorthogonal barcode-sequencing (BOB-seq v0.1) drug screen identifies transcriptome-wide on- and off-target RNA splicing effects and outperforms existing multiplexing RNA sequencing methods in workflow simplicity, sample-to-sample variability, and barcoding accuracy. Bobcodes add isoform resolution to scalable multiplexed RNA sequencing.

genomics↗