bioRxiv Science⌕ Search

bioRxiv · 10.64898/2026.07.14.738588

Toward routine health phenotyping: High-throughput prediction of metabolic, immune, and inflammatory biomarkers from milk mid-infrared spectroscopy in early-lactation dairy cows

Abstract

This study evaluated the potential of milk mid-infrared (MIR) spectroscopy, combined with routinely available on-farm variables, for predicting serum metabolic, immune, and inflammatory biomarkers in early-lactation cows. Data included 5,936 blood samples from 4,442 cows across 23 Australian dairy herds, with paired milk MIR spectra and serum measurements for up to 14 biomarkers. Prediction models were developed using partial least squares regression and evaluated using nested 10-fold random cross-validation and leave-one-herd-out validation. The results show that while basic herd-test data, including milk fat, protein, and lactose concentration, as well as on-farm variables, including DIM, calving age, breed, and herd could predict serum biomarkers, combining MIR spectra with these on-farm variables produced the best overall performance. In random cross-validation, blood urea nitrogen (BUN) was predicted most accurately (R2 = 0.78), while {beta}-hydroxybutyrate (BHB) and nonesterified fatty acids (NEFA) showed moderate accuracy (R2 = 0.56 and 0.44, respectively). BUN also showed the strongest external validation performance, with leave-one-herd-out R2 = 0.58 and comparable accuracy for predicting records collected after 70 days in milk (R2 = 0.65). BHB and NEFA had moderate leave-one-herd-out accuracy but did not transfer beyond early lactation. Most other biomarkers showed low or inconsistent external validation performance. Overall, MIR spectroscopy combined with on-farm variables shows promise for routine prediction of BUN, BHB and NEFA, which can be used for monitoring and genetic evaluation of, for example, ketosis and energy deficit. Initial random cross-validation results for glucose, bilirubin and cholesterol were promising, but more data is needed to improve the prediction accuracy and robustness of the predictions. HighlightsO_LIMilk MIR can predict several serum biomarkers in early-lactation dairy cows. C_LIO_LIMIR-predicted blood urea nitrogen shows the greatest accuracy and robustness. C_LIO_LIMIR-predicted {beta}-hydroxybutyrate and nonesterified fatty acids show moderate accuracy. C_LIO_LIMost mineral, hepatic, and inflammatory biomarkers had limited accuracy. C_LIO_LIRoutine MIR phenotyping is most promising for BUN, BHB, and NEFA. C_LI SummaryMilk mid-infrared (MIR) spectroscopy and on-farm variables are evaluated as a high-throughput tool to predict health-related serum biomarkers in early-lactation dairy cows. The data include 5,936 paired blood and milk samples from 4,442 cows across 23 Australian dairy herds and up to 14 serum biomarkers. Prediction models are developed using partial least squares regression with nested random cross-validation and leave-one-herd-out validation. Blood urea nitrogen (BUN) shows the greatest and most transferable prediction accuracy across herds and lactation stages. {beta}-hydroxybutyrate (BHB) and nonesterified fatty acids (NEFA) are predicted with moderate accuracy, but only during early lactation. Random cross-validation results for glucose and bilirubin are promising, but larger datasets are needed for robust external validation. Most other mineral, hepatic, and inflammatory biomarkers show limited external prediction accuracy. These results indicate that MIR-based routine health phenotyping is most promising for BUN, BHB, and NEFA.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ho, P., Hemsworth, J., Reich, C., Bath, C., Liu, Z., Rochfort, S., Khansefid, M., Tahir, S., HaileMariam, M., Goddard, M. E., Marett, L., Williams, R., Ho, C., Berkhout, M., Xiang, R., Chamberlain, A.. 2026-07-20. Toward routine health phenotyping: High-throughput prediction of metabolic, immune, and inflammatory biomarkers from milk mid-infrared spectroscopy in early-lactation dairy cows. https://doi.org/10.64898/2026.07.14.738588

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗

Large language model-based bibliometric evaluation of population descriptors in human genetics

As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.

genetics↗

Mitigating biases of rescaling in forward-in-time population genetic simulations

Forward-in-time population genetic simulations are widely used in evolutionary analyses, but simulating large populations and long genomic regions remains computationally demanding. To reduce this cost, parameter rescaling is widely employed, in which the original evolutionary process is approximated by one with a smaller population size and fewer generations. Recently, several studies using the SLiM simulator have raised concerns about the accuracy of this rescaling approach. In this study, we show that many of the biases reported in these studies can be mitigated by using a different simulation algorithm. These results reveal that the accuracy of parameter rescaling depends on how well the simulation algorithm preserves diffusion-limit properties under rescaling.

genetics↗