bioRxiv Science⌕ Search

Biology subjects

Caliebe, A.

Publications and source records attributed to Caliebe, A..

3 recordsLinked to original sources

TrACES of Time: Towards estimating time-of-day of bloodstain deposition by targeted RNA sequencing

BackgroundIn forensic molecular biology, the main task consists of identifying individuals who contributed to biological traces recovered from (potential) crime scenes. However, to support evidence-based reconstruction of the course of activities having taken place at the scene, contextualising information regarding how and when a biological trace was deposited is oftentimes required. ResultsHere we present the development of a forensic molecular biological analysis procedure for the prediction of the time-of-day at which a bloodstain has been deposited by targeted quantification of selected mRNA markers. Time-of-day candidate prediction markers with diurnally rhythmic expression have previously been identified by whole transcriptome sequencing. Here, we build on our previous findings by establishing a targeted cDNA sequencing protocol on an Ion S5 massively parallel sequencing device for the targeted gene expression quantification of 74 time-of-day candidate prediction markers. Based on expression measurements of these markers in 408 blood samples (from 51 individuals deposited at eight time points over a day), we establish and compare different statistical methods to predict time of deposition. The most suitable model employing penalised regression achieved a root mean squared error of 3 hours and 44 minutes with 78 % of predictions being correct within +/- 4 h (evaluated by five-fold cross-validation). ConclusionsOur study provides the first prediction model for time-of-day of bloodstain deposition based on targeted RNA sequencing and thus represents an important step towards forensic trace deposition timing. It thereby relevantly contributes to the growing knowledge on Transcriptomic Analyses for the Contextualisation of Evidential Stains (TrACES).

molecular biology↗

Neolithic introgression of IL23R-related protection against chronic inflammatory bowel diseases in modern Europeans

BackgroundThe hypomorphic variant rs11209026-A in the IL23R gene provides significant protection against immune-related diseases in Europeans, notably inflammatory bowel disease (IBD). Today, the A-allele occurs with an average frequency of 5% in Europe. MethodsThis study comprised 251 ancient genomes from Europe spanning over 14,000 years. In these samples, the investigation focused on admixture informed analyses and selection scans of rs11209026-A and its haplotypes. Findingsrs11209026-A was found at high frequencies in Anatolian Farmers (AF, 18%) where it was likely under weak positive selection. AF later introduced the allele into the ancient European gene-pool. Subsequent admixture caused its frequency to decrease and formed the current southwest-to-northeast allele frequency cline in Europe. The geographic distribution of rs11209026-A may influence the gradient in IBD incidence rates that are highest in northern and eastern Europe. InterpretationGiven the dramatic changes from hunting and gathering to agriculture during the Neolithic, AF might have been exposed to selective pressures from a pro-inflammatory lifestyle and diet. Therefore, the protective A-allele may have increased survival by reducing intestinal inflammation and microbiome dysbiosis. The adaptively evolved function of the variant likely contributes to the high efficacy and low side-effects of modern IL-23 neutralization therapies for chronic inflammatory diseases. This study highlights how evolutionary informed research can provide promising targets for new therapeutic strategies. FundingDeutsche Forschungsgemeinschaft (DFG German Research Foundation) under Germanys Excellence Strategy - EXC 2167 390884018 and EXC 2150 390870439.

evolutionary biology↗

Adaptive predictor-set linear model: an imputation-free method for linear regression prediction on datasets with missing values

Linear regression (LR) is vastly used in data analysis for continuous outcomes in biomedicine and epidemiology. Despite its popularity, LR is incompatible with missing data, which frequently occur in health sciences. For parameter estimation, this short-coming is usually resolved by complete-case analysis or imputation. Both workarounds, however, are inadequate for prediction, since they either fail to predict on incomplete records or ignore missingness-induced reduction in prediction accuracy and rely on (unrealistic) assumptions about the missing mechanism. Here, we derive adaptive predictor-set linear model (aps-lm), capable of making predictions for incomplete data without the need for imputation. It is derived by using a predictor-selection operation, the Moore-Penrose pseudoinverse and the reduced QR-decomposition. aps-lm is an LR generalization that inherently handles missing values. It is applied on a reference dataset, where complete predictors and outcome are available, and yields a set of privacy-preserving parameters. In a second stage, these are shared for making predictions of the outcome on external datasets with missing entries for predictors without imputation. Moreover, aps-lm computes prediction errors that account for the pattern of missing values even under extreme missingness. We benchmark aps-lm in a simulation study. aps-lm showed greater prediction accuracy and reduced bias compared to popular imputation strategies under a wide range of scenarios including variation of sample size, goodness-of-fit, missing value type and covariance structure. Finally, as a proof-of-principle, we apply aps-lm in the context of epigenetic aging clocks, linear models that predict a persons biological age from epigenetic data with promising clinical applications.

genetics↗