bioRxiv Science⌕ Search

Biology subjects

Vidaki, A.

Publications and source records attributed to Vidaki, A..

4 recordsLinked to original sources

qBiCo: A method to assess global DNA conversion performance in epigenetics via single-copy genes and repetitive elements

Human DNA methylation profiling offers great promises in various biomedical applications, including ageing, cancer and even forensics. So far, most DNA methylation techniques are based on a chemical process called sodium bisulfite conversion, which specifically converts non-methylated cytosines into uracils. However, despite the popularity of this approach, it is known to cause DNA fragmentation and loss affecting standardization, while incomplete conversion may result in potential misinterpretation of methylation-based outcomes. To offer the community a solution, we developed qBiCo - a novel quality-control method to address the quantity and quality of bisulfite-converted DNA. qBiCo is a 5-plex, TaqMan(R) probe-based, quantitative (q)PCR assay that amplifies single- and multi-copy DNA fragments of converted and non-converted nature. It estimates four parameters: converted DNA concentration, fragmentation, global conversion efficiency, and potential PCR inhibition. We optimized qBiCo using synthetic DNA standards and assessed it using standard developmental validation criteria, showcasing that qBiCo is reliable, robust and sensitive down to picogram level. We also evaluated its performance by testing decreasing DNA amounts using several commercial bisulfite conversion kits. Depending on the starting DNA quantity, bisulfite-converted DNA recoveries ranged from 8.5 - 100 %, conversion efficiencies from 78 - 99.9 %, while certain kits highly fragment DNA, demonstrating large variability in their performance. Towards building a prototype tool, we further optimized key functionalities, for example, by replacing the poorest performing single-plex assay and creating a more representative DNA standard. Aiming to scale-up and move towards implementation, we successfully transferred and validated our novel method in six different qPCR platforms from different major manufacturers. Overall, with the present study, we offer researchers in the epigenetic field a novel long-awaited QC tool that for the first time allows them to measure key quality and quantity parameters of the most popular DNA conversion process. The tool also enables standardization to prevent inconsistent data and false outcomes in the future, regardless of the downstream experimental analysis of DNA methylation-based research and applications across different fields of biology and biomedicine.

genetics↗

Adaptive predictor-set linear model: an imputation-free method for linear regression prediction on datasets with missing values

Linear regression (LR) is vastly used in data analysis for continuous outcomes in biomedicine and epidemiology. Despite its popularity, LR is incompatible with missing data, which frequently occur in health sciences. For parameter estimation, this short-coming is usually resolved by complete-case analysis or imputation. Both workarounds, however, are inadequate for prediction, since they either fail to predict on incomplete records or ignore missingness-induced reduction in prediction accuracy and rely on (unrealistic) assumptions about the missing mechanism. Here, we derive adaptive predictor-set linear model (aps-lm), capable of making predictions for incomplete data without the need for imputation. It is derived by using a predictor-selection operation, the Moore-Penrose pseudoinverse and the reduced QR-decomposition. aps-lm is an LR generalization that inherently handles missing values. It is applied on a reference dataset, where complete predictors and outcome are available, and yields a set of privacy-preserving parameters. In a second stage, these are shared for making predictions of the outcome on external datasets with missing entries for predictors without imputation. Moreover, aps-lm computes prediction errors that account for the pattern of missing values even under extreme missingness. We benchmark aps-lm in a simulation study. aps-lm showed greater prediction accuracy and reduced bias compared to popular imputation strategies under a wide range of scenarios including variation of sample size, goodness-of-fit, missing value type and covariance structure. Finally, as a proof-of-principle, we apply aps-lm in the context of epigenetic aging clocks, linear models that predict a persons biological age from epigenetic data with promising clinical applications.

genetics↗

Simultaneous mapping of epigenetic inter-haplotype, inter-cell and inter-individual variation via the discovery of jointly regulated CpGs in pooled sequencing data

In the post-GWAS era, great interest has arisen in the mapping of epigenetic inter-individual variation towards investigating the emergence of phenotype in health and disease. Relevant DNA methylation methodologies - epigenome-wide association studies (EWAS), methylation quantitative trait loci (mQTL) mapping and allele-specific methylation (ASM) analysis - can each map certain sources of epigenetic variation and all depend on matching phenotypic/genotypic data. Here, to avoid these requirements, we developed Binokulars, a novel randomization test that identifies signatures of joint CpG regulation from reads spanning multiple CpGs. We tested and benchmarked our novel approach against EWAS and ASM on pooled whole-genome bisulfite sequencing (WGBS) data from whole blood, sperm and combined. As a result, Binokulars simultaneously discovered regions associated with imprinting, cell type- and tissue-specific regulation, mQTL, ageing and other (still unknown) epigenetic processes. To verify examples of mQTL and polymorphic imprinting, we developed JRC_sorter, another novel tool that classifies regions based on epigenotype models, which we deployed on non-pooled WGBS data from cord blood. In the future, this approach can be applied on larger pools to simultaneously map and characterise inter-haplotype, inter-cell and inter-individual variation in DNA methylation in a cost-effective fashion, a relevant pursuit towards phenome-mapping in the post-GWAS era.

genomics↗

Impact of SNP microarray analysis of compromised DNA on kinship classification success in the context of investigative genetic genealogy

Single nucleotide polymorphism (SNP) data generated with microarray technologies have been used to solve murder cases via investigative leads obtained from identifying relatives of the unknown perpetrator included in accessible genomic databases, referred to as investigative genetic genealogy (IGG). However, SNP microarrays were developed for relatively high input DNA quantity and quality, while SNP microarray data from compromised DNA typically obtainable from crime scene stains are largely missing. By applying the Illumina Global Screening Array (GSA) to 264 DNA samples with systematically altered quantity and quality, we empirically tested the impact of SNP microarray analysis of deprecated DNA on kinship classification success, as relevant in IGG. Reference data from manufacturer-recommended input DNA quality and quantity were used to estimate genotype accuracy in the compromised DNA samples and for simulating data of different degree relatives. Although stepwise decrease of input DNA amount from 200 nanogram to 6.25 picogram led to decreased SNP call rates and increased genotyping errors, kinship classification success did not decrease down to 250 picogram for siblings and 1st cousins, 1 nanogram for 2nd cousins, while at 25 picogram and below kinship classification success was zero. Stepwise decrease of input DNA quality via increased DNA fragmentation resulted in the decrease of genotyping accuracy as well as kinship classification success, which went down to zero at the average DNA fragment size of 150 base pairs. Combining decreased DNA quantity and quality in mock casework and skeletal samples further highlighted possibilities and limitations. Overall, GSA analysis achieved maximal kinship classification success from 800-200 times lower input DNA quantities than manufacturer-recommended, although DNA quality plays a key role too, while compromised DNA produced false negative kinship classifications rather than false positive ones. Author SummaryInvestigative genetic genealogy (IGG), i.e., identifying unknown perpetrators of crime via genomic database-tracing of their relatives by means of microarray-based single nucleotide polymorphism (SNP) data, is a recently emerging field. However, SNP microarrays were developed for much higher DNA quantity and quality than typically available from crime scenes, while SNP microarray data on quality and quantity compromised DNA are largely missing. As first attempt to investigate how SNP microarray analysis of quantity and quality compromised DNA impacts kinship classification success in the context of IGG, we performed systematic SNP microarray analyses on DNA samples below the manufacturer-recommended quantity and quality as well as on mock casework samples and on skeletal remains. In addition to IGG, our results are also relevant for any SNP microarray analysis of compromised DNA, such as for the DNA prediction of appearance and biogeographic ancestry in forensics and anthropology and for other purposes.

genetics↗