bioRxiv Science⌕ Search

Biology subjects

Marsico, F.

Publications and source records attributed to Marsico, F..

4 recordsLinked to original sources

Using all available evidence to solve kinship cases

Kinship cases, ranging from standard paternity tests to complex disaster victim identifications, are typically evaluated using likelihood ratios (LR) based on forensic genetic markers. However, in some contexts, genetic information alone is not enough to reach conclusive results. This is common when establishing distant familial connections using large DNA-databases, or even in simple cases such as determining which individual is the parent and which is the child in a relationship pair. Although forensic practitioners frequently incorporate additional evidence (SE), such as age, biological sex, or phenotypic traits, in these cases, this integration typically occurs informally, without rigorous probability estimation, compromising procedural transparency and reliability. Here, we present a comprehensive methodological framework that formally synthesizes forensic DNA evidence (FDE) with SE through Markov chain models and customized transition matrices designed for various biological traits. This approach generates combined likelihood assessments expressed as LRs or posterior probabilities. Validation through simulated and real-world case studies demonstrates that systematic incorporation of SE improves resolution accuracy in kinship determinations. To facilitate adoption, we have implemented this methodology in mispitools, an open-source R package.

genetics↗

Identity-by-descent captures Shared Environmental Factors at Biobank Scale

Genealogical ties leave traces beyond shared genetic material: relatives tend to share environmental exposures and cultural practices that persist across generations. Using genomic and electronic health record data from 13,143 individuals in the Biorepository for Integrative Genomics, we applied a hierarchical community detection algorithm to group individuals into four major communities and 17 subcommunities based on the proportion of their genome shared identical by descent (IBD). This approach captured a fine-scale demographic structure beyond traditional ancestry classifications. By integrating neighborhood-level geographic data with census-derived environmental metrics, we revealed unequal exposure to environmental stressors between subcommunities, which correlated with differential rates of health conditions, including elevated respiratory condition prevalence in subcommunities under greater pollution burden. Through an interventional mediation analysis we found that excess risk at the subcommunity-level substantially exceeded what measured environmental factors alone could explain. thus showing the contribution of unmeasured shared exposures that IBD-based grouping captures. Subcommunity-level disparities persisted in sensitivity analyses adjusting for self-reported race, indicating that IBD-based clustering captures health-relevant information beyond demographic categories commonly used as proxies for social determinants of health. We implemented an open-source dashboard for interactive exploration of IBD-defined communities, disease prevalence, and environmental exposures. Together, we demonstrate that IBD-based clustering jointly captures genetic and environmental determinants of health, offering a scalable framework for precision health and population genetics that translates biobank data into actionable insights.

genomics↗

Mispitools: An R Package for Comprehensive Statistical Methods in Kinship Inference

The search for missing persons is a complex process that involves the comparison of data from two entities: unidentified persons (UP), who may be alive or deceased, and missing persons (MP), whose whereabouts are unknown. Although existing tools support DNA-based kinship analyses for the search, they typically do not integrate or statistically evaluate diverse lines of evidence collected throughout the investigative process. Examples of alternative lines of evidence are pigmentation traits, biological sex, and age, among others. The package Mispitools fills this gap by providing comprehensive statistical methods adapted to a holistic investigation workflow. Mispitools systematically assesses the data from each investigative stage, computing the statistical weight of various types of evidence through a likelihood ratio (LR) approach. It also provides models for combining obtained LRs. Furthermore, Mispitools offers customized visualizations and a user-friendly interface, broadening its applicability among forensic practitioners and genealogical researchers.

genetics↗

Likelihood Ratios for physical traits in forensicinvestigations

Recent years have seen significant advances in DNA phenotyping, which predicts the physical traits of an unknown person, such as hair, eyes, and skin color, using DNA data. This technique is increasingly used in forensic investigations to identify missing persons, disaster victims, and suspects of crimes. A key contribution of DNA phenotyping is that it allows researchers to search through lists of individuals with similar characteristics, often gathered from testimonies, photographs, and social media data. However, despite their growing relevance, current methods lack comprehensive mathematical models to calculate likelihood ratios that accurately assess the statistical weight of evidence. Our work bridges this gap by developing new likelihood ratio models, validated through computational simulations. In addition, we demonstrate the ability of these models to improve forensic investigations in real-world scenarios. Furthermore, we introduce the R package forensicolors, freely available on CRAN, to facilitate the application of the methodologies developed.

genetics↗