Sampling bias in large healthcare claims databases
Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum's de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.
bioinformatics↗