bioRxiv Science⌕ Search

Biology subjects

Fowler, E. E. E.

Publications and source records attributed to Fowler, E. E. E..

1 recordsLinked to original sources

Techniques to Produce and Evaluate Realistic Multivariate Synthetic Data

BackgroundData modeling in biomedical-healthcare research requires a sufficient sample size for exploration and reproducibility purposes. A small sample size can inhibit model performance evaluations (i.e., the small sample problem). ObjectiveA synthetic data generation technique addressing the small sample size problem is evaluated. We show: (1) from the space of arbitrarily distributed samples, a subgroup (class) has a latent multivariate normal characteristic; (2) synthetic populations (SPs) of unlimited size can be generated from this class with univariate kernel density estimation (uKDE) followed by standard normal random variable generation techniques; and (3) samples drawn from these SPs are statistically like their respective samples. MethodsThree samples (n = 667), selected pseudo-randomly, were investigated each with 10 input variables (i.e., X). uKDE (optimized with differential evolution) was used to augment the sample size in X (i.e., the input variables). The enhanced sample size was used to construct maps that produced univariate normally distributed variables in Y (mapped input variables). Principal component analysis in Y produced uncorrelated variables in T, where the univariate probability density functions (pdfs) were approximated as normal with specific variances; a given SP in T was generated with normally distributed independent random variables with these specified variances. Reversing each step produced the respective SPs in Y and X. Synthetic samples of the same size were drawn from these SPs for comparisons with their respective samples. Multiple tests were deployed: to assess univariate and multivariate normality; to compare univariate and multivariate pdfs; and to compare covariance matrices. ResultsOne sample was approximately multivariate normal in X and all samples were approximately multivariate normal in Y, permitting the generation of unlimited sized SPs. Uni/multivariate pdf and covariance comparisons (in X, Y and T) showed similarity between samples and synthetic samples. ConclusionsThe work shows that a class of multivariate samples has a latent normal characteristic; for such samples, our technique is a simplifying mechanism that offers an approximate solution to the small sample problem by generating similar synthetic data. Further studies are required to understand this latent normal class, as two samples exhibited this characteristic in the study.

bioinformatics↗