bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

ORAL HEALTH BEHAVIOUR PERCEPTION SCALE APPLIED AMONG A SAMPLE OF PORTUGUESE ADOLESCENTS

IntroductionThe application of a scale can be particularly useful for the epidemiological studies comparing different populations and for analysis of the influence of distinct aspects of oral health on the development of certain health conditions. The aim of this study consists in the creation of a scale to classify the level of perception of the oral health behaviors applicable to the sample of Portuguese adolescents.\n\nMaterials and methodsAn observational cross-sectional study was designed with a total of 649 adolescents between the ages of 12 and 18 years old from five public schools in the Viseu and Guarda districts, in Portugal. Data was collected by the application of a self-administered questionnaire and, after analysis of data collection, the newly Universidade Catolica Portuguesa (UCP) oral health perception scale was created.\n\nResultsAnalyzing the sample included in the present study, we verified, by the UCP oral health perception scale created, that 67.9% of the sample presented a poor perception of their oral health behaviors, 23.9% intermediate/sufficient and only 8.2% presented what is considered in the scale as having a good classification in terms of oral health behaviors perception respecting the assumptions defined for the elaboration of the present scale.\n\nConclusionsFor this purpose, through the scale to classify the level of oral health behaviors applicable to the sample of portuguese adolescents, it is possible to compare the data of several samples and understand what are the most frequent oral or eating habits among adolescents.

scientific communication and education

Assessing the potential of plains zebra to maintain African horse sickness in the Western Cape Province, South Africa

African horse sickness (AHS) is a disease of equids that results in a non-tariff barrier to the trade of live equids from affected countries. AHS is endemic in South Africa except for a controlled area in the Western Cape Province (WCP) where sporadic outbreaks have occurred in the past 2 decades. There is potential that the presence of zebra populations, thought to be the natural reservoir hosts for AHS, in the WCP could maintain AHS virus circulation in the area and act as a year-round source of infection for horses. However, it remains unclear whether the epidemiology or the ecological conditions present in the WCP would enable persistent circulation of AHS in the local zebra populations.\n\nHere we developed a hybrid deterministic-stochastic vector-host compartmental model of AHS transmission in plains zebra (Equus quagga), where host populations are age- and sex-structured and for which population and AHS transmission dynamics are modulated by rainfall and temperature conditions. Using this model, we showed that populations of plains zebra present in the WCP are not sufficiently large for AHS introduction events to become endemic and that coastal populations of zebra need to be >2500 individuals for AHS to persist >2 years, even if zebras are infectious for more than 50 days. AHS cannot become endemic in the coastal population of the WCP unless the zebra population involves at least 50,000 individuals. Finally, inland populations of plains zebra in the WCP may represent a risk for AHS to persist but would require populations of at least 500 zebras or show unrealistic duration of infectiousness for AHS introduction events to become endemic.\n\nOur results provide evidence that the risk of AHS persistence from a single introduction event in a given plains zebra population in the WCP is extremely low and it is unlikely to represent a long-term source of infection for local horses.

ecology

Computational Pan-genome Mapping and pairwise SNP-distance improve Detection of Mycobacterium tuberculosis Transmission Clusters

Next-generation sequencing based base-by-base distance measures have become an integral complement to epidemiological investigation of infectious disease outbreaks. This study introduces PANPASCO, a computational pan-genome mapping based, pairwise distance method that is highly sensitive to differences between cases, even when located in regions of lineage specific reference genomes. We show that our approach is superior to previously published methods in several datasets and across different Mycobacterium tuberculosis lineages, as its characteristics allow the comparison of a high number of diverse samples in one analysis - a scenario that becomes more and more likely with the increased usage of whole-genome sequencing in transmission surveillance.\n\nAuthor summaryTuberculosis still is a threat to global health. It is essential to detect and interrupt transmissions to stop the spread of this infectious disease. With the rising use of next-generation sequencing methods, its application in the surveillance of Mycobacterium tuberculosis has become increasingly important in the last years. The main goal of molecular surveillance is the identification of patient-patient transmission and cluster detection. The mutation rate of M. tuberculosis is very low and stable. Therefore, many existing methods for comparative analysis of isolates provide inadequate results since their resolution is too limited. There is a need for a method that takes every detectable difference into account. We developed PANPASCO, a novel approach for comparing pairs of isolates using all genomic information available for each pair. We combine improved SNP-distance calculation with the use of a pan-genome incorporating more than 100 M. tuberculosis reference genomes for read mapping prior to variant detection. We thereby enable the collective analysis and comparison of similar and diverse isolates associated with different M. tuberculosis strains.

bioinformatics

Sensitivity and robustness of comorbidity network analysis

1Summary and KeywordsO_ST_ABSBackgroundC_ST_ABSComorbidity network analysis (CNA) is an increasingly popular approach in systems medicine, in which mathematical graphs encode epidemiological correlations (links) between diseases (nodes) inferred from their occurrence in an underlying patient population. A variety of methods have been used to infer properties of the constituent diseases or underlying populations from the network structure, but few have been validated or reproduced.\n\nObjectivesTo test the robustness and sensitivity of several common CNA techniques to the source of population health data and the method of link determination.\n\nMethodsWe obtained six sources of aggregated disease co-occurrence data, coded using varied ontologies, most of which were provided by the authors of CNAs. We constructed families of comorbidity networks from these data sets, in which links were determined using a range of statistical thresholds and measures of association. We calculated degree distributions, single-value statistics, and centrality rankings for these networks and evaluated their sensitivity to the source of data and link determination parameters. From two open-access sources of patient-level data, we constructed comorbidity networks using several multivariate models in addition to comparable pairwise models and evaluated differences between correlation estimates and network structure.\n\nResultsGlobal network statistics vary widely depending on the underlying population. Much of this variation is due to network density, which for our six data sets ranged over three orders of magnitude. The statistical threshold for link determination also had strong effects on global statistics, though at any fixed threshold the same patterns distinguished our six populations. The association measure used to quantify comorbid relations had smaller but discernible effects on global structure. Co-occurrence rates estimated using multivariate models were increasingly negative-shifted as models accounted for more effects. However, only associations between the most prevalent disorders were consistent from model to model. Centrality rankings were likewise similar when based on the same dataset using different constructions; but they were difficult to compare, and very different when comparable, between data sets, especially those using different ontologies. The most central disease codes were particular to the underlying populations and were often broad categories, injuries, or non-specific symptoms.\n\nConclusionsCNAs can improve robustness and comparability by accounting for known limitations. In particular, we urge comorbidity network analysts (a) to include, where permissible, disaggregated disease occurrence data to allow more targeted reproduction and comparison of results; (b) to report differences in results obtained using different association measures, including both one of relative risk and one of correlation; (c) when identifying centrally located disorders, to carefully decide the most suitable ontology for this purpose; and, (d) when relevant to the interpretation of results, to compare them to those obtained using a multivariate model.

bioinformatics

Winter is coming: pathogen emergence in seasonal environments

Many infectious diseases exhibit seasonal dynamics driven by periodic fluctuations of the environment. Predicting the risk of pathogen emergence at different points in time is key for the development of effective public health strategies. Here we study the impact of seasonality on the probability of emergence of directly transmitted pathogens under different epidemiological scenarios. We show that when the period of the fluctuation is large relative to the duration of the infection, the probability of emergence varies dramatically with the time at which the pathogen is introduced in the host population. In particular, we identify a new effect of seasonality (the winter is coming effect) where the probability of emergence is vanishingly small even though pathogen transmission is high. We use this theoretical framework to compare the impact of different control strategies on the average probability of emergence. We show that, when pathogen eradication is not attainable, the optimal strategy is to act intensively in a narrow time interval. Interestingly, the optimal control strategy is not always the strategy minimizing R0, the basic reproduction ratio of the pathogen. This theoretical framework is extended to study the probability of emergence of vector borne diseases in seasonal environments and we show how it can be used to improve risk maps of Zika virus emergence.

ecology

Toxicants Associated with Spontaneous Abortion in the Comparative Toxicogenomics Database (CTD)

BackgroundUp to 70% of all pregnancies result in either implantation failure or spontaneous abortion (SA). Many events occur before women are aware of their pregnancy and we lack a comprehensive understanding of high-risk SA chemicals. In epidemiologic research, failure to account for a toxicants impact on SA can also bias toxicant-birth outcome associations. Our goal was to identify chemicals with a high number of interactions with SA genes, based on known toxicogenomic responses.\n\nMethodsWe used reference SA (MeSH: D000022) and chemical gene lists from the Comparative Toxicogenomics Database in three species (human, mouse, and rat). We prioritized chemicals (n=25) found in maternal blood/urine samples or in groundwater, tap water, or Superfund sites. For chemical-disease gene sets of sufficient size (n=13 chemicals, n=20 comparisons), chi-squared enrichment tests and proportional reporting ratios (PRR) were calculated. We then cross-validated enrichment results. Finally, among the SA genes, we assessed enrichment for gene ontology biological processes and for chemicals associated with SA in humans, we visualized specific gene-chemical interactions.\n\nResultsThe number of unique genes annotated to a chemical ranged from 2 (bromacil) to 5,607 (atrazine), and 121 genes were annotated to SA. In humans, all chemicals tested were highly enriched for SA gene overlap (all p<0.001; parathion PRR=7, cadmium PRR=6.5, lead PRR=3.9, arsenic PRR=3.5, atrazine PRR=2.8). In mice, highest enrichment (p<0.001) was observed for naphthalene (PRR=16.1), cadmium (PRR=12.8), arsenic (PRR=11.6), and carbon tetrachloride (PRR=7.7). In rats, we observed highest enrichment (p<0.001) for cadmium (PRR=8.7), carbon tetrachloride (PRR=8.3), and dieldrin (PRR=5.3). Our findings were robust to 1,000 permutations each of gene sets ranging in size from 100 to 10,000. SA genes were overrepresented in biological processes: inflammatory response (q=0.001), collagen metabolic process (q=1x10-13), cell death (q=0.02), and vascular development (q=0.005).\n\nConclusionWe observed chemical gene sets (parathion, cadmium, naphthalene, carbon tetrachloride, arsenic, lead, dieldrin, and atrazine) were highly enriched for SA genes. Exposures to chemicals linked to SA, thus linked to probability of live birth, may deplete fetuses susceptible to adverse birth outcomes. Our findings have critical public health implications for successful pregnancies as well as the interpretation of environmental pregnancy cohort analyses.

pharmacology and toxicology

Compilation of 29-year postmortem examinations identifies a \"Millennium bug\" in equine parasite communities

Horses are infected by a wide range of parasite species that form complex communities. Parasite control imposes significant constraints on parasite communities whose monitoring remains however difficult to track through time. Postmortem examination is a reliable method to quantify parasite communities. Here, we compiled 1,673 necropsy reports accumulated over 29 years, in the reference necropsy centre from Normandy (France). The burden of non-strongylid species was quantified and the presence of strongylid species was noted. Details of horse deworming history and the cause of death were registered. Building on these data, we investigated the temporal trend in non-strongylids epidemiology and we determined the contribution of parasites to the death of horses throughout the study period. Data analyses revealed the seasonal variations of non-strongylid parasite abundance and reduced worm burden in race horses. Beyond these observations, we found a shift in the species responsible for fatal parasitic infection from the year 2000 onward, whereby fatal cyathostominosis and Parascaris spp. infection have replaced death cases caused by S. vulgaris and tapeworms. Concomitant break in the temporal trend of parasite species prevalence was also found within a 10-year window (1998-2007) that has seen the rise of Parascaris spp. and the decline of both Gasterophilus spp. and tapeworms. A few cases of parasite persistence following deworming were identified that all occurred after 2000. Altogether, these findings provide insights into major shifts in non-strongylid parasite prevalence and abundance over the last 29 years. They also underscore the critical importance of Parascaris spp. in young equids.

zoology

Association between the 4p16 genomic locus and different types of congenital heart disease: results from adult survivors in the UK Biobank

Congenital heart disease is the most common birth defect in newborns and the leading cause of death in infancy, affecting nearly 1% of live births. A locus in chromosome 4p16, adjacent to MSX1 and STX18, has been associated with atrial septal defects (ASD) in multiple European and Chinese cohorts. Here, genotyping data from the UK Biobank was used to test for associations between this locus and congenital heart disease in adult survivors of left ventricular outflow tract obstruction (n=164) and ASD (n=223), with a control sample of 332,788 individuals, and a meta-analysis of the new and existing ASD data was performed.\n\nThe results show an association between the previously reported markers at 4p16 and risk for either ASD or left ventricular outflow tract obstruction, with effect sizes similar to the published data (OR between 1.27-1.45; all p<0.05). Differences in allele frequencies remained constant through the studied age range (40-70 years), indicating that the variants themselves do not drive lethal genetic defects. Meta-analysis shows an OR of 1.35 (95% CI: 1.25-1.46; p<10-4) for the association with ASD.\n\nThe findings show that the genetic associations with ASD can be generalized to adult survivors of both ASD and left ventricular lesions. Although the 4p16 associations are statistically compelling, the mentioned alleles confer only a small risk for disease and their frequencies in this adult sample are the same as in children, likely limiting their clinical significance. Further epidemiological and functional studies may elicit factors triggering disease in interaction with the risk alleles.

genetics

JAM-A functions as a female microglial tumor suppressor in glioblastoma

Glioblastoma (GBM) remains refractory to treatment. In addition to its cellular and molecular heterogeneity, epidemiological studies indicate the presence of additional complexity associated with biological sex. GBM is more prevalent and aggressive in male compared to female patients, suggesting the existence of sex-specific growth, invasion, and therapeutic resistance mechanisms. While sex-specific molecular mechanisms have been reported at a tumor cell-intrinsic level, sex-specific differences in the tumor microenvironment have not been investigated. Using transgenic mouse models, we demonstrate that deficiency of junctional adhesion molecule-A (JAM-A) in female mice enhances microglia activation, GBM cell proliferation, and tumor growth. Mechanistically, JAM-A suppresses anti-inflammatory/pro-tumorigenic gene activation via interferon-activated gene 202b (Ifi202b) and found in inflammatory zone (Fizz1) in female microglia. Our findings suggest that cell adhesion mechanisms function to suppress pathogenic microglial activation in the female tumor microenvironment, which highlights an emerging role for sex differences in the GBM microenvironment and suggests that sex differences extend beyond previously reported tumor cell intrinsic differences.\n\nSummaryTuraga et al. demonstrate that female microglia drive a more aggressive glioblastoma phenotype in the context of JAM-A deficiency. These findings highlight a sex-specific role for JAM-A and represent the first evidence of sexual dimorphism in the glioblastoma microenvironment.

cancer biology

Physical activity and risks of breast and colorectal cancer: A Mendelian randomization analysis

Physical activity has been associated with lower risks of breast and colorectal cancer in epidemiological studies; however, it is unknown if these associations are causal or confounded. In two-sample Mendelian randomization analyses, using summary genetic data from the UK Biobank and GWA consortia, we found that a one standard deviation increment in average acceleration was associated with lower risks of breast cancer (odds ratio [OR]: 0.59, 95% confidence interval [CI]: 0.42 to 0.84, P-value=0.003) and colorectal cancer (OR: 0.66, 95% CI: 0.53 to 0.82, P-value=2*E-4). We found similar magnitude inverse associations by breast cancer subtype and by colorectal cancer anatomical site. Our results support a potentially causal relationship between higher physical activity levels and lower risks of breast cancer and colorectal cancer. Based on these data, the promotion of physical activity is probably an effective strategy in the primary prevention of these commonly diagnosed cancers.\n\nDisclaimerWhere authors are identified as personnel of the International Agency for Research on Cancer / World Health Organization, the authors alone are responsible for the views expressed in this article and they do not necessarily represent the decisions, policy or views of the International Agency for Research on Cancer / World Health Organization.

genetics

A new high-throughput tool to screen mosquito-borne viruses in Zika virus endemic/epidemic areas

Mosquitoes are vectors of arboviruses affecting animal and human health. Arboviruses circulate primarily within an enzootic cycle and recurrent spillovers contribute to the emergence of human-adapted viruses able to initiate an urban cycle involving anthropophilic mosquitoes. The increasing volume of travel and trade offers multiple opportunities for arbovirus introduction in new regions. This scenario has been exemplified recently with the Zika pandemic. To incriminate a mosquito as vector of a pathogen, several criteria are required such as the detection of natural infections in mosquitoes. In this study, we used a high-throughput chip based on the BioMarkTM Dynamic arrays system capable of detecting 64 arboviruses in a single experiment. A total of 17,958 mosquitoes collected in Zika-endemic/epidemic countries (Brazil, French Guiana, Guadeloupe, Suriname, Senegal, and Cambodia) were analyzed. Here we show that this new tool can detect endemic and epidemic viruses in different mosquito species in an epidemic context. Thus, this fast and low-cost method can be suggested as a novel epidemiological surveillance tool to identify circulating arboviruses.

molecular biology

A robust estimation of mosaic loss of chromosome Y from genotype-array-intensity data to improve disease risk associations and transcriptional effects

Accurate protocols and methods to robustly detect the mosaic loss of chromosome Y (mLOY) are needed given its reported role in cancer, several age-related disorders and overall male mortality. Intensity SNP-array data have been used to infer mLOY status and to determine its prominent role in male disease. However, discrepancies of reported findings can be due to the uncertainty and variability of the methods used for mLOY detection and to the differences in the tissue-matrix used. We proposed MADloy, the first publicly available software tool that incorporates previous methods and includes a new robust approach, allowing efficient calling in large studies and comparisons between methods. The new method implemented in MADloy optimizes mLOY calling by correctly modeling the underlying reference population with no-mLOY status and incorporating B-deviation information. We observed improvements in the calling accuracy with respect to previous methods, using experimentally validated samples, and an increment in the statistical power to detect associations with disease and mortality, using simulation studies and real dataset analyses. We applied MADloy to detect the increment of mLOY cellularity in blood on 18 individuals after 3 years, and to confirm that its detection in saliva was sub-optimal (41%). We illustrate the use of MADloy to detect the down-regulation genes in the chromosome Y in kidney and bladder tumors with mLOY, and to perform pathway analyses for the detection of mLOY in blood. MADloy is a new software tool implemented in R for easy and robust calling of mLOY status in men aimed to facilitate its study in large epidemiological studies.

genomics

Bayesian hierarchical models for disease mapping applied to contagious pathologies

Disease mapping aims to determine the underlying disease risk from scattered epidemiological data and to represent it on a smoothed colored map. This methodology is based on Bayesian inference and is classically dedicated to non-infectious diseases whose incidence is low and whose cases distribution is spatially (and eventually temporally) structured. Over the last decades, disease mapping has received many major improvements to extend its scope of application: integrating the temporal dimension, dealing with missing data, taking into account various a prioris (environmental and population covariates, assumptions concerning the repartition and the evolution of the risk), dealing with overdispersion, etc. We aim to adapt this approach to rare infectious diseases. In the context of a contagious disease, the outcome of a primary case can in addition generate secondary occurrences of the pathology in a close spatial and temporal neighborhood; this can result in local overdispersion and in higher spatial and temporal dependencies due to direct and/or indirect transmission. We have proposed and tested 60 Bayesian hierarchical models on 400 simulated datasets and bovine tuberculosis real data. This analysis shows the relevance of the CAR (Conditional AutoRegressive) processes to deal with the structure of the risk. We can also conclude that the negative binomial models outperform the Poisson models with a Gaussian noise to handle overdispersion. In addition our study provided relevant maps which are congruent with the real risk (simulated data) and with the knowledge concerning bovine tuberculosis (real data).\n\nAuthor summaryDisease mapping is dedicated to non-infectious diseases whose incidence is low and whose distribution is spatially (and eventually temporally) structured. In this paper, we aim to adapt this approach to rare infectious pathologies. In the context of a contagious disease, the outcome of a primary case can in addition generate secondary occurrences of the pathology in a close spatial and temporal neighborhood, resulting in local overdispersion and in high spatial and temporal dependencies. We thus explored different adapted spatial, temporal and spatiotemporal links and highlight the most adapted to likely risk structures for infectious diseases. We also conclude that the negative binomial models outperform the Poisson models with a Gaussian noise to handle overdispersion. Our study also provided relevant maps which are congruent with the real risk (in case of simulated data) and with the knowledge concerning bovine tuberculosis (when applying to real data). Thus disease mapping appears as a promising way to investigate rare infectious diseases.

scientific communication and education

Understanding the role of urban design in disease spreading

Cities are complex systems whose characteristics impact the health of people who live in them. Nonetheless, urban determinants of health often vary within spatial scales smaller than the resolution of epidemiological datasets. Thus, as cities expand and their inequalities grow, the development of theoretical frameworks that explain health at the neighborhood level is becoming increasingly critical. To this end, we developed a methodology that uses census data to introduce urban geography as a leading-order predictor in the spread of influenza-like pathogens. Here, we demonstrate our framework using neighborhood-level census data for Guadalajara (GDL, Western Mexico). Our simulations were calibrated using weekly hospitalization data from the 2009 A/H1N1 influenza pandemic and show that daily mobility patterns drive neighborhood-level variations in the basic reproduction number R0, which in turn give rise to robust spatiotemporal patterns in the spread of disease. To generalize our results, we ran simulations in hypothetical cities with the same population, area, schools and businesses as GDL but different land use zoning. Our results demonstrate that the agglomeration of daily activities can largely influence the growth rate, size and timing of urban epidemics. Overall, these findings support the view that cities can be redesigned to limit the geographic scope of influenza-like outbreaks and provide a general mathematical framework to study the mechanisms by which local and remote health consequences result from characteristics of the physical environment. Author summaryEnvironmental, social and economic factors give rise to health inequalities among the inhabitants of a city, prompting researchers to propose smart urban planning as a tool for public health. Here, we present a mathematical framework that relates the spatial distributions of schools and economic activities to the spatiotemporal spread of influenza-like outbreaks. First, we calibrated our model using city-wide data for Guadalajara (GDL, Western Mexico) and found that a persons place of residence can largely influence their role and vulnerability during an epidemic. In particular, the higher contact rates of people living near major activity hubs can give rise to predictable patterns in the spread of disease. To test the universality of our findings, we redesigned GDL by redistributing houses, schools and businesses across the city and ran simulations in the resulting geographies. Our results suggest that, through its impact on the agglomeration of economic activities, urban planning may be optimized to inhibit epidemic growth. By predicting health inequalities at the neighborhood-level, our methodology may help design public health strategies that optimize resources and target those who are most vulnerable. Moreover, it provides a mathematical framework for the design and analysis of experiments in urban health research.

systems biology

When does a minor outbreak become a major epidemic? Linking the risk from invading pathogens to practical definitions of a major epidemic

Forecasting whether or not initial reports of disease will be followed by a severe epidemic is an important component of disease management. Standard epidemic risk estimates involve assuming that infections occur according to a branching process and correspond to the probability that the outbreak persists beyond the initial stochastic phase. However, an alternative assessment is to predict whether or not initial cases will lead to a severe epidemic in which available control resources are exceeded. We show how this risk can be estimated by considering three practically relevant potential definitions of a severe epidemic; namely, an outbreak in which: i) a large number of hosts are infected simultaneously; ii) a large total number of infections occur; and iii) the pathogen remains in the population for a long period. We show that the probability of a severe epidemic under these definitions often coincides with the standard branching process estimate for the major epidemic probability. However, these practically relevant risk assessments can also be different from the major epidemic probability, as well as from each other. This holds in different epidemiological systems, highlighting that careful consideration of what constitutes a severe epidemic in an ongoing outbreak is vital for accurate risk quantification.

ecology

Weak-Instrument Robust Tests in Two-Sample Summary-Data Mendelian Randomization

Mendelian randomization (MR) is a popular method in genetic epidemiology to estimate the effect of an exposure on an outcome using genetic variants as instrumental variables (IV), with two-sample summary-data MR being the most popular due to privacy. Unfortunately, many MR methods for two-sample summary data are not robust to weak instruments, a common phenomena with genetic instruments; many of these methods are biased and no existing MR method has Type I error control under weak instruments. In this work, we propose test statistics that are robust to weak instruments by extending Anderson-Rubin, Kleibergen, and conditional likelihood ratio tests in econometrics to the two-sample summary data setting. We conclude with a simulation and an empirical study and show that the proposed tests control size and have better power than current methods.

genetics

Future Preventive Gene Therapy of Polygenic Diseases from a Population Genetics Perspective

With the accumulation of scientific knowledge of the genetic causes of common diseases and continuous advancement of gene-editing technologies, gene therapies to prevent polygenic diseases may soon become possible. This study endeavored to assess population genetics consequences of such therapies. Computer simulations were used to evaluate the heterogeneity in causal alleles for polygenic diseases that could exist among geographically distinct populations. The results show that although heterogeneity would not be easily detectable by epidemiological studies following population admixture, even significant heterogeneity would not impede the outcomes of preventive gene therapies. Preventive gene therapies designed to correct causal alleles to a naturally-occurring neutral state of nucleotides would lower the prevalence of polygenic early- to middle-age-onset diseases in proportion to the decreased population relative risk attributable to the edited alleles. The outcome would manifest differently for late-onset diseases, for which the therapies would result in a delayed disease onset and decreased lifetime risk, however the lifetime risk would increase again with prolonging population life expectancy, which is a likely consequence of such therapies. If gene therapies that prevent heritable diseases were to be applied on a large scale, the decreasing frequency of risk alleles in populations would reduce the disease risk or delay the age of onset, even with a fraction of the population receiving such therapies. With ongoing population admixture, all groups would benefit over generations.

genetics

%svy_freqs: A generic SAS macro for cross-tabulation between a factor and a by-group variable given a third variable and creating publication-quality tables using data from complex surveys

IntroductionIn epidemiological studies, cross-tabulations are a simple but important tool for understanding the distribution of socio-demographic characteristics among study participants. They become more useful when comparisons are presented using a by-group variable such as key demographic characteristic or an outcome status; for instance, sex or the presence or absence of a disease status. Most available statistical analysis software can easily perform cross-tabulations, however, output from these must be processed further to make it readily available for review and use in a publication. In addition, performing three-way cross-tabulations of complex survey data such as those required to show the distribution of disease prevalence across multiple factors and a by-group variable is not easily implemented directly using available standard procedures of commonly used statistical software.\n\nMethodsWe developed a generic SAS macro, %svy_freqs, to create quality publication-ready tables from cross-tabulations between a factor and a by-group variable given a third variable using survey or non-survey data. The SAS macro also performs classical two-way cross-tabulations and refines output into publication-quality tables. It provides extra features not available in existing procedures such as ability to incorporate parameters for survey design and replication-based variance estimation methods, performing validation checks for input parameters, transparently formatting character variable values into numeric ones and allowing for generalizability.\n\nResultsWe demonstrate the application of the SAS macro in the analysis of data from the 2013-2014 National Health and Nutrition Examination Survey (NHANES), a complex survey designed to assess the health and nutritional status of adults and children in the United States (U.S.).\n\nConclusionThe SAS code use to develop the macro is simple yet comprehensive, easy to follow, straightforward for the end user and simple for a SAS programmer to extend. The SAS macro has shown to shorten turn-around time for statistical analysis, eliminate errors when preparing output, and support reproducible research.

bioinformatics