bioRxiv ScienceSearch

bioRxiv · 10.1101/367037

Using machine learning to guide targeted and locally-tailored empiric antibiotic prescribing in a children’s hospital in Cambodia

Abstract

BackgroundEarly and appropriate empiric antibiotic treatment of patients suspected of having sepsis is associated with reduced mortality. The increasing prevalence of antimicrobial resistance risks eroding the benefits of such empiric therapy. This problem is particularly severe for children in developing country settings. We hypothesized that by applying machine learning approaches to readily collected patient data, it would be possible to obtain actionable and patient-specific predictions for antibiotic-susceptibility. If sufficient discriminatory power can be achieved, such predictions could lead to substantial improvements in the chances of choosing an appropriate antibiotic for empiric therapy, while minimizing the risk of increased selection for resistance due to use of antibiotics usually held in reserve.\n\nMethods and FindingsWe analyzed blood culture data collected from a 100-bed childrens hospital in North-West Cambodia between February 2013 and January 2016. Clinical, demographic and living condition information for each child was captured with 35 independent variables. Using these variables, we used a suite of machine learning algorithms to predict Gram stains and whether bacterial pathogens could be treated with standard empiric antibiotic therapies: i) ampicillin and gentamicin; ii) ceftriaxone; iii) at least one of the above.\n\n243 cases of bloodstream infection were available for analysis. We used 195 (80%) to train the algorithms, and 48 (20%) for evaluation. We found that the random forest method had the best predictive performance overall as assessed by the area under the receiver operating characteristic curve (AUC), though support vector machine with radial kernel had similar performance for predicting Gram stain and ceftriaxone susceptibility. Predictive performance of logistic regression, simple and boosted decision trees and k-nearest neighbors were poor in comparison. The random forest method gave an AUC of 0.91 (95%CI 0.81-1.00) for predicting susceptibility to ceftriaxone, 0.75 (0.60-0.90) for susceptibility to ampicillin and gentamicin, 0.76 (0.59-0.93) for susceptibility to neither, and 0.69 (0.53-0.85) for Gram stain result. The most important variables for predicting susceptibility were time from admission to blood culture, patient age, hospital versus community-acquired infection, and age-adjusted weight score.\n\nConclusionsApplying machine learning algorithms to patient data that are readily available even in resource-limited hospital settings can provide highly informative predictions on susceptibilities of pathogens to guide appropriate empiric antibiotic therapy. Used as a decision support tool, such approaches have the potential to lead to better targeting of empiric therapy, improve patient outcomes and reduce the burden of antimicrobial resistance.\n\nAuthor summaryO_LSTWhy was this study done?C_LSTO_LIEarly and appropriate antibiotic treatment of patients with life-threatening bacterial infections is thought to reduce the risk of mortality.\nC_LIO_LIIn hospitals that have a microbiology laboratory, it takes 3-4 days to get results which indicate which antibiotics are likely to be effective; before this information is available antibiotics have to be prescribed empirically i.e. without knowledge of the causative organism.\nC_LIO_LIIncreasing resistance to antibiotics amongst bacteria makes finding an appropriate antibiotic to use empirically difficult; this problem is particularly severe for children in developing country settings.\nC_LIO_LIIf we could predict which antibiotics were likely to be effective at the time of starting antibiotic therapy, we might be able to improve patient outcomes and reduce resistance.\nC_LI\n\nO_LSTWhat Did the Researchers Do and Find?C_LSTO_LIWe evaluated the ability of a number of different algorithms (i.e. sets of step-by-step instructions) to predict susceptibility to commonly-used antibiotics using routinely available patient data from a childrens hospital in Cambodia.\nC_LIO_LIWe found that an algorithm called random forests enabled surprisingly accurate predictions, particularly for predicting whether the infection was likely to be treatable with ceftriaxone, the most commonly used empiric antibiotic at the study hospital.\nC_LIO_LIUsing this approach it would be possible to correctly predict when a different antibiotic would be needed for empiric treatment over 80% of the time, while recommending a different antibiotic when ceftriaxone would suffice less than 20% of the time.\nC_LI\n\nO_LSTWhat Do These Findings Mean?C_LSTO_LIUsing readily available patient information, sophisticated algorithms can enable good predictions of whether antibiotics are likely to be effective several days before laboratory tests are available.\nC_LIO_LIAlgorithms would need to be trained with local hospital data, but our study shows that even with relatively limited data from a small hospital, good predictions can be obtained.\nC_LIO_LIUsed as part of a decision support system such algorithms could help choose appropriate antibiotics for empiric therapy; this would be expected to translate into better patient outcomes and may help to reduce resistance.\nC_LIO_LISuch as a decision support system would have very low costs and be easy to implement in low- and middle-income countries.\nC_LI

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Oonsivilai, M., Yin, M., Luangasanatip, N., Lubell, Y., Miliya, T., Tan, P., Loeuk, L., Turner, P., Cooper, B. S.. 2018-07-13. Using machine learning to guide targeted and locally-tailored empiric antibiotic prescribing in a children’s hospital in Cambodia. https://doi.org/10.1101/367037

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Translating surveillance data into incidence estimates

Monitoring a population for a disease requires the hosts to be sampled and tested for the pathogen. This results in sampling series from which to estimate the disease incidence, i.e. the proportion of hosts infected. Existing estimation methods assume that disease incidence is not changing between monitoring rounds, resulting in underestimation of the disease incidence. In this paper we develop an incidence estimation model accounting for epidemic growth with monitoring rounds sampling varying incidence. We also show how to accommodate the asymptomatic period characteristic to most diseases. For practical use, we produce an approximation of the model, which is subsequently shown accurate for relevant epidemic and sampling parameters. Both the approximation and the full model are applied to stochastic spatial simulations of epidemics. The results prove their consistency for a very wide range of situations.

epidemiology

The Swiss Primary Ciliary Dyskinesia registry: objectives, methods and first results

Primary Ciliary Dyskinesia (PCD) is a rare hereditary, multi-organ disease caused by defects in ciliary structure and function. It results in a wide range of clinical manifestations, most commonly in the upper and lower airways. Central data collection in national and international registries is essential to studying the epidemiology of rare diseases and filling in gaps in knowledge of diseases such as PCD. For this reason, the Swiss Primary Ciliary Dyskinesia Registry (CH-PCD) was founded in 2013 as a collaborative project between epidemiologists and adult and paediatric pulmonologists.\n\nThe registry records patients of any age, suffering from PCD, who are treated and resident in Switzerland. It collects information from patients identified through physicians, diagnostic facilities, and patient organisations. The registry dataset contains data on diagnostic evaluations, lung function, microbiology and imaging, symptoms, treatments, and hospitalizations.\n\nBy May 2018, CH-PCD has contacted 566 physicians of different specialties and identified 134 patients with PCD. At present this number represents an overall 1 in 63,000 prevalence of people diagnosed with PCD in Switzerland. Prevalence differs by age and region; it is highest in children and adults younger than 30 years, and in Espace Mittelland. The median age of patients in the registry is 25 years (range 5-73), and 49 patients have a definite PCD diagnosis based on recent international guidelines. Data from CH-PCD are contributed to international collaborative studies and the registry facilitates patient identification for nested studies.\n\nCH-PCD has proven to be a valuable research tool that already has highlighted weaknesses in PCD clinical practice in Switzerland. Development of centralised diagnostic and management centres and adherence to international guidelines are needed to improve diagnosis and management--particularly for adult PCD patients.

epidemiology

Perfect Counterfactuals for Epidemic Simulations

Simulation studies are often used to predict the expected impact of control measures in infectious disease outbreaks. Typically, two independent sets of simulations are conducted, one with the intervetnion, and one without, and epidemic sizes (or some related metric) are compared to estimate the effect of the intervention. Since it is possible that controlled epidemics are larger than uncontrolled ones if there is substantial stochastic variation between epidemics, uncertainty intervals from this approach can include a negative effect even for an effective intervention. To more precisely estimate the number of cases an intervention will prevent within a single epidemic, here we develop a single world approach to matching simulations of controlled epidemics to their exact uncontrolled counterfac-tual. Our method borrows concepts from percolation approaches prune out possible epidemic histories and create potential epidemic graph that can be realized to create perfectly matched controlled and uncontrolled epidemics. We present an implementation of this method for a common class of compartmental models, and its application in a simple SIR model. Results illustrate how, at the cost of some computation time, this method substantially narrows confidence intervals and avoids non-sensical inferences.

epidemiology