bioRxiv ScienceSearch

SEARCH · bioRxiv Science

Results for “Epidemiology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,243 records · Page 69Linked to original sources

Improved Automatic Pharmacovigilance: An Enhancement to the MedWatcher Social System for Monitoring Adverse Events

Traditional pharmacovigilance systems rely on adverse event reports received by regulatory authorities such as the United States Food and Drug Administration (FDA). These traditional systems suffer from underreporting and are not timely due to their reliance on third-party sentinels. To address these issues, the MedWatcher Social system for monitoring adverse events through automated processing of digital social media data and crowdsourcing was launched in 2012 by Boston Childrens Hospital and the FDA. The system is rooted in the well-established FDA MedWatch system.\n\nMedWatcher Social uses an indicator score approach to identify adverse events. This study evaluates the MedWatcher Social adverse event classifiers performance on Twitter data and proposes an enhancement to the indicator score method that results in improved adverse event identification.\n\nOur research suggests that automatic pharmacovigilance systems using the original indicator score approach should be updated. Careful consideration of modeling assumptions is critical when designing algorithms for computational epidemiology, and algorithms should be regularly reevaluated to identify enhancements and to remedy concept drift.

bioinformatics

Causal relationship of cerebrospinal fluid biomarkers with the risk of Alzheimer’s disease: A two-sample Mendelian randomization study

Whether the epidemiological association of amyloid beta (A{beta}) and tau pathology with Alzheimers disease (AD) is causal remains unclear. The recent failures to demonstrate the efficacy of several amyloid beta-modifying drugs may indicate the possibility that the observed association is not causal. These failures also led to efforts to develop tau-directed treatments whose efficacy is still tentative. Herein, we conducted a two-sample Mendelian randomization analysis to determine whether the relationship between the cerebrospinal fluid (CSF) biomarkers for amyloid and tau pathology and the risk of AD is causal. We used the summary statistics of a genome-wide association study (GWAS) for CSF biomarkers (A{beta}1-42, phosphorylated tau 181 [p-tau], and total tau [t-tau]) in 3,146 individuals and for late-onset AD (LOAD) in 21,982 LOAD cases and 41,944 cognitively normal controls. We tested the association between the change in the genetically predicted CSF biomarkers and LOAD risk. We found a modest decrease in the LOAD risk per one standard deviation (SD) increase in the genetically predicted CSF A{beta} (odds ratio [OR], 0.63 for AD; 95% confidence interval [CI], 0.38-0.87; P = 0.02). In contrast, we observed a significant increase in the LOAD risk per one SD increase in the genetically predicted CSF p-tau (OR, 2.37; 95% CI, 1.46-3.28; P = 1.09x10-5). However, no causal association was observed of the CSF t-tau with the LOAD risk (OR, 1.15; 95% CI, 0.85-1.45; P = 0.29). Our findings need to be validated in future studies with more genetic variants identified in larger GWASs for CSF biomarkers.

genomics

Topological data analysis to uncover the shape of immune responses during co-infection

Co-infections by multiple pathogens have important implications in many aspects of health, epidemiology and evolution. However, how to disentangle the contributing factors of the immune response when two infections take place at the same time is largely unexplored. Using data sets of the immune response during influenza-pneumococcal co-infection in mice, we employ here topological data analysis to simplify and visualise high dimensional data sets.\n\nWe identified persistent shapes of the simplicial complexes of the data in the three infection scenarios: single viral infection, single bacterial infection, and co-infection. The immune response was found to be distinct for each of the infection scenarios and we uncovered that the immune response during the co-infection has three phases and two transition points. During the first phase, its dynamics is inherited from its response to the primary (viral) infection. The immune response has an early (few hours post co-infection) and then modulates its response to finally react against the secondary (bacterial) infection. Between 18 to 26 hours post co-infection the nature of the immune response changes again and does no longer resembles either of the single infection scenarios.\n\nAuthor summaryThe mapper algorithm is a topological data analysis technique used for the qualitative analysis, simplification and visualisation of high dimensional data sets. It generates a low-dimensional image that captures topological and geometric information of the data set in high dimensional space, which can highlight groups of data points of interest and can guide further analysis and quantification.\n\nTo understand how the immune system evolves during the co-infection between viruses and bacteria, and the role of specific cytokines as contributing factors for these severe infections, we use Topological Data Analysis (TDA) along with an extensive semi-unsupervised parameter value grid search, and k-nearest neighbour analysis.\n\nWe find persistent shapes of the data in the three infection scenarios, single viral and bacterial infections and co-infection. The immune response is shown to be distinct for each of the infections scenarios and we uncover that the immune response during the co-infection has three phases and two transition points, a previously unknown property regarding the dynamics of the immune response during co-infection.

immunology

SNAPPy: a snakemake pipeline for scalable HIV-1 subtyping by phylogenetic pairing

Human immunodeficiency virus 1 (HIV-1) genome sequencing is routinely done for drug resistance monitoring in hospitals worldwide. Subtyping these extensive datasets of HIV-1 sequences is a critical first step in molecular epidemiology and surveillance studies. The clinical relevance of HIV-1 subtypes is increasingly recognized. Several studies suggest subtype-related differences in disease progression, transmission route efficiency, immune evasion, and even therapeutic outcomes. HIV-1 subtyping is mainly done using web servers. These tools have limitations in scalability and potential noncompliance with data protection legislation. Thus, the aim of this work was to develop an efficient method for local and high-throughput HIV-1 subtyping. We designed SNAPPy: a snakemake pipeline for scalable HIV-1 subtyping by phylogenetic pairing. It contains several tasks of phylogenetic inference and BLAST queries, which can be executed sequentially or in parallel, taking advantage of multiple-core processing units. Although it was built for subtyping, SNAPPy is also useful to perform extensive HIV-1 alignments. This tool facilitates large-scale sequence-based HIV-1 research by providing a local, resource efficient and scalable alternative for HIV-1 subtyping. It is capable of analysing full-length genomes or partial HIV-1 genomic regions (GAG, POL, ENV) and recognizes more than 90 circulating recombinant forms. SNAPPy is freely available at: https://github.com/PMMAraujo/snappy.

bioinformatics

Comparative analysis of MOBQ4 plasmids demonstrates that MOBQ is a cis-acting-enriched relaxase protein family

A group of small mobilizable plasmids is increasingly being reported in epidemiology surveys of enterobacteria. Some of them encode colicins, while others are cryptic. All of them encode a relaxase belonging to a previously non-described group of the MOBQ class, MOBQ4. While highly similar in their mobilization module, two families with unrelated replicons can be distinguished, MOBQ41 and MOBQ42. Members of both groups were compatible between them and stably maintained in E. coli. MOBQ4 plasmids were mobilized by conjugation. They contain two transfer genes, mobA coding for the MOBQ4 relaxase and mobC, which was non-essential but enhanced the plasmid mobilization frequency. The origin of transfer was located between these two divergently transcribed mob genes. MPFI conjugative plasmids were the most efficient helpers for MOBQ4 conjugative transmission. No interference in mobilization was observed when both MOBQ41 and MOBQ42 were present in the same donor cell. Remarkably, MOBQ4 relaxases exhibited a cis-acting preference for their oriTs, a feature already observed in other MOBQ plasmids. These findings indicate that MOBQ4 plasmids can efficiently spread among enterobacteria aided by coresident IncI1, IncK and IncL/M plasmids, while ensuring their self-dissemination over highly-related elements.\n\nIMPORTANCEPlasmids are key vehicles of horizontal gene transfer and contribute greatly to bacterial genome plasticity. A group of plasmids, called mobilizable, is able to disseminate aided by helper conjugative plasmids. Here, we studied a group of phylogenetically-related mobilizable plasmids, MOBQ4, commonly found in clinically-relevant enterobacteria, uncovering the helper plasmids responsible for their dissemination. We found that the two plasmid species encompassed in the MOBQ4 group can coexist and transfer orthogonally, despite origin-of-transfer cross-recognition by their relaxases. Specific discrimination among their highly similar oriT sequences is guaranteed by the preferential cis activity of the MOBQ4 relaxases. Such strategy would be biologically relevant in a scenario of co-residence of non-divergent elements to favor self-dissemination.

microbiology

Dynamics of Livestock-Associated Methicillin Resistant Staphylococcus aureus in pig farms networks: insight from mathematical modeling and French data

Livestock-associated methicillin resistant Staphylococcus aureus (LA-MRSA) colonizes livestock animals worldwide, especially pigs and calves. Although frequently carried asymptomatically, LA-MRSA can cause severe infections in humans. It is therefore important to better understand LA-MRSA spreading dynamics within pig farms and over pig farms networks, and to compare different strategies of control and surveillance. For this purpose, we propose a stochastic meta-population model of LA-MRSA spread along the French pig-farm network (n=10,542 farms), combining within- and between-farms dynamics, based on detailed data on breeding practices and pig exchanges between holdings. We calibrate the model using French epidemiological data. We then identify farm-level factors associated with the spreading potential of LA-MRSA in the network. We also show that, assuming control measures applied in a limited (n=100) number of farms, targeting farms depending on their centrality in the network is the only way to significantly reduce LA-MRSA global prevalence. Finally, we investigate the scenario of emergence of a new LA-MRSA strain, and find that the farms with the highest indegree would be the best sentinels for a targeted surveillance of such a strains introduction.

microbiology

Global genomic population structure of Clostridioides difficile

Clostridioides difficile is the primary infectious cause of antibiotic-associated diarrhea. Local transmissions and international outbreaks of this pathogen have been previously elucidated by bacterial whole-genome sequencing, but comparative genomic analyses at the global scale were hampered by the lack of specific bioinformatic tools. Here we introduce EnteroBase, a publicly accessible database (http://enterobase.warwick.ac.uk) that automatically retrieves and assembles C. difficile short-reads from the public domain, and calls alleles for core-genome multilocus sequence typing (cgMLST). We demonstrate that the identification of highly related genomes is 89% consistent between cgMLST and single-nucleotide polymorphisms. EnteroBase currently contains 13,515 quality-controlled genomes which have been assigned to hierarchical sets of single-linkage clusters by cgMLST distances. Hierarchical clustering can be used to identify populations of C. difficile at all epidemiological levels, from recent transmission chains through to pandemic and endemic strains, and is largely compatible with prior ribotyping. Hierarchical clustering thus enables comparisons to earlier surveillance data and will facilitate communication among researchers, clinicians and public-health officials who are combatting disease caused by C. difficile.

microbiology

Cryptococcus neoformans recovered from olive trees (Olea europaea) in Turkey reveal allopatry with African and South American lineages

Cryptococcus species are life-threatening human fungal pathogens that cause cryptococcal meningoencephalitis in both immunocompromised and healthy hosts. The natural environmental niches of Cryptococcus include pigeon (Columba livia) guano, soil, and a variety of tree species such as Eucalyptus camaldulensis, Ceratonia siliqua, Platanus orientalis, and Pinus spp. Genetic and genomic studies of extensive sample collections have provided insights into the population distribution and composition of different Cryptococcus species in geographic regions around the world. However, few such studies examined Cryptococcus in Turkey. We sampled 388 Olea europaea (olive) and 132 E. camaldulensis trees from 7 locations in coastal and inland areas of the Aegean region of Anatolian Turkey in September 2016 to investigate the distribution and genetic diversity present in the natural Cryptococcus population. We isolated 84 Cryptococcus neoformans strains (83 MAT and 1 MATa) and 3 Cryptococcus deneoformans strains (all MATa) from 87 (22.4% of surveyed) O. europaea trees; a total of 32 C. neoformans strains were isolated from 32 (24.2%) of the E. camaldulensis trees, all of which were MAT. A statistically significant difference was observed in the frequency of C. neoformans isolation between coastal and inland areas (P < 0.05). Thus, O. europaea trees could represent a novel niche for C. neoformans. Interestingly, the MATa C. neoformans isolate was fertile in laboratory crosses with VNI and VNB MAT tester strains and produced robust hyphae, basidia, and basidiospores, thus suggesting potential sexual reproduction in the natural population. Sequencing analyses of the URA5 gene identified at least 5 different genotypes among the isolates. Population genetics and genomic analyses revealed that most of the isolates in Turkey belong to the VNBII lineage of C. neoformans, which is predominantly found in southern Africa; these isolates are part of a distinct minor clade within VNBII that includes several isolates from Zambia and Brazil. Our study provides insights into the geographic distribution of different C. neoformans lineages in the Mediterranean region and highlights the need for wider geographic sampling to gain a better understanding of the natural habitats, migration, epidemiology, and evolution of this important human fungal pathogen.

genetics

Genomic variant identification methods alter Mycobacterium tuberculosis transmission inference

Pathogen genomic data are increasingly used to characterize global and local transmission patterns of important human pathogens and to inform public health interventions. Yet there is no current consensus on how to measure genomic variation. We investigated the effects of variant identification approaches on transmission inferences for M. tuberculosis by comparing variants identified by five different groups in the same sequence data from a clonal outbreak. We then measured the performance of commonly used variant calling approaches in recovering variation in a simulated tuberculosis outbreak and tested the effect of applying increasingly stringent filters on transmission inferences and phylogenies. We found that variant calling approaches used by different groups do not recover consistent sets of variants, often leading to conflicting transmission inferences. Further, performance in recovering true outbreak variation varied widely across approaches. Finally, stringent filters rapidly eroded the accuracy of transmission inferences and quality of phylogenies reconstructed from outbreak variation. We conclude that measurements of genetic distance and phylogenetic structure are dependent on variant calling approach. Variant calling algorithms trained upon true sequence data outperform other approaches and enable inclusion of repetitive regions typically excluded from genomic epidemiology studies, maximizing the information gleaned from outbreak genomes.

genomics

Variant antigen diversity in Trypanosoma vivax is not driven by recombination

African trypanosomes are vector-borne haemoparasites that cause African trypanosomiasis in humans and animals. Parasite survival in the bloodstream depends on immune evasion, achieved by antigenic variation of the Variant Surface Glycoprotein (VSG) coating the trypanosome cell surface. Recombination, or rather directed gene conversion, is fundamental in Trypanosoma brucei, as both a mechanism of VSG gene switching and of generating antigenic diversity during infections. Trypanosoma vivax is a related, livestock pathogen also displaying antigenic variation, but whose VSG lack key structures necessary for gene conversion in T. brucei. Thus, this study tests a long-standing prediction that T. vivax has a more restricted antigenic repertoire. Here we show that global VSG repertoire is broadly conserved across diverse T. vivax clinical strains. We use sequence mapping, coalescent approaches and experimental infections to show that recombination plays little, if any, role in diversifying T. vivax VSG sequences. These results explain interspecific differences in disease, such as propensity for self-cure, and indicate that either T. vivax has an alternate mechanism for immune evasion or else a distinct transmission strategy that reduces its reliance on long-term persistence. The lack of recombination driving antigenic diversity in T. vivax has immediate consequences for both the current mechanistic model of antigenic variation in African trypanosomes and species differences in virulence and transmission strategy, requiring us to reconsider the wider epidemiology of animal African trypanosomiasis.

microbiology

Population-Level Disease Dynamics Reflect Individual Heterogeneities in Transmission

Host heterogeneity in disease transmission is widespread and presents a major hurdle to predicting and minimizing pathogen spread. Using the Drosophila melanogaster model system infected with Drosophila C virus, we integrate experimental measurements of individual host heterogeneity in social aggregation, virus shedding, and disease-induced mortality into an epidemiological framework that simulates outbreaks of infectious disease. We use these simulations to calculate individual variation in disease transmission and apportion this variation to specific components of transmission: social network degree distribution, infectiousness, and infection duration. The experimentally-observed variation produces substantial differences in individual transmission potential, providing evidence for genetic and sex-specific effects on disease dynamics at a population level. Manipulating variation in social network connectivity, infectiousness, and infection duration in simulated populations reveals that these components affect disease transmission in clear and distinct ways. We consider the implications of this genetic and sex-specific variation in disease transmission and discuss implications for appropriate control methods given the relative contributions made by social aggregation, virus shedding, and infection duration to transmission in other host-pathogen systems.

ecology

Megacities as drivers of national outbreaks: the role of holiday travel in the spread of infectious diseases

Human mobility connects populations and can lead to large fluctuations in population density, both of which are important drivers of epidemics. Measuring population mobility during infectious disease outbreaks is challenging, but is a particularly important goal in the context of rapidly growing and highly connected urban centers in low and middle income countries, which can act to amplify and spread local epidemics nationally and internationally. Here, we combine estimates of population movement from mobile phone data for over 4 million subscribers in the megacity of Dhaka, Bangladesh, one of the most densely populated cities globally. We combine mobility data with epidemiological data from a household survey, to understand the role of population mobility on the spatial spread of the mosquito-borne virus chikungunya within and outside Dhaka city during a large outbreak in 2017. The peak of the 2017 chikungunya outbreak in Dhaka coincided with the annual Eid holidays, during which large numbers of people traveled from Dhaka to their native region in other parts of the country. We show that regular population fluxes around Dhaka city played a significant role in determining disease risk, and that travel during Eid was crucial to the spread of the infection to the rest of the country. Our results highlight the impact of large-scale population movements, for example during holidays, on the spread of infectious diseases. These dynamics are difficult to capture using traditional approaches, and we compare our results to a standard diffusion model, to highlight the value of real-time data from mobile phones for outbreak analysis, forecasting, and surveillance.

ecology

Determining the serotype composition of mixed samples of pneumococcus using whole genome sequencing

Serotyping of Streptococcus pneumoniae is a critical tool in the surveillance of the pathogen and development and evaluation of vaccines. Whole-genome DNA sequencing and analysis is becoming increasingly common and is an effective method for pneumococcal serotype identification of pure isolates. However, because of the complexities of the pneumococcal capsular loci, current analysis software requires samples to be pure (or nearly pure) and only contain a single pneumococcal serotype. We introduce a new software tool called SeroCall, which can identify and quantitate the serotypes present in samples, even when several serotypes are present. The sample preparation, library preparation and sequencing follow standard laboratory protocols. The software runs as fast or faster than existing identification tools on typical computing servers and is freely available under an open source license at https://github.com/knightjimr/serocall. Using samples with known concentrations of different serotypes as well as blinded samples, we were able to accurately quantify the abundance of different serotypes of pneumococcus in mixed cultures, with 100% accuracy for detecting the major serotype and up to 86% accuracy for detecting minor serotypes. We were also able to track changes in serotype frequency over time in an experimental setting. This approach could be applied in both epidemiologic field studies of pneumococcal colonization as well as in experimental lab studies and could provide a cheaper and more efficient method for serotyping than alternative approaches.

bioinformatics

Essential omega-3 fatty acids tune microglial phagocytosis of synaptic elements in the developing brain

Omega-3 fatty acids (n-3 polyunsaturated fatty acids; n-3 PUFAs) are essential for the functional maturation of the brain. Westernization of dietary habits in both developed and developing countries is accompanied by a progressive reduction in dietary intake of n-3 PUFAs. Low maternal intake of n-3 PUFAs has been linked to neurodevelopmental diseases in epidemiological studies, but the mechanisms by which a n-3 PUFA dietary imbalance affects CNS development are poorly understood. Active microglial engulfment of synaptic elements is an important process for normal brain development and altered synapse refinement is a hallmark of several neurodevelopmental disorders. Here, we identify a molecular mechanism for detrimental effects of low maternal n-3 PUFA intake on hippocampal development. Our results show that maternal dietary n-3 PUFA deficiency increases microglial phagocytosis of synaptic elements in the developing hippocampus, through the activation of 12/15- lipoxygenase (LOX)/12-HETE signaling, which alters neuronal morphology and affects cognition in the postnatal offspring. While women of child bearing age are at higher risk of dietary n-3 PUFA deficiency, these findings provide new insights into the mechanisms linking maternal nutrition to neurodevelopmental disorders.\n\nOne Sentence SummaryLow maternal omega-3 fatty acids intake impairs microglia-mediated synaptic refinement via 12-HETE pathway in the developing brain.

neuroscience

Mitochondrial DNA variation of Leber’s Hereditary Optic Neuropathy (LHON) in Western Siberia

Lebers hereditary optic neuropathy (LHON) is a form of disorder caused by pathogenic mutations in a mitochondrial DNA. LHON is maternally inherited disease, which manifests mainly in young adults, affecting predominantly males. Clinically LHON has a manifestation as painless central vision loss, resulting in early onset of disability. Epidemiology of LHON has not been fully investigated yet. In this study, we report 44 genetically unrelated families with LHON manifestation. We performed whole mtDNA genome sequencing and provided genealogical and molecular genetic data on mutations and haplogroup background of LHON patients in the Western Siberia population. Known \"primary\" pathogenic mtDNA mutations (MITOMAP) were found in 32 families: m.11778G>A represents 53,10% (17/32), m.3460G>A - 21,90% (7/32), m.14484T>C - 18,75% (6/32), and rare m.10663T>C and m.3635G>A represent 6,25% (2/32). We describe potentially pathogenic m.4659G>A in one subject without known pathogenic mutations, and potentially pathogenic m.9444C>T, m.6261G>A, m.9921G>A, m.8551T>C, m.8412T>C, m.15077G>A in families with known pathogenic mutations confirmed. We suppose these mutations could contribute to the pathogenesis of optic neuropathy development. Our results indicate that haplogroup affiliation and mutational spectrum of the Western Siberian LHON cohort substantially deviate from those of European populations.

molecular biology

A comprehensive analysis of racial disparities in chemical biomarker concentrations in United States women, 1999-2014

BackgroundStark racial disparities in disease incidence among American women remains a persistent public health challenge. These disparities likely result from complex interactions between genetic, social, lifestyle, and environmental risk factors. The influence of environmental risk factors, such as chemical exposure, however, may be substantial and is poorly understood.\n\nObjectivesWe quantitatively evaluated chemical-exposure disparities by race/ethnicity and age in United States (US) women by using biomarker data for 143 chemicals from the National Health and Nutrition Examination Survey (NHANES) 1999-2014.\n\nMethodsWe applied a series of survey-weighted, generalized linear models using data from the entire NHANES women population and age-group stratified subpopulations. The outcome was chemical biomarker concentration and the main predictor was race/ethnicity with adjustment for age, socioeconomic status, smoking habits, and NHANES cycle.\n\nResultsThe highest disparities across non-Hispanic Black, Mexican American, Other Hispanic, and other race/multiracial women were observed for pesticides and their metabolites, including 2,5-dichlorophenol, o,p-DDE, beta-hexachlorocyclohexane, and 2,4-dichlorophenol, along with personal care and consumer product compounds. The latter included parabens, monoethyl phthalate, and several metals, such as mercury and arsenic. Moreover, for Mexican American, Other Hispanic, and non-Hispanic black women, there were several exposure disparities that persisted across age groups, such as higher 2,4- and 2,5-dichlorophenol concentrations. Exposure differences for methyl and propyl parabens, however, were the starkest between non-Hispanic black and non-Hispanic white children with average differences exceeding 4 folds.\n\nDiscussionsWe systematically evaluated differences in chemical exposures across women of various race/ethnic groups and across age groups. Our findings could help inform chemical prioritization in designing epidemiological and toxicological studies. In addition, they could help guide public health interventions to reduce environmental and health disparities across populations.

bioinformatics

Efficient whole genome sequencing of influenza A viruses

The constant threat of emergence for novel pathogenic influenza A viruses with pandemic potential, makes full-genome characterization of circulating influenza viral strains a high priority, allowing detection of novel and re-assorting variants. Sequencing the full-length genome of influenza A virus traditionally required multiple amplification rounds, followed by the subsequent sequencing of individual PCR products. The introduction of high-throughput sequencing technologies has made whole genome sequencing easier and faster. We present a simple protocol to obtain whole genome sequences of hypothetically any influenza A virus, even with low quantities of starting genetic material. The complete genomes of influenza A viruses of different subtypes and from distinct sources (clinical samples of pdmH1N1, tissue culture-adapted H3N2 viruses, or avian influenza viruses from cloacal swabs) were amplified with a single multisegment reverse transcription-PCR reaction and sequenced using Illumina sequencing platform. Samples with low quantity of genetic material after initial PCR amplification were re-amplified by an additional PCR using random primers. Whole genome sequencing was successful for 66% of the samples, whilst the most relevant genome segments for epidemiological surveillance (corresponding to the hemagglutinin and neuraminidase) were sequenced with at least 93% coverage (and a minimum 10x) for 98% of the samples. Low coverage for some samples is likely due to an initial low viral RNA concentration in the original sample. The proposed methodology is especially suitable for sequencing a large number of samples, when genetic data is urgently required for strains characterization, and may also be useful for variant analysis.

microbiology

Rapid detection of identity-by-descent tracts for mega-scale datasets

The ability to identify segments of genomes identical-by-descent (IBD) is a part of standard workflows in both statistical and population genetics. However, traditional methods for finding local IBD across all pairs of individuals scale poorly leading to a lack of adoption in very large-scale datasets. Here, we present iLASH, IBD by LocAlity-Sensitive Hashing, an algorithm based on similarity detection techniques that shows equal or improved accuracy in simulations compared to the current leading method and speeds up analysis by several orders of magnitude on genomic datasets, making IBD estimation tractable for hundreds of thousands to millions of individuals. We applied iLASH to the Population Architecture using Genomics and Epidemiology (PAGE) dataset of [~]52,000 multi-ethnic participants, including several founder populations with elevated IBD sharing, which identified IBD segments on a single machine in an hour ([~]3 minutes per chromosome compared to over 6 days per chromosome for a state-of-the-art algorithm). iLASH is able to efficiently estimate IBD tracts in very large-scale datasets, as demonstrated via IBD estimation across the entire UK Biobank ([~]500,000 individuals), detecting nearly 13 billion pairwise IBD tracts shared between [~]11% of participants. In summary, iLASH enables fast and accurate detection of IBD, an upstream step in applications of IBD for population genetics and trait mapping.

genomics