bioRxiv ScienceSearch

Biology subjects

Che, H.

Publications and source records attributed to Che, H..

3 recordsLinked to original sources

De novo assembly of two Swedish genomes reveals missing segments from the human GRCh38 reference and improves variant calling of population-scale sequencing data

We have performed de novo assembly of two Swedish genomes using long-read sequencing and optical mapping, resulting in total assembly sizes of nearly 3 Gb and hybrid scaffold N50 values of over 45 Mb. A further analysis revealed over 10 Mb of sequences absent from the human GRCh38 reference in each individual. Around 6 Mb of these novel sequences (NS) are shared with a Chinese personal genome. The NS are highly repetitive, have elevated GC-content and are primarily located in centromeric or telomeric regions. A BLAST search showed that 31% of the NS are different from any sequences deposited in nucleotide databases. The remaining NS correspond to human (62%) or primate (6%) nucleotide entries, while 1% of hits show the highest similarity to other species, including mouse and a few different classes of parasitic worms. Up to 1 Mb of NS can be assigned to chromosome Y, and large segments are missing from GRCh38 also at chromosomes 14, 17 and 21. Inclusion of these novel sequences into the GRCh38 reference radically improves the alignment and variant calling of whole-genome sequencing data at several genomic loci. Through a re-analysis of 200 samples from a Swedish population-scale sequencing project, we obtained over 75,000 putative novel SNVs per individual when using a custom version of GRCh38 extended with 17.3 Mb of NS. In addition, about 10,000 false positive SNV calls per individual were removed from the GRCh38 autosomes and sex chromosomes in the re-analysis, with some of them located in protein coding regions.

genomics

Genome-wide analysis of ivermectin response by Onchocerca volvulus reveals that genetic drift and soft selective sweeps contribute to loss of drug sensitivity

BackgroundTreatment of onchocerciasis using mass ivermectin administration has reduced morbidity and transmission throughout Africa and Central/South America. Mass drug administration is likely to exert selection pressure on parasites, and phenotypic and genetic changes in several Onchocerca volvulus populations from Cameroon and Ghana - exposed to more than a decade of regular ivermectin treatment - have raised concern that sub-optimal responses to ivermectins anti-fecundity effect are becoming more frequent and may spread.\n\nMethodology/Principal FindingsPooled next generation sequencing (Pool-seq) was used to characterise genetic diversity within and between 108 adult female worms differing in ivermectin treatment history and response. Genome-wide analyses revealed genetic variation that significantly differentiated good responder (GR) and sub-optimal responder (SOR) parasites. These variants were not randomly distributed but clustered in ~31 quantitative trait loci (QTLs), with little overlap in putative QTL position and gene content between countries. Published candidate ivermectin SOR genes were largely absent in these regions; QTLs differentiating GR and SOR worms were enriched for genes in molecular pathways associated with neurotransmission, development, and stress responses. Finally, single worm genotyping demonstrated that geographic isolation and genetic change over time (in the presence of drug exposure) had a significantly greater role in shaping genetic diversity than the evolution of SOR.\n\nConclusions/SignificanceThis study is one of the first genome-wide association analyses in a parasitic nematode, and provides insight into the genomics of ivermectin response and population structure of O. volvulus. We argue that ivermectin response is a polygenically-determined quantitative trait in which identical or related molecular pathways but not necessarily individual genes likely determine the extent of ivermectin response in different parasite populations. Furthermore, we propose that genetic drift rather than genetic selection of SOR is the underlying driver of population differentiation, which has significant implications for the emergence and potential spread of SOR within and between these parasite populations.\n\nAuthor summaryOnchocerciasis is a human parasitic disease endemic across large areas of Sub-Saharan Africa, where more that 99% of the estimated 100 million people globally at-risk live. The microfilarial stage of Onchocerca volvulus causes pathologies ranging from mild itching to visual impairment and ultimately, irreversible blindness. Mass administration of ivermectin kills microfilariae and has an anti-fecundity effect on adult worms by temporarily inhibiting the development in utero and/or release into the skin of new microfilariae, thereby reducing morbidity and transmission. Phenotypic and genetic changes in some parasite populations that have undergone multiple ivermectin treatments in Cameroon and Ghana have raised concern that sub-optimal response to ivermectins anti-fecundity effect may increase in frequency, reducing the impact of ivermectin-based control measures. We used next generation sequencing of small pools of parasites to define genome-wide genetic differences between phenotypically characterised good and sub-optimal responder parasites from Cameroon and Ghana, and identified multiple genomic regions differentiating the response types. These regions were largely different between parasites from both countries but revealed common molecular pathways that might be involved in determining the extent of response to ivermectins anti-fecundity effect. These data reveal a more complex than previously described pattern of genetic diversity among O. volvulus populations that differ in their geography and response to ivermectin treatment.

genomics

SweGen: A whole-genome map of genetic variability in a cross-section of the Swedish population

Here we describe the SweGen dataset, a high-quality map of genetic variation in the Swedish population. This data represents a basic resource for clinical genetics laboratories as well as for sequencing-based association studies, by providing information on the frequencies of genetic variants in a cohort that is well matched to national patient cohorts. To select samples for this study, we first examined the genetic structure of the Swedish population using high-density SNP-array data from a nation-wide population based cohort of over 10,000 individuals. From this sample collection, 1,000 individuals, reflecting a cross-section of the population and capturing the main genetic structure, were selected for whole genome sequencing (WGS). Analysis pipelines were developed for automated alignment, variant calling and quality control of the sequencing data. This resulted in a whole-genome map of aggregated variant frequencies in the Swedish population that we hereby release to the scientific community.

genetics