bioRxiv ScienceSearch

Biology subjects

Tan, P.

Publications and source records attributed to Tan, P..

9 recordsLinked to original sources

Large-scale whole-genome sequencing of three diverse Asian populations in Singapore

Asian populations are currently underrepresented in human genetics research. Here we present whole-genome sequencing data of 4,810 Singaporeans from three diverse ethnic groups: 2,780 Chinese, 903 Malays, and 1,127 Indians. Despite a medium depth of 13.7x, we achieved essentially perfect (>99.8%) sensitivity and accuracy for detecting common variants and good sensitivity (>89%) for detecting extremely rare variants with <0.1% allele frequency. We found 89.2 million single-nucleotide polymorphisms (SNPs) and 9.1 million small insertions and deletions (INDELs), more than half of which have not been cataloged in dbSNP. In particular, we found 126 common deleterious mutations (MAF>0.01) that were absent in the existing public databases, highlighting the importance of local population reference for genetic diagnosis. We describe fine-scale genetic structure of Singapore populations and their relationship to worldwide populations from the 1000 Genomes Project. In addition to revealing noticeable amounts of admixture among three Singapore populations and a Malay-related novel ancestry component that has not been captured by the 1000 Genomes Project, our analysis also identified some fine-scale features of genetic structure consistent with two waves of prehistoric migration from south China to Southeast Asia. Finally, we demonstrate that our data can substantially improve genotype imputation not only for Singapore populations, but also for populations across Asia and Oceania. These results highlight the genetic diversity in Singapore and the potential impacts of our data as a resource to empower human genetics discovery in a broad geographic region.

genetics

Whole-Genome Genomics Correlates of Response To Anti-PD1 Therapy in Relapsed/Refractory Natural Killer/T Cell Lymphoma

AbstractThis study aims to identify recurrent genetic alterations in relapsed or refractory (RR) natural-killer/T-cell lymphoma (NKTL) patients who have achieved complete response (CR) with programmed cell death 1 (PD-1) blockade therapy. Seven of the eleven patients treated with pembrolizumab achieved CR while the remaining four had progressive disease (PD). Using whole genome sequencing (WGS), we found recurrent clonal structural rearrangements (SR) of the PD-L1 gene in four of the seven (57%) CR patients pretreated tumors. These PD-L1 SRs consist of inter-chromosomal translocations, tandem duplication and micro-inversion that disrupted the suppressive function of PD-L1 3UTR. Interestingly, recurrent JAK3-activating (p.A573V) mutations were also validated in two CR patients tumors that did not harbor the PD-L1 SR. Importantly, these mutations were absent in the four PD cases. With immunohistochemistry (IHC), PD-L1 positivity could not discriminate patients who archived CR (range: 6%-100%) from patients who had PD (range: 35%-90%). PD-1 blockade with pembrolizumab is a potent strategy for RR NKTL patients and genomic screening could potentially accompany PD-L1 IHC positivity to better select patients for anti-PD-1 therapy.

genomics

Using machine learning to guide targeted and locally-tailored empiric antibiotic prescribing in a children’s hospital in Cambodia

BackgroundEarly and appropriate empiric antibiotic treatment of patients suspected of having sepsis is associated with reduced mortality. The increasing prevalence of antimicrobial resistance risks eroding the benefits of such empiric therapy. This problem is particularly severe for children in developing country settings. We hypothesized that by applying machine learning approaches to readily collected patient data, it would be possible to obtain actionable and patient-specific predictions for antibiotic-susceptibility. If sufficient discriminatory power can be achieved, such predictions could lead to substantial improvements in the chances of choosing an appropriate antibiotic for empiric therapy, while minimizing the risk of increased selection for resistance due to use of antibiotics usually held in reserve.\n\nMethods and FindingsWe analyzed blood culture data collected from a 100-bed childrens hospital in North-West Cambodia between February 2013 and January 2016. Clinical, demographic and living condition information for each child was captured with 35 independent variables. Using these variables, we used a suite of machine learning algorithms to predict Gram stains and whether bacterial pathogens could be treated with standard empiric antibiotic therapies: i) ampicillin and gentamicin; ii) ceftriaxone; iii) at least one of the above.\n\n243 cases of bloodstream infection were available for analysis. We used 195 (80%) to train the algorithms, and 48 (20%) for evaluation. We found that the random forest method had the best predictive performance overall as assessed by the area under the receiver operating characteristic curve (AUC), though support vector machine with radial kernel had similar performance for predicting Gram stain and ceftriaxone susceptibility. Predictive performance of logistic regression, simple and boosted decision trees and k-nearest neighbors were poor in comparison. The random forest method gave an AUC of 0.91 (95%CI 0.81-1.00) for predicting susceptibility to ceftriaxone, 0.75 (0.60-0.90) for susceptibility to ampicillin and gentamicin, 0.76 (0.59-0.93) for susceptibility to neither, and 0.69 (0.53-0.85) for Gram stain result. The most important variables for predicting susceptibility were time from admission to blood culture, patient age, hospital versus community-acquired infection, and age-adjusted weight score.\n\nConclusionsApplying machine learning algorithms to patient data that are readily available even in resource-limited hospital settings can provide highly informative predictions on susceptibilities of pathogens to guide appropriate empiric antibiotic therapy. Used as a decision support tool, such approaches have the potential to lead to better targeting of empiric therapy, improve patient outcomes and reduce the burden of antimicrobial resistance.\n\nAuthor summaryO_LSTWhy was this study done?C_LSTO_LIEarly and appropriate antibiotic treatment of patients with life-threatening bacterial infections is thought to reduce the risk of mortality.\nC_LIO_LIIn hospitals that have a microbiology laboratory, it takes 3-4 days to get results which indicate which antibiotics are likely to be effective; before this information is available antibiotics have to be prescribed empirically i.e. without knowledge of the causative organism.\nC_LIO_LIIncreasing resistance to antibiotics amongst bacteria makes finding an appropriate antibiotic to use empirically difficult; this problem is particularly severe for children in developing country settings.\nC_LIO_LIIf we could predict which antibiotics were likely to be effective at the time of starting antibiotic therapy, we might be able to improve patient outcomes and reduce resistance.\nC_LI\n\nO_LSTWhat Did the Researchers Do and Find?C_LSTO_LIWe evaluated the ability of a number of different algorithms (i.e. sets of step-by-step instructions) to predict susceptibility to commonly-used antibiotics using routinely available patient data from a childrens hospital in Cambodia.\nC_LIO_LIWe found that an algorithm called random forests enabled surprisingly accurate predictions, particularly for predicting whether the infection was likely to be treatable with ceftriaxone, the most commonly used empiric antibiotic at the study hospital.\nC_LIO_LIUsing this approach it would be possible to correctly predict when a different antibiotic would be needed for empiric treatment over 80% of the time, while recommending a different antibiotic when ceftriaxone would suffice less than 20% of the time.\nC_LI\n\nO_LSTWhat Do These Findings Mean?C_LSTO_LIUsing readily available patient information, sophisticated algorithms can enable good predictions of whether antibiotics are likely to be effective several days before laboratory tests are available.\nC_LIO_LIAlgorithms would need to be trained with local hospital data, but our study shows that even with relatively limited data from a small hospital, good predictions can be obtained.\nC_LIO_LIUsed as part of a decision support system such algorithms could help choose appropriate antibiotics for empiric therapy; this would be expected to translate into better patient outcomes and may help to reduce resistance.\nC_LIO_LISuch as a decision support system would have very low costs and be easy to implement in low- and middle-income countries.\nC_LI

epidemiology

Naa10p promotes metastasis by stabilizing matrix metalloproteinase-2 protein in human osteosarcomas

N--Acetyltransferase 10 protein (Naa10p) mediates N-terminal acetylation of nascent proteins. Oncogenic or tumor suppressive roles of Naa10p were reported in cancers. Here, we report an oncogenic role of Naa10p in promoting metastasis of osteosarcomas. Higher NAA10 transcripts were observed in metastatic osteosarcoma tissues compared to non-metastatic tissues and were also correlated with a worse prognosis of patients. Knockdown and overexpression of Naa10p in osteosarcoma cells respectively led to decreased and increased cell migratory/invasive abilities. Re-expression of Naa10p, but not an enzymatically inactive mutant, relieved suppression of the invasive ability in vitro and metastasis in vivo imposed by Naa10p-knockdown. According to protease array screening, we identified that matrix metalloproteinase (MMP)-2 was responsible for the Naa10p-induced invasive phenotype. Naa10p was directly associated with MMP-2 protein through its acetyltransferase domain and maintained MMP-2 protein stability via NatA complex activity. MMP-2 expression levels were also significantly correlated with Naa10p levels in osteosarcoma tissues. These results reveal a novel function of Naa10p in the regulation of cell invasiveness by preventing MMP-2 protein degradation that is crucial during osteosarcoma metastasis.

cancer biology

Transmission patterns of hyper-endemic multi-drug resistant Klebsiella pneumoniae in a Cambodian neonatal unit: a longitudinal study with whole genome sequencing

BackgroundKlebsiella pneumoniae is an important and increasing cause of life-threatening disease in hospitalised neonates. Third generation cephalosporin resistance (3GC-R) is frequently a marker of multi-drug resistance, and can complicate management of infections. 3GC-R K. pneumoniae is hyper-endemic in many developing country settings, but its epidemiology is poorly understood and prospective studies of endemic transmission are lacking. We aimed to determine the transmission dynamics of 3GC-R K. pneumoniae in a newly opened neonatal unit (NU) in Cambodia.\n\nMethodsWe performed a prospective longitudinal study between September and November 2013. Rectal swabs from 37 consented patients were collected upon NU admission and every three days thereafter. Morphologically different colonies from swabs growing cefpodoxime-resistant K. pneumoniae were selected for whole-genome sequencing (WGS).\n\nResults32/37 (86%) patients screened positive for 3GC-R K. pneumoniae and 93 colonies from 119 swabs were sequenced. Isolates were resistant to a median of six (range 3-9) antimicrobials. WGS revealed high diversity; pairwise distances between isolates from the same patient were either 0-1 SNV or >1,000 SNVs; 19/32 colonized patients harboured K. pneumoniae colonies differing by >1000 SNVs. Diverse lineages accounted for 18 probable importations to the NU and nine probable transmission clusters involving 19/37 (51%) of screened patients. Median cluster size was 5 patients (range 3-9).\n\nConclusionsThe epidemiology of 3GC-R K. pneumoniae was characterised by multiple introductions and a dense network of cross-infection, with half of screened neonates part of a transmission cluster. Efforts to reduce the 3GC-R K. pneumoniae disease burden should consider targeting both processes.

epidemiology

Osteo-Oto-Hepato-Enteric Syndrome (O2HE) is caused by loss of function mutations in UNC45A

Despite the rapid discovery of genes for rare genetic disorders, we continue to encounter individuals presenting with hitherto unknown syndromic manifestations. Here, we have studied four affected people in three families presenting with cholestasis, congenital diarrhea, impaired hearing and bone fragility, a clinical entity we have termed O2HE (Osteo-Oto-Hepato-enteric) syndrome. Whole exome sequencing of all affected individuals and their parents identified biallelic mutations in Unc-45 Myosin Chaperone A (UNC45A), as a likely driver for this disorder. Subsequent in vitro and in vivo functional studies of the candidate gene indicated a loss of function paradigm, wherein mutations attenuated or abolished protein activity with concomitant defects in gut development and function.

genetics

A pan cancer analysis of promoter activity highlights the regulatory role of alternative transcription start sites and their association with noncoding mutations

Most human protein-coding genes are regulated by multiple, distinct promoters, suggesting that the choice of promoter is as important as its level of transcriptional activity. While the role of promoters as driver elements in cancer has been recognized, the contribution of alternative promoters to regulation of the cancer transcriptome remains largely unexplored. Here we infer active promoters using RNA-Seq data from 1,188 cancer samples with matched whole genome sequencing data. We find that alternative promoters are a major contributor to context-specific regulation of isoform expression and that alternative promoters are frequently deregulated in cancer, affecting known cancer-genes and novel candidates. Our study suggests that a highly dynamic landscape of active promoters shapes the cancer transcriptome, opening many opportunities to further explore the interplay of regulatory mechanism and noncoding somatic mutations with transcriptional aberrations in cancer.

genomics

Revealing unidentified heterogeneity in different epithelial cancers using heterocellular subtype classification

Cancers are currently diagnosed, categorised, and treated based on their tissue of origin. However, how different cellular compartments of tissues (e.g., epithelial, immune and stem cells) are similar across cancer types is unknown. Here we used colorectal cancer subtypes and their signatures representing different colonic crypt cell types as surrogates to classify different epithelial cancers into five heterotypic cellular (heterocellular) subtypes. The stem-like and inflammatory heterocellular subtypes are ubiquitous across epithelial cancers so capture intrinsic, tissue-independent properties. Conversely, well-differentiated/specialized goblet-like/enterocyte heterocellular subtypes differ across cancer types due to their colorectum-specific genes. The transit-amplifying heterocellular subtype shows a dynamic range of cellular differentiation with shared common pathways (Wnt, FGFR) in certain cancer types. Importantly, this approach revealed previously unrecognised heterogeneity in pancreatic, breast, microsatellite-instability enriched and KRAS mutation-dependent cancers. Immune cell-type differences are common and useful for patient stratification for immunotherapy. This unique approach identifies cell type-dependent but tissue-independent heterogeneity in different cancers for precision medicine.

bioinformatics

Genome-Scale Mutational Signatures Of Aflatoxin In Cells, Mice And Human Tumors

Aflatoxin B1 (AFB1) is a mutagen and IARC Group 1 carcinogen that causes hepatocellular carcinoma (HCC). Here we present the first whole genome data on the mutational signatures of AFB1 exposure from a total of > 40,000 mutations in four experimental systems: two different human cell lines, and in liver tumors in wild-type mice and in mice that carried a hepatitis B surface antigen transgene - this to model the multiplicative effects of aflatoxin exposure and hepatitis B in causing HCC. AFB1 mutational signatures from all four experimental systems were remarkably similar. We integrated the experimental mutational signatures with data from newly-sequenced HCCs from Qidong County, China, a region of well-studied aflatoxin exposure. This indicated that COSMIC mutational signature 24, previously hypothesized to stem from aflatoxin exposure, indeed likely represents AFB1 exposure, possibly combined with other exposures. Among published somatic mutation data, we found evidence of AFB1 exposure in 0.7% of HCCs treated in North America, 1% of HCCs from Japan, but 16% of HCCs from Hong Kong. Thus, aflatoxin exposure apparently remains a substantial public health issue in some areas. This aspect of our study exemplifies the promise of future widespread resequencing of tumor genomes in providing new insights into the contribution of mutagenic exposures to cancer incidence.

cancer biology