bioRxiv ScienceSearch

Biology subjects

Lee, J.

Publications and source records attributed to Lee, J..

At least 55 records · Page 3Linked to original sources

Improved Aedes aegypti mosquito reference genome assembly enables biological discovery and vector control

Female Aedes aegypti mosquitoes infect hundreds of millions of people each year with dangerous viral pathogens including dengue, yellow fever, Zika, and chikungunya. Progress in understanding the biology of this insect, and developing tools to fight it, has been slowed by the lack of a high-quality genome assembly. Here we combine diverse genome technologies to produce AaegL5, a dramatically improved and annotated assembly, and demonstrate how it accelerates mosquito science and control. We anchored the physical and cytogenetic maps, resolved the size and composition of the elusive sex-determining \"M locus\", significantly increased the known members of the glutathione-S-transferase genes important for insecticide resistance, and doubled the number of chemosensory ionotropic receptors that guide mosquitoes to human hosts and egg-laying sites. Using high-resolution QTL and population genomic analyses, we mapped new candidates for dengue vector competence and insecticide resistance. We predict that AaegL5 will catalyse new biological insights and intervention strategies to fight this deadly arboviral vector.

genomics

Directed evolution of CRISPR-Cas9 to increase its specificity

The use of CRISPR-Cas9 as a therapeutic reagent is hampered by its off-target effects. Although rationally designed S. pyogenes Cas9 (SpCas9) variants that display higher specificities than the wild-type SpCas9 protein are available, these attenuated Cas9 variants are often poorly efficient in human cells. Here, we have used a directed evolution approach in E. coli to obtain Sniper-Cas9, which shows high specificities without sacrificing on-target activities in human cells.

biochemistry

A modular transcriptional signature identifies phenotypic heterogeneity of human tuberculosis infection

Whole blood transcriptional signatures distinguishing active tuberculosis patients from asymptomatic latently infected individuals exist. Consensus has not been achieved regarding the optimal reduced gene sets as diagnostic biomarkers that also achieve discrimination from other diseases. Here we show a blood transcriptional signature of active tuberculosis using RNA-Seq, confirming microarray results, that discriminates active tuberculosis from latently infected and healthy individuals, validating this signature in an independent cohort. Using an advanced modular approach, we utilise information from the entire transcriptome, which includes over-abundance of type I interferon-inducible genes and under-abundance of IFNG and TBX21, to develop a signature that discriminates active tuberculosis patients from latently infected individuals, or those with acute viral and bacterial infections. We suggest methods targeting gene selection across multiple discriminant modules can improve development of diagnostic biomarkers with improved performance. Finally, utilising the modular approach we demonstrate dynamic heterogeneity in a longitudinal study of recent tuberculosis contacts.

immunology

Machine Learning of the Cardiac Phenome and Skin Transcriptome to Categorize Heart Disease in Systemic Sclerosis

BackgroundCardiac involvement is a leading cause of death in systemic sclerosis (SSc/scleroderma). The complexity of SSc cardiac manifestations is not fully captured by the current clinical SSc classification, which is based on extent of skin involvement and specific autoantibodies. Therefore, we sought to develop a clinically relevant SSc cardiac disease classification to improve clinical care and increase understanding of SSc cardiac disease pathobiology. We hypothesized that machine learning could identify novel SSc cardiac disease subgroups, and that gene expression assessment of skin could provide insights into molecular pathogenesis of these SSc pheno-groups.\n\nMethodsWe used unsupervised model-based clustering (phenomapping) of SSc patient echocardiographic and clinical data to identify clinically relevant SSc pheno-groups in a discovery cohort (n=316), and validated these findings in an external SSc validation cohort (n=67). Cox regression was used to evaluate survival differences among groups. Gene expression profiles from skin biopsies from a subset of SSc patients (n=68) and controls (n=18) were analyzed with weighted gene co-expression network analyses to identify gene modules that were associated with cardiac pheno-groups and echocardiographic parameters.\n\nResultsFour SSc cardiac pheno-groups were identified with distinct profiles. Pheno-group #1 displayed a predominant cutaneous phenotype without cardiac involvement; pheno-group #2 had long-standing SSc with limited skin and cardiac involvement; pheno-group #3 had diffuse skin involvement, a high frequency of interstitial lung disease (88%), and significant right heart remodeling/dysfunction; and pheno-group #4 had prolonged SSc disease duration, limited skin involvement, and marked biventricular cardiac involvement. After multivariable adjustment, pheno-group #3 (hazard ratio [HR] 7.8, 95% confidence interval [CI] 1.5-33.0) and pheno-group #4 (HR 10.5, 95% CI 2.1-52.7) remained associated with mortality (P<0.05). The addition of pheno-group classification was additive to conventional survival models (P<0.05 by likelihood ratio test for all models), a finding that was replicated in the validation cohort. Skin gene expression analysis identified 2 gene modules (representing fibrosis and skin integrity, respectively) that differed among the cardiac pheno-groups and were associated with specific echocardiographic parameters.\n\nConclusionsMachine learning of echocardiographic and skin gene expression data in SSc identifies clinically relevant subgroups with distinct cardiac phenotypes, survival, and associated molecular pathways in skin.

genomics

Biopipe: A Lightweight System Enabling Comparison of Bioinformatics Tools and Workflows

Analyzing next generation sequencing data always requires researchers to install many tools, prepare input data compliant to the required data format, and execute the tools in specific orders. Such tool installation and workflow execution process is tedious and error-prone, and becomes very challenging when researchers need to compare multiple alternative tool chains. To mitigate this problem, we developed a new lightweight and portable system, Biopipe, to simplify the creation and execution of bioinformatics tools and workflows, and to further enable the comparison between alternative tools or workflows. Biopipe allows users to create and edit workflows with user-friendly web interfaces, and automates tool installation as well as workflow synthesis by downloading and executing predefined Docker images. With Biopipe, biologists can easily experiment with and compare different bioinformatics tools and workflows without much computer science knowledge. There are mainly two parts in Biopipe: a web application and a standalone Java application. They are freely available at http://bench.cs.vt.edu:8282/Biopipe-Workflow-Editor-0.0.1/index.xhtml and https://code.vt.edu/saima5/Biopipe-Run-Workflow\n\nContactnm8247@cs.vt.edu\n\nSupplementary informationSupplementary data are available online.

bioinformatics

Multi-platform discovery of haplotype-resolved structural variation in human genomes

The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, and strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three human parent-child trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs ([&ge;]50 bp) per human genome. We also discover 156 inversions per genome--most of which previously escaped detection. Fifty-eight of the inversions we discovered intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The method and the dataset serve as a gold standard for the scientific community and we make specific recommendations for maximizing structural variation sensitivity for future large-scale genome sequencing studies.

genomics

Anoctamin 9/TMEM16J is a Cation Channel Activated by cAMP/PKA Signal

Anoctamins are membrane proteins that consist of 10 homologs. ANO1 and ANO2 are anion channels activated by intracellular calcium that meditate numerous physiological functions. ANO6 is a scramblase that redistributes phospholipids across the cell membrane. However, the others are not well characterized. We found ANO9/TMEM16J is a cation channel activated by a cAMP-dependent PKA. Intracellular cAMP activated robust currents in whole-cells expressing ANO9 and inhibited by PKA blockers. A cholera toxin and purified PKA also activated ANO9. The cAMP-induced ANO9 currents were permeable to cations. The mutation of a possible phosphorylation site at Ser245 elicited a block of the cAMP-dependent activation. High levels of Ano9 transcripts were found in intestines. Human intestinal SW480 cells showed cAMP-dependent currents. We conclude that ANO9 is a cation channel activated by the cAMP/PKA pathway and could play a role in intestine function.

biophysics

Genome expansion and lineage-specific genetic innovations in the world’s largest organisms (Armillaria)

Armillaria species are both devastating forest pathogens and some of the largest terrestrial organisms on Earth. They forage for hosts and achieve immense colony sizes using rhizomorphs, root-like multicellular structures of clonal dispersal. Here, we sequenced and analyzed genomes of four Armillaria species and performed RNA-Seq and quantitative proteomic analysis on seven invasive and reproductive developmental stages of A. ostoyae. Comparison with 22 related fungi revealed a significant genome expansion in Armillaria, affecting several pathogenicity-related genes, lignocellulose degrading enzymes and lineage-specific genes likely involved in rhizomorph development. Rhizomorphs express an evolutionarily young transcriptome that shares features with the transcriptomes of fruiting bodies and vegetative mycelia. Several genes show concomitant upregulation in rhizomorphs and fruiting bodies and shared cis-regulatory signatures in their promoters, providing genetic and regulatory insights into complex multicellularity in fungi. Our results suggest that the evolution of the unique dispersal and pathogenicity mechanisms of Armillaria might have drawn upon ancestral genetic toolkits for wood-decay, morphogenesis and complex multicellularity.

evolutionary biology

The evolutionary history of 2,658 cancers

Cancer develops through a process of somatic evolution. Here, we use whole-genome sequencing of 2,778 tumour samples from 2,658 donors to reconstruct the life history, evolution of mutational processes, and driver mutation sequences of 39 cancer types. The early phases of oncogenesis are driven by point mutations in a small set of driver genes, often including biallelic inactivation of tumour suppressors. Early oncogenesis is also characterised by specific copy number gains, such as trisomy 7 in glioblastoma or isochromosome 17q in medulloblastoma. By contrast, increased genomic instability, a nearly four-fold diversification of driver genes, and an acceleration of point mutation processes are features of later stages. Copy-number alterations often occur in mitotic crises leading to simultaneous gains of multiple chromosomal segments. Timing analysis suggests that driver mutations often precede diagnosis by many years, and in some cases decades, providing a window of opportunity for early cancer detection.

cancer biology

Habitat preference of an herbivore shapes the habitat distribution of its host plant

Plant distributions can be limited by habitat-biased herbivory, but the proximate causes of such biases are rarely known. Distinguishing plant-centric from herbivore-centric mechanisms driving differential herbivory between habitats is difficult without experimental manipulation of both plants and herbivores. Here we tested alternative hypotheses driving habitat-biased herbivory in bittercress (Cardamine cordifolia), which is more abundant under shade of shrubs and trees (shade) than in nearby meadows (sun) where herbivory is intense from the specialist fly Scaptomyza nigrita. This system has served as a textbook example of habitat-biased herbivory driving a plants distribution across an ecotone, but the proximate mechanisms underlying differential herbivory are still unclear. First, we found that higher S. nigrita herbivory in sun habitats contrasts sharply with their preference to attack plants from shade habitats in laboratory choice experiments. Second, S. nigrita strongly preferred leaves in simulated sun over simulated shade habitats, regardless of plant source habitat. Thus, herbivore preference for brighter, warmer habitats overrides their preference for more palatable shade plants. This promotes the sun-biased herbivore pressure that drives the distribution of bittercress into shade habitats.

ecology

Welfare of zebra finches used in research

Over the past 50 years, songbirds have become a valuable model organism for scientists studying vocal communication from its behavioral, hormonal, neuronal, and genetic perspectives. Many advances in our understanding of vocal learning result from research using the zebra finch, a close-ended vocal learner. We review some of the manipulations used in zebra finch research, such as isolate housing, transient/irreversible impairment of hearing/vocal organs, implantation of small devices for chronic electrophysiology, head fixation for imaging, aversive song conditioning using sound playback, and mounting of miniature backpacks for behavioral monitoring. We highlight the use of these manipulations in scientific research, and estimate their impact on animal welfare, based on the literature and on data from our past and ongoing work. The assessment of harm-benefits tradeoffs is a legal prerequisite for animal research in Switzerland. We conclude that a diverse set of known stressors reliably lead to suppressed singing rate, and that by contraposition, increased singing rate may be a useful indicator of welfare. We hope that our study can contribute to answering some of the most burning questions about zebra finch welfare in research on vocal behaviors.

animal behavior and cognition

YASS: Yet Another Spike Sorter

Spike sorting is a critical first step in extracting neural signals from large-scale electrophysiological data. This manuscript describes an efficient, reliable pipeline for spike sorting on dense multi-electrode arrays (MEAs), where neural signals appear across many electrodes and spike sorting currently represents a major computational bottleneck. We present several new techniques that make dense MEA spike sorting more robust and scalable. Our pipeline is based on an efficient multi-stage \"triage-then-cluster-then-pursuit\" approach that initially extracts only clean, high-quality waveforms from the electrophysiological time series by temporarily skipping noisy or \"collided\" events (representing two neurons firing synchronously). This is accomplished by developing a neural network detection method followed by efficient outlier triaging. The clean waveforms are then used to infer the set of neural spike waveform templates through nonparametric Bayesian clustering. Our clustering approach adapts a \"coreset\" approach for data reduction and uses efficient inference methods in a Dirichlet process mixture model framework to dramatically improve the scalability and reliability of the entire pipeline. The \"triaged\" waveforms are then finally recovered with matching-pursuit deconvolution techniques. The proposed methods improve on the state-of-the-art in terms of accuracy and stability on both real and biophysically-realistic simulated MEA data. Furthermore, the proposed pipeline is efficient, learning templates and clustering much faster than real-time for a [~=] 500-electrode dataset, using primarily a single CPU core.

neuroscience

High Performance Virtual Screening by Targeting a High-resolution RNA Dynamic Ensemble

Dynamic ensembles that capture the flexibility of RNA three-dimensional (3D) structures hold great promise in advancing RNA-targeted drug discovery. Here, we experimentally screened the transactivation response element (TAR) RNA from human immunodeficiency virus type-1 (HIV-1) against ~100,000 small molecules. We used this dataset, along with 240 known hit molecules, to evaluate virtual screening (VS) against a high-resolution TAR ensemble determined by combining NMR spectroscopy and molecular dynamics (MD) simulations. Ensemble-based VS (EBVS) scores molecules with an area under the receiver operator characteristic curve (ROC AUC) of 0.87 with ~50% of all hits falling within the top 2% of scored molecules, and also correctly predicts the different TAR inter-helical structures when bound to six molecules. The prediction accuracy decreased significantly with decreasing accuracy of the target ensemble or when docking against a single RNA structure. These results demonstrate that experime ...

biochemistry

Methods For Estimation Of Model Accuracy In CASP12

Methods for reliably estimating the quality of 3D models of proteins are essential drivers for the wide adoption and serious acceptance of protein structure predictions by life scientists. In this paper, the most successful groups in CASP12 describe their latest methods for Estimates of Model Accuracy (EMA). We show that pure single model accuracy estimation methods has shown clear progress since CASP11; the three top methods (MESHI, ProQ3, SVMQA) all perform better than the top method of CASP11 (ProQ2). The pure single model accuracy estimation methods outperform quasi-single (ModFOLD6 variations) and consensus methods (Pcons, ModFOLDclust2, Pcomb-domain and Wallner) in model selection, but are still not as good as those methods in absolute model quality estimation and predictions of local quality. Finally, we show that when using contact based model quality measures (CAD, 1DDT) the single model quality methods perform relatively better.

bioinformatics

The LDB1 complex co-opts CTCF for erythroid lineage specific long-range enhancer interactions

Lineage-specific transcription factors are critical for long-range enhancer interactions but direct or indirect contributions of architectural proteins such as CTCF to enhancer function remain less clear. The LDB1 complex mediates enhancer-gene interactions at the {beta}-globin locus through LDB1 self-interaction. We find that a novel LDB1-bound enhancer upstream of carbonic anhydrase 2 (Car2) activates its expression by interacting directly with CTCF at the gene promoter. Both LDB1 and CTCF are required for enhancer-Car2 looping and the domain of LDB1 contacted by CTCF is necessary to rescue Car2 transcription in LDB1 deficient cells. Genome wide studies and CRISPR/Cas9 genome editing indicate that LDB1-CTCF enhancer looping underlies activation of a substantial fraction of erythroid genes. Our results provide a mechanism by which long-range interactions of architectural protein CTCF can be tailored to achieve a tissue-restricted pattern of chromatin loops and gene expression.

molecular biology

Heritable Small RNAs Regulate Nematode Benzimidazole Resistance

Parasitic nematodes impose a debilitating health and economic burden across much of the world. Nematode resistance to anthelmintic drugs threatens parasite control efforts in both human and veterinary medicine. Despite this threat, the genetic landscape of potential resistance mechanisms to these critical drugs remains largely unexplored. Here, we exploit natural variation in the model nematodes Caenorhabditis elegans and Caenorhabditis briggsae to discover quantitative trait loci (QTL) that control sensitivity to benzimidazoles widely used in human and animal medicine. High-throughput phenotyping of albendazole, fenbendazole, mebendazole, and thiabendazole responses in panels of recombinant lines led to the discovery of over 15 QTL in C. elegans and four QTL in C. briggsae associated with divergent responses to these anthelmintics. Many of these QTL are conserved across benzimidazole derivatives, but others show drug and dose specificity. We used near-isogenic lines to recapitulate and narrow the C. elegans albendazole QTL of largest effect and identified candidate variants correlated with the resistance phenotype. These QTL do not overlap with known benzimidazole resistance genes from parasitic nematodes and present specific new leads for the discovery of novel mechanisms of nematode benzimidazole resistance. Analyses of orthologous genes reveal significant conservation of candidate benzimidazole resistance genes in medically important parasitic nematodes. These data provide a basis for extending these approaches to other anthelmintic drug classes and a pathway towards validating new markers for anthelmintic resistance that can be deployed to improve parasite disease control.\n\nAuthor SummaryThe treatment of roundworm (nematode) infections in both humans and animals relies on a small number of anti-parasitic drugs. Resistance to these drugs has appeared in veterinary parasite populations and is a growing concern in human medicine. A better understanding of the genetic basis for parasite drug resistance can be used to help maintain the effectiveness of anti-parasitic drugs and to slow or to prevent the spread of drug resistance in parasite populations. This goal is hampered by the experimental intractability of nematode parasites. Here, we use non-parasitic model nematodes to systematically explore responses to the critical benzimidazole class of anti-parasitic compounds. Using a quantitative genetics approach, we discovered unique genomic intervals that control drug effects, and we identified differences in the genetic architectures of drug responses across compounds and doses. We were able to narrow a major-effect genomic region associated with albendazole resistance and to establish that candidate genes discovered in our genetic mappings are largely conserved in important human and animal parasites. This work provides new leads for understanding parasite drug resistance and contributes a powerful template that can be extended to other anti-parasitic drug classes.

genetics

Multiple reference genome sequences of hot pepper reveal the massive evolution of plant disease resistance genes by retroduplication

Transposable elements (TEs) provide major evolutionary forces leading to new genome structure and species diversification. However, the role of TEs in the expansion of disease resistance gene families has been unexplored in plants. Here, we report high-quality de novo genomes for two peppers (Capsicum baccatum and C. chinense) and an improved reference genome (C. annuum). Dynamic genome rearrangements involving translocations among chromosome 3, 5 and 9 were detected in comparison between C. baccatum and the two other peppers. The amplification of athila LTR-retrotransposons, members of the gypsy superfamily, led to genome expansion in C. baccatum. In-depth genome-wide comparison of genes and repeats unveiled that the copy numbers of NLRs were greatly increased by LTR-retrotransposon-mediated retroduplication. Moreover, retroduplicated NLRs exhibited great abundance across the angiosperms, with most cases lineage-specific and thus recent events. Our study revealed that retroduplication has played key roles in the emergence of new disease-resistance genes in plants.

plant biology

Flexibility to contingency changes distinguishes habitual and goal-directed strategies in humans

Decision-making in the real world presents the challenge of requiring flexible yet prompt behavior, a balance that has been characterized in terms of a trade-off between a slower, prospective goal-directed model-based (MB) strategy and a fast, retrospective habitual model-free (MF) strategy. Theory predicts that flexibility to changes in both reward values and transition contingencies can determine the relative influence of the two systems in reinforcement learning, but few studies have manipulated the latter. Therefore, we developed a novel two-level contingency change task in which transition contingencies between states change every few trials; MB and MF control predict different responses following these contingency changes, allowing their relative influence to be inferred. Additionally, we manipulated the rate of contingency changes in order to determine whether contingency change volatility would play a role in shifting subjects between a MB and MF strategy. We found that human subjects employed a hybrid MB/MF strategy on the task, corroborating the parallel contribution of MB and MF systems in reinforcement learning. Further, subjects did not remain at one level of MB/MF behavior but rather displayed a shift towards more MB behavior over the first two blocks that was not attributable to the rate of contingency changes but rather to the extent of training. We demonstrate that flexibility to contingency changes can distinguish MB and MF strategies, with human subjects utilizing a hybrid strategy that shifts towards more MB behavior over blocks, consequently corresponding to a higher payoff.\n\nAuthor SummaryTo make good decisions, we must learn to associate actions with their true outcomes. Flexibility to changes in action/outcome relationships, therefore, is essential for optimal decision-making. For example, actions can lead to outcomes that change in value - one day, your favorite food is poorly made and thus less pleasant. Alternatively, changes can occur in terms of contingencies - ordering a dish of one kind and instead receiving another. How we respond to such changes is indicative of our decision-making strategy; habitual learners will continue to choose their favorite food even if the quality has gone down, whereas goal-directed learners will soon learn it is better to choose another dish. A popular paradigm probes the effect of value changes on decision making, but the effect of contingency changes is still unexplored. Therefore, we developed a novel task to study the latter. We find that humans used a mixed habitual/goal-directed strategy in which they became more goal-directed over the course of the task, and also earned more rewards with increasing goal-directed behavior. This shows that flexibility to contingency changes is adaptive for learning from rewards, and indicates that flexibility to contingency changes can reveal which decision-making strategy is used.

animal behavior and cognition