bioRxiv ScienceSearch

Biology subjects

Shetty, A.

Publications and source records attributed to Shetty, A..

5 recordsLinked to original sources

Cost Effective, Experimentally Robust Differential Expression Analysis for Human/Mammalian, Pathogen, and Dual-Species Transcriptomics

As sequencing read length has increased, researchers have quickly adopted longer reads for their experiments. Here, we examine host-pathogen interaction studies to assess if using longer reads is warranted. Six diverse datasets encountered in studies of host-pathogen interactions were used to assess what genomic attributes might affect the outcome of differential gene expression analysis including: gene density, operons, gene length, number of introns/exons, and intron length. Principal components analysis, hierarchical clustering with bootstrap support, and regression analyses of pairwise comparisons were undertaken on the same reads, looking at all combinations of paired and unpaired reads trimmed to 36,54,72, and 101-bp. For E coli, 36-bp single end reads performed as well as any other read length and as well as paired end reads. For all other comparisons, 54-bp and 72-bp reads were typically equivalent and different from 36-bp and 101-bp reads. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data. However, experiments with an a priori expected greater biological complexity had more variable results with all read lengths relative to those with decreased complexity. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results across all comparisons. However, using 36-bp single end reads may be desirable for bacterial samples, although possibly only if the transcriptional response is expected a priori to be robust.\n\nDATA SUMMARYO_LIThe human only CSHL Encode data set (1) was downloaded from ftp://hgdownload.cse.ucsc.edu/goldenPath/hgl9/encodeDCC/wgEncodeCshlLongRnaSeq/.\nC_LIO_LIThe data from mice vaginas infected with Candida albicans (2) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP057050).\nC_LIO_LIThe data from Aspergillus fumigatus cells in contact with human cells was downloaded from the SRA (url - https://www.ncbi.nlm.nih.gov/bioproject/399754).\nC_LIO_LIThe data from a strand-specific library from a study comparing C. albicans cells in contact with human cells with those in media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from C. albicans in culture media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from Escherichia coli grown in different media (4) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP056578).\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTAs sequencing technologies improve, sequencing costs decrease and read lengths increase. We examine host-pathogen interaction studies to assess if using these longer reads is warranted given their increased cost relative to using the same number of shorter reads. To this end we compared the use of various read lengths and read pairing for six diverse host-pathogen datasets with varying genomic attributes including: gene density, operons, gene length, number of introns/exons, and intron length. We find that in the bacterial sample, 36-bp single end reads performed as well as any other read length and as well as paired end reads. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results for all other comparisons. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data.

genomics

A plug-and-play system for enzyme production at commercially viable levels in fed-batch cultures of Escherichia coli BL21 (DE3)

Commercial exploitation of enzymes in biotransformation necessitates a robust method for enzyme production that yields high enzyme titer. Nitrilases are a family of hydrolases that can transform nitriles to enantiopure carboxylic acids, which are important pharmaceutical intermediates. Here, we report a fed-batch method that uses a defined medium and involves growth under carbon limiting conditions using DO-stat feeding approach combined with an optimized post-induction strategy, yielding high cell densities and maximum levels of active and soluble enzyme. This strategy affords strict control of nutrient feeding and growth rates, and ensures sustained protein synthesis over a longer period. The method was optimized for highest titer of nitrilase reported so far (247 kU/l) using recombinant E. coli expressing the Alcaligenes sp. ECU0401 nitrilase. The fed-batch protocol presented here can also be employed as template to produce a wide variety of enzymes with minimal modification, as demonstrated for alcohol dehydrogenase and formate dehydrogenase.

bioengineering

Targeted enrichment outperforms other enrichment techniques and enables more multi-species RNA-Seq analyses

Enrichment methodologies enable analysis of minor members in multi-species transcriptomic analyses. We compared standard enrichment of bacterial and eukaryotic mRNA to targeted enrichment with Agilent SureSelect (AgSS) capture for Brugia malayi, Aspergillus fumigatus, and the Wolbachia endosymbiont of B. malayi (wBm). Without introducing significant systematic bias, the AgSS quantitatively enriched samples, resulting in more reads mapping to the target organism. The AgSS-enriched libraries consistently had a positive linear correlation with its unenriched counterpart (r2=0.559-0.867). Up to a 2,242-fold enrichment of RNA from the target organism was obtained following a power law (r2=0.90), with the greatest fold enrichment achieved in samples with the largest ratio difference between the major and minor members. While using a single total library for prokaryote and eukaryote in a single sample could be beneficial for samples where RNA is limiting, we observed a decrease in reads mapping to protein coding genes and an increase of multi-mapping reads to rRNAs in AgSS enrichments from eukaryotic total RNA libraries as opposed to eukaryotic poly(A)-enriched libraries. Our results support a recommendation of using Agilent SureSelect targeted enrichment on poly(A)-enriched libraries for eukaryotic captures and total RNA libraries for prokaryotic captures to increase the robustness of multi-species transcriptomic studies.

genomics

Genome-wide association study of asthma in individuals of African ancestry reveals novel asthma susceptibility loci

BACKGROUNDAsthma is a complex disease with striking disparities across racial and ethnic groups, which may be partly attributable to genetic factors. One of the main goals of the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to discover genes conferring risk to asthma in populations of African descent.\n\nMETHODSWe performed a genome-wide meta-analysis of asthma across 11 CAAPA datasets (4,827 asthma cases and 5,397 controls), genotyped on the African Diaspora Power Chip (ADPC) and including existing GWAS array data. The genotype data were imputed up to a whole genome sequence reference panel from n=880 African ancestry individuals for a total of 61,904,576 SNPs. Statistical models appropriate to each study design were used to test for association, and results were combined using the weighted Z-score method. We also used admixture mapping as a complementary approach to identify loci involved in asthma pathogenesis in subjects of African ancestry.\n\nRESULTSSNPs rs787160 and rs17834780 on chromosome 2q22.3 were significantly associated with asthma (p=6.57 x 10-9 and 2.97 x 10-8, respectively). These SNPs lie in the intergenic region between the Rho GTPase Activating Protein 15 (ARHGAP15) and Glycosyltransferase Like Domain Containing 1 (GTDC1) genes. Four low frequency variants on chromosome 1q21.3, which may be involved in the \"atopic march\" and which are not polymorphic in Europeans, also showed evidence for association with asthma (1.18 x10-6 [≤] p [≤] 3.06 x10-6). SNP rs11264909 on chromosome 1q23.1, close to a region previously identified by the EVE asthma meta-analysis as having a putative African ancestry specific effect, only showed differences in counts in subjects homozygous for alleles of African ancestry. Admixture mapping also identified a significantly associated region on chromosome 6q23.2, which includes the Transcription Factor 21 (TCF21) gene, previously shown to be differentially expressed in bronchial tissues of asthmatics and non-asthmatics.\n\nCONCLUSIONSWe have identified a number of novel asthma association signals warranting further investigation.

bioinformatics

Dual RNA sequencing (dRNA-Seq) of bacteria and their host cells

Bacterial pathogens subvert host cells by manipulating cellular pathways for survival and replication; in turn, host cells respond to the invading pathogen through cascading changes in gene expression. Deciphering these complex temporal and spatial dynamics to identify novel bacterial virulence factors or host response pathways is crucial for improved diagnostics and therapeutics. Dual RNA sequencing (dRNA-Seq) has recently been developed to simultaneously capture host and bacterial transcriptomes from an infected cell. This approach builds on the high sensitivity and resolution of RNA-Seq technology and is applicable to any bacteria that interact with eukaryotic cells, encompassing parasitic, commensal or mutualistic lifestyles. We pioneered dRNA-Seq to simultaneously capture prokaryotic and eukaryotic expression profiles of cells infected with bacteria, using in vitro Chlamydia-infected epithelial cells as proof of principle. Here we provide a detailed laboratory and bioinformatics protocol for dRNA-seq that is readily adaptable to any host-bacteria system of interest.

genomics