bioRxiv ScienceSearch

Biology subjects

Chakraborty, S.

Publications and source records attributed to Chakraborty, S..

13 recordsLinked to original sources

Anemia Diagnosis on a Simple Paper-based Assay

In developing countries, the maternal and neonatal mortality rate is often affected by prenatal period anemia, a preventable and ubiquitous impairment attributed due to low hemoglobin (Hgb) concentration. We report the development of a simple, frugal (~ 0.02 $ per test), rapid and high fidelity paper-based colorimetric microfluidic device for point-of-care (POC) detection of anemia. We validate our findings with 32 blood samples collected from different patients covering a wide spectrum of anemia and subsequently, compare with standard pathological results measured using a hematology analyzer. POC based Hgb estimates are correlated with the pathological gold standard estimates of Hgb levels (r = 0.909), and the POC test method yielded similar sensitivity and specificity for detecting mild anemia (n = 8) (<11 g/dl) (sensitivity: 87.5%, specificity: 100 %) and for severe anemia (n = 3) (<7 g/dl) (sensitivity: 100 %, specificity: 100 %). The estimated Hgb levels are, within 1.5 g/dl from the pathological estimate, for 91 % of the blood samples. Results demonstrate the elevated efficacy and viability of this POC colorimetric diagnostic test, in comparison to the state-of-the-art complex and expensive diagnostic tests for anemia detection.

bioengineering

The wheat Sr22, Sr33, Sr35 and Sr45 genes confer resistance against stem rust in barley

In the last 20 years, stem rust caused by the fungus Puccinia graminis f. sp. tritici (Pgt), has re-emerged as a major threat to wheat and barley cultivation in Africa and Europe. In contrast to wheat with 82 designated stem rust (Sr) resistance genes, barleys genetic variation for stem rust resistance is very narrow with only seven resistance genes genetically identified. Of these, only one locus consisting of two genes is effective against Ug99, a strain of Pgt which emerged in Uganda in 1999 and has since spread to much of East Africa and parts of the Middle East. The objective of this study was to assess the functionality, in barley, of cloned wheat Sr genes effective against Ug99. Sr22, Sr33, Sr35 and Sr45 were transformed into barley cv. Golden Promise using Agrobacterium-mediated transformation. All four genes were found to confer effective stem rust resistance. The barley transgenics remained susceptible to the barley leaf rust pathogen Puccinia hordei, indicating that the resistance conferred by these wheat Sr genes was specific for Pgt. Cloned Sr genes from wheat are therefore a potential source of resistance against wheat stem rust in barley.

plant biology

Reverse transcriptase fused CRISPR-Cas1 locus with RNA-seq expression necessitates revisiting hypothesis on acquisition of antibiotic resistance genes in multidrug-resistant Enterococcus faecalis V583

The emergence of drug-resistance in Enterococcus faecalis V583 through acquisition of resistance genes has been correlated to the absence of CRISPR-loci. Here, the presence of a bona-fide CRISPR locus in E. faecalis V583 (Accid:NC_004668.1) at 2238156 with a single 20 nt repeat is demonstrated. The presence of a putative endonuclease Cas1 13538 nucleotides away from the repeat substantiates this claim. This Cas1 (628 aa) is highly homologous (Eval:5e-34) to a Cas1 from Pseudanabaena biceps (Accid:WP 009625648.1, 697 aa), which belongs to the enigmatic family of RT-CRISPR locus. Such significant similarity to a Cas protein, the presence of a topoisomerase, other DUF (domain of unknown function) proteins as is often seen in CRISPR loci, and other hypothetical proteins indicates that this is a bona-fide CRISPR locus. Further corroboration is provided by expression of both the repeat and the Cas1 gene in existing RNA-seq data (SRX3438611). Since so little is known of even well-studied species like E. faecalis V583 having many hypothetical proteins, computational absence of evidence should not be taken as evidence of absence (both crisprfinder and PILER-CR do not report this as a CRISPR locus). It is unlikely that bacteria would completely give up defense against its primeval enemies (viruses) to bolster its fight against the newly introduced antibiotics.

genomics

An atypical CRISPR-Cas locus in Symbiobacterium thermophilum flanked by a transposase, a reverse transcriptase, the endonuclease MutS2 and a putative Cas9-like protein

Clustered regularly interspaced short palindromic repeats (CRISPR) is a prokaryotic adaptive defense system that assimilates short sequences of invading genomes (spacers) within repeats, and uses nearby effector proteins (Cas), one of which is an endonuclease (Cas9), to cleave homologous nucleic acid during future infections from the same or closely related organisms. Here, a novel CRISPR locus with uncharacterized Cas proteins, is reported in Symbiobacterium thermophilum (Accid:NC 006177.1) around loc.1248561. Credence to this assertion is provided by four arguments. First, the presence of an exact repeat (CACGTGGGGTTCGGGTCGGACTG, 23 nucleotides) occurs eight times encompassing fragments about 83 nucleotides long. Second, comparison to a known CRISPR-Cas locus in the same organism (loc.355482) with an endonuclease Cas3 (WP 011194444.1, 729 aa) [~]10000 nt upstream shows the presence of a known MutS2 endonuclease (WP 011195247.1, 801 aa) in approximately the same distance in loc.1248561. Thirdly, and remarkably, an uncharacterized protein (1357 aa) long is uncannily close in length to known Cas9 proteins (1368 for Streptococcus pyogenes). Lastly, the presence of transposases and reverse transcriptase (RT) downstream of the repeat indicates this is one of an enigmatic RT-CRISPR locus, Also, the MutS2 endonuclease is not characterized as a CRISPR-endonuclease to the best of my knowledge. Interestingly, this locus was not among the four loci (three confirmed, one probable) reported by crisperfinder (http://crispr.i2bc.paris-saclay.fr/Server), indicating that the search algorithm needs to be revisited. This finding begs the question - how many such CRISPR-Cas loci and Cas9-like proteins lie undiscovered within bacterial genomes?

genomics

Detection and accurate False Discovery Rate control of differentially methylated regions from Whole Genome Bisulfite Sequencing

With recent advances in sequencing technology, it is now feasible to measure DNA methylation at tens of millions of sites across the entire genome. In most applications, biologists are interested in detecting differentially methylated regions, composed of multiple sites with differing methylation levels among populations. However, current computational approaches for detecting such regions do not provide accurate statistical inference. A major challenge in reporting uncertainty is that a genome-wide scan is involved in detecting these regions, which needs to be accounted for. A further challenge is that sample sizes are limited due to the costs associated with the technology. We have developed a new approach that overcomes these challenges and assesses uncertainty for differentially methylated regions in a rigorous manner. Region-level statistics are obtained by fitting a generalized least squares (GLS) regression model with a nested autoregressive correlated error structure for the effect of interest on transformed methylation proportions. We develop an inferential approach, based on a pooled null distribution, that can be implemented even when as few as two samples per population are available. Here we demonstrate the advantages of our method using both experimental data and Monte Carlo simulation. We find that the new method improves the specificity and sensitivity of list of regions and accurately controls the False Discovery Rate (FDR).

genomics

Ambiguous specification of EGFR mutations compounded by nil or negligible fragmented gene counts and erroneous application of the Kappa statistic reiterates doubts on the veracity of the TEP-study

Final amendment noteThis paper had raised two issues - the error-prone classification and mistaken application of the Kappa statistic. The classification critique still holds, and is being taken up with other criticisms at http://www.biorxiv.org/content/early/2017/07/02/146134. The Kappa statistic was an error on my part since I had failed to see another page in Table S1. Please consider this pre-print closed.\n\nOriginal abstractThe use of RNA-seq from tumor-educated platelets (TEP) as a liquid biopsy source [1] has been refuted recently (http://biorxiv.org/content/early/2017/06/05/146134, not peer-reviewed). The TEP-study also mentioned that mutant epidermal growth factor receptor (EGFR) was accurately distinguished using surrogate TEP mRNA profiles, which is contested here. It is shown that only 10 out of 24 (a smaller sample set, original study has 60) non-small cell lung carcinoma (NSCLC) samples here has any expression at all. Even there the number of reads (101 bp) are [1, 4, 1, 14, 9, 1, 2, 19, 21, 6], and do not even add up to one complete EGFR gene (about 6000 bp). EGFR mutations have been painstakingly collated in www.mycancergenome.org/content/disease/lung-cancer/egfr. In stark contrast, the TEP study has no specification of the EGFR mutant used. The TEP study found EGFR mutations in 17/21 (81%), and EGFR wild-type in 4/39 (10%) for NSCLC samples (Table S7, reflected in Fig 3, Panel E in percentages). A major flaw is the assumption that a non \"EGFR wild-type\" is a \"EGFR mutant\" since cases zero with EGFR reads (which are almost half of the samples) could be either. The application of the Kappa statistic to this data is erroneous for two reasons. First, the Kappa statistic does not handle \"unknowns\", as is the case for samples with zero expression. Secondly, interobserver variation can be measured in any situation in which two or more independent observers are evaluating the same thing [2]. The 90% (Fig 3, Panel E) is just the percentage of samples (35/39) that are not \"EGFT WT\" in one observation. It is not qualified to be in the Kappa matrix, where it translates to 35, leading to a Kappa=0.707, which implies \"substantial agreement\" [2]. The other observation (looking for EGFR mutation) is in a different set. To summarize, this work reiterates negligible expression of EGFR reads in NSCLC samples, and finds serious shortcomings in the statistical analysis of subsequent mutational analysis from these reads in the TEP-study.

genomics

A plausible explanation for in silico reporting of erroneous MET gene expression in tumor-educated platelets (TEP) intended for "liquid biopsy" of non-small cell lung carcinoma still refutes the TEP-study

Final amendment noteThis paper had proposed a plausible way for detecting large quantities of MET, which the authors have clarified was not done :the possible explanation proposed for this erroneous MET gene expression does bypass the filtering step we perform in the data processing pipeline, i.e. selection of intron-spanning reads, as can be read in the main text\" comments in http://www.biorxiv.org/content/early/2017/07/02/146134, where a continuing critique of the TEP study continues. Please consider this pre-print closed.\n\nOriginal abstractThe reported over-expression of MET genes in non-small cell lung carcinoma (NSCLC) from an analysis of the RNA-seq data from tumor-educated platelets (TEP), intended to supplement existing liquid biopsy techniques [1], has been refuted recently (http://biorxiv.org/content/early/2017/06/05/146134, not peer-reviewed). The MET proto-oncogene (Accid:NG 008996.1, RefSeqGene LRG 662 on chromosome 7, METwithintrons) encodes 21 exons resulting in a 6710 bps MET gene (Accid: NM 001127500.2, METonlyexons). METwithintrons has multiple matches in the RNA-seq derived reads of lung cancer samples (for example: SRR1982756.11853382). Unfortunately, these are non-specific sequences in the intronic regions, matching to multiple genes on different chromosomes with 100% identity (KIF6 on chr6, COL6A6 on chr3, MYO16 on chr13, etc. for SRR1982756.11853382). In contrast, METonlyexons has few matches in the reads, if at all [2]. However, even RNA-seq from healthy donors have similar matches for METwithintrons so the computation behind the over-expression statistic remains obscure, even if METwithintrons was used as the search gene. In summary, this work re-iterates the lack of reproducibility in the bioinformatic analysis that establishes TEP as a possible source for \"liquid biopsy\".

genomics

No evidence of MET and HER2 over-expression in non-small cell lung carcinoma and breast cancer, respectively, raises serious doubts on using RNA-seq profiles of tumor-educated platelets as a ‘liquid biopsy’ source

In this detailed critique of the study proposing using RNA-seq from tumor-educated platelets (TEP) as a liquid biopsy source [1], several flawed assumptions leave little biological basis behind the statistical computations. First, there is no supporting evidence provided for the FFPE based classification of METoverexpression and EGFR mutation on tumor-tissues. Considering that raw reads of MET expression in a subset of healthy [N=21, mean=112, sd=77] and NSCLC [N=24, mean=11, sd=12] samples (typically with millions of reads) translates into over-expression in reality, providing the data for such computations is vital for future validation. A similar criticism applies for classifying samples based on EGFR mutations (the study uses only exon 20 and 21 from a wide range of possible mutations) with negligible counts [N=24, mean=3, sd=6]. While Ofner et. al, 2017 faced major problems associated with FFPE DNA, it is also true that Fassunke, et al., 2015 found concordance in 26 out of 26 samples for EGFR mutations in another FFPE-based study. However, Fassunke, et al., 2015 have been meticulous in describing the EGFR amplicons (exon 18 and 19 are missing in the TEP-study). Any error in initial classification renders downstream computations error-prone. The low counts of MET in the RNA-seq firmly establishes that inclusion of genes with such low counts in the set of 1100 discriminatory genes (Table S4) makes no sense as the \"real\" counts could vary wildly. Yet, TRAT1 is an example of one discriminator gene with counts of healthy [N=21, mean=164, sd=375] and NSCLC [N=24, mean=53, sd=176]. There are many such genes which should be excluded. Moving on to a discriminator with high counts (F13A1) in both healthy [N=21, mean=28228, sd=48581] and NSCLC [N=24, mean=98336, sd=74574] samples, a bonafide platelet gene that \"encodes the coagulation factor XIII A subunit\". Platelets do not have a nucleus, and thus the blue-print (chromosomes and related machinery) for making or regulating mRNA. They are boot-strapped with mRNA, like F13A1, during origination and then just go on keep collecting mRNA during circulation (which is the premise of their use in liquid biopsy). The assumption that these genes are differentially spliced in huge numbers is highly speculative without providing experimental proof. The discovery of spliceosomes in anucleate platelets [2] in 2005, 30 years after splicing was discovered in the nucleus by Sharp and Robert, probably indicates that spliceosomes are not dominant in platelets. Zucker, et al., 2017 have shown for another gene F11 that it is present in platelets as pre-mRNA and is spliced upon platelet activation [3]. Any study using the F13A1 gene as a discriminator ought to show the same two things, followed by differential counts in TEP. Ironically, F11 is not present in the discriminator set. Another blood coagulation related gene (TFPI) shows slight over-expression in NSCLC (moderate counts, healthy [N=21, mean=1352, sd=592] and NSCLC [N=24, mean=1854, sd=846]), agreeing with Iversen, et al., 1998 [4], but in contrast to Fei, et al., 2017 [5], demonstrating that the jury is still out on the levels of many such genes. Thus, circulating mRNA from tumor tissues are not discriminatoryif MET is degraded to such levels in platelets educated by NSCLC tumors, why not other possible mRNA that might have been picked during the same class? Furthermore, high count genes can only be bona-fide platelet genes, and have no supporting experimental proof of splicing differences (any one gene would suffice to instill some confidence). In conclusion, looking past the statistical smoke surrounding \"surrogate signatures\", one finds no biological relevance.

cancer biology

Cataloguing Over-Expressed Genes In Epstein Barr Virus Immortalized Lymphoblastoid Cell Lines Through Consensus Analysis Of PacBio Transcriptomes Corroborates Hypomethylation Of Chromosome 1

The ability of Epstein Barr Virus (EBV) to transform resting cell B-cells into immortalized lymphoblastoid cell lines (LCL) provides a continuous source of peripheral blood lymphocytes that are used to model conditions in which these lymphocytes play a key role. Here, the PacBio generated transcriptome of three LCLs from a parent-daughter trio (SRAid:SRP036136) provided by a previous study [1] were analyzed using a kmer-based version of YeATS (KEATS). The set of over-expressed genes in these cell lines were determined based on a comparison with the PacBio transcriptome of twenty tissues provided by another study (hOPTRS) [2]. MIR155 long non-coding RNA (MIR155HG), Fc fragment of IgE receptor II (FCER2), T-cell leukemia/lymphoma 1A (TCL1A), and germinal center associated signaling and motility (GCSAM) were genes having the highest expression counts in the three LCLs with no expression in hOPTRS. Other over-expressed genes, having low expression in hOPTRS, were membrane spanning 4-domains A1 (MS4A1) and ribosomal protein S2 pseudogene 55 (RPS2P55). While some of these genes are known to be over-expressed in LCLs, this study provides a comprehensive cataloguing of such genes. A recent work involving a patient with EBV-positive large B-cell lymphoma was unusually lacking various B-cell markers, but over-expressing CD30 [3] - a gene ranked 79 among uniquely expressed genes here. Hypomethylation of chromosome 1 observed in EBV immortalized LCLs [4, 5] is also corroborated here by mapping the genes to chromosomes. Extending previous work identifying un-annotated genes [6], 80 genes were identified which are expressed in the three LCLs, not in hOPTRS, and missing in the GENCODE, RefSeq and RefSeqGene databases. KEATS introduces a method of determining expression counts based on a partitioning of the known annotated genes, has runtimes of a few hours on a personal workstation and provides detailed reports enabling proper debugging.

genomics

Shorter unreported sequences in a RACE-Seq study involving seven tissues confirms ~150 novel transcripts identified in MCF-7 cell line PacBio transcriptome, leaving ~100 non-redundant transcripts exclusive to the cancer cell line

PacBio sequencing generates much longer reads compared to second-generation sequencing technologies, with a trade-off of lower throughput, higher error rate and more cost per base. The PacBio transcriptome of the breast cancer cell line MCF-7 was found to have [~]300 transcripts un-annotated in the current GENCODE (v25) or RefSeq, and missing in the liver, heart and brain PacBio transcriptomes [1]. RACE-sequencing (RACE-seq [2]) extends a well-established method of characterizing cDNA molecules generated by rapid amplification of cDNA ends (RACE [3]) using high-throughput sequencing technologies, reducing costs compared to PacBio. Here, shorter fragments of [~]150 transcripts were found to be present in seven tissues analyzed in a recent RACE-seq study (Accid:ERP012249) [4]. These transcripts were not among the [~]2500 novel transcripts reported in that study, tested separately here using the genomic coordinates provided, although all curated novel isoforms were incorporated into the human GENCODE set (v22) in that study. Non-redundancy analysis of the exclusive transcripts identified one transcript mapping to Chr1 with seven different splice variants, and erroneously mapped to Chr15 (PAC clone 15q11-q13) from the Prader-Willi/Angelman Syndrome region (Accid:AC004137.1). Finally, there are [~]100 non-redundant transcripts missing in the seven tissues, in addition to other three tissues analyzed previously. Their absence in GENCODE and RefSeq databases rule them out as commonly transcribed regions, further increasing their likelihood as biomarkers.

genomics

MCF-7 breast cancer cell line PacBio generated transcriptome has ~300 novel transcribed regions, un-annotated in both RefSeq and GENCODE, and absent in the liver, heart and brain transcriptomes

Illuminating the dark regions of the human genome remains an ongoing effort, a decade and a half after the human genome was sequenced - RefSeq and GENCODE being two of the major annotation databases. Pacific Biosciences (PacBio) has provided open access to the transcriptome of MCF-7, a breast cancer cell line that has provided significant therapeutic advancement in breast cancer research since the 1970s. PacBio sequencing generates much longer reads compared to second-generation sequencing technologies, with a trade-off of lower throughput, higher error rate and more cost per base. Here, this transcriptome was analyzed using the YeATS pipeline, with additionally introduced kmer based algorithms, reducing computational times to a few hours on a simple workstation. Out of ~300 transcripts that have no match in both RefSeq and GENCODE, ~250 are absent in the transcriptomes of the heart, liver and brain, also provided by PacBio. Also, ~200 transcripts are absent in a recent catalogue of un-annotated long non-coding RNAs from 6,503 samples (~43 Terabases of sequence data) [1], and among 2,556 novel transcripts reported in an experimental workflow RACE-Seq [2]. 65 transcripts have >100 amino acid open reading frames, and have the potential of being protein coding genes. ORF based annotation also identified few bacterial transcripts in the PacBio database mapped to the human genome, and one human transcript that has been annotated as bacterial in the NCBI database. The current work reiterates the under-utilization of transcriptomes for annotating genomes. It also provides new leads for investigating breast cancer by virtue of exclusively expressed transcripts not expressed in other tissues, which have the prospects of breast cancer biomarkers based on further investigations.

genomics

YeATSAM analysis of the chloroplast genome of walnut reveals several putative un-annotated genes and mis-annotation of the trans-spliced rps12 gene in other organisms

An open reading frame (ORF) is genomic sequence that can be translated into amino acids, and does not contain any stop codon. Previously, YeATSAM analyzed ORFs from the RNA-seq derived transcriptome of walnut, and revealed several genes that were not annotated by widely-used methods. Here, a similar ORF-based method is applied to the chloroplast genome from walnut (Accid:KT963008). This revealed, in addition to the ~84 protein coding genes, ~100 additional putative protein coding genes with homology to RefSeq proteins. Some of these genes have corresponding transcripts in the previously derived transcriptome from twenty different tissues, establishing these as bona fide genes. Other genes have introns, and need to be manually annotated. Importantly, this analysis revealed the mis-annotation of the rps12 gene in several organisms which have used an automated annotation flow. This gene has three exons - exon1 is ~28kbp away from exon2 and exon3 - and is assembled by trans-splicing. Automated annotation tools are more likely to select an ORF closer to exon2 to complete a possible protein, and are unlikely to properly annotate trans-spliced genes. A database of trans-spliced genes would greatly benefit annotations. Thus, the current work continues previous work establishing the proper identification of ORFs as a simple and important step in many applications, and the requirement of validation of annotations.

genomics

Unifying the two different classes of plant non-specific lipid-transfer proteins allergens classified in the WHO/IUIS allergen database through a motif with conserved sequence, structural and electrostatic features

The ubiquitously occuring non-specific lipid-transfer proteins (nsLTPs) in plants are implicated in key processes like biotic and abiotic stress, seed development and lipid transport. Additionally, they constitute a panallergen multigene family present in both food and pollen. Presently there are 49 nsLTP entries in the WHO/IUIS allergen database (http://allergen.org/). Analysis of full-length allergens identified only two major classes (nsLTP1,n=32 and nsLTP2,n=2), although nsLTPs are classified into many other groups. nsLTP1 and nsLTP2 are differentiated by their sequences, molecular weights, pattern of the conserved disulphide bonds and volume of the hydrophobic cavity. The conserved R44 is present in all full length nsLTP1 allergens (only Par j 2 from Parietaria judaica has K44), while D43 is present in all but Par j 1/2 from P. judaica (residue numbering based on PDBid:2ALGA). Although, the importance of these residues is well-established in nsLTP1, the corresponding residues in nsLTP2 remain unknown. A structural motif comprising of two cysteines with a disulphide bond (C3-C50), R44 and D43 identified a congruent motif (C3/C35/R47/D42) in a nsLTP2 protein from rice (PDBid:1L6HA), using the CLASP methodology. This also provides a quantitative method to assess the cross-reactivity potential of different proteins through congruence of an epitope and its neighbouring residues. Future work will involve obtaining the PDB structure of an nsLTP2 allergen and Par j 1/2 nsLTP1 sequences with a missing D43, determine whether nsLTP from other groups beside nsLTP1/2 are allergens, and determine nsLTP allergens from other plants commonly responsible for causing allergic reactions (chickpea, walnut, etc.) based on a genome wide identification of genes with conserved allergen features and their in vitro characterization.

immunology