bioRxiv ScienceSearch

Biology subjects

Pal Choudhury, P.

Publications and source records attributed to Pal Choudhury, P..

7 recordsLinked to original sources

Comparative validation of breast cancer risk prediction models and projections for future risk stratification

BackgroundWell-validated risk models are critical for risk stratified breast cancer prevention. We used the Individualized Coherent Absolute Risk Estimation (iCARE) tool for comparative model validation of five-year risk of invasive breast cancer in a prospective cohort, and to make projections for population risk stratification.\n\nMethodsPerformance of two recently developed models, iCARE-BPC3 and iCARE-Lit, were compared with two established models (BCRAT, IBIS) based on classical risk factors in a UK-based cohort of 64,874 women (863 cases) aged 35-74 years. Risk projections in US White non-Hispanic women aged 50-70 years were made to assess potential improvements in risk stratification by adding mammographic breast density (MD) and polygenic risk score (PRS).\n\nResultsThe best calibrated models were iCARE-Lit (expected to observed number of cases (E/O)=0.98 (95% confidence interval [CI]=0.87 to 1.11)) for women younger than 50 years; and iCARE-BPC3 (E/O=1.00 (0.93 to 1.09)) for women 50 years or older. Risk projections using iCARE-BPC3 indicated classical risk factors can identify ~500,000 women at moderate to high risk (>3% five-year risk). Additional information on MD and a PRS based on 172 variants is expected to increase this to ~3.6 million, and among them, ~155,000 invasive breast cancer cases are expected within five years.\n\nConclusionsiCARE models based on classical risk factors perform similarly or better than BCRAT or IBIS. Addition of MD and PRS can lead to substantial improvements in risk stratification. Independent prospective validation of integrated models is needed prior to clinical evaluation risk stratified breast cancer screening and prevention.

epidemiology

Chemical Characterization of Interacting Genes in Few Subnetworks of Alzheimer’s Disease

A number of genes have been identified as a key player in Alzheimers disease (AD). Topological analysis of co-expression network reveals that key genes are mostly central or hub genes. The association between a hub gene and its neighbour genes can be derived easily using relative abundance of their expression levels. However, it is still unexplored fact that whether any hub and its neighbour genes within a sub-network exhibits any kind of proximity with respect to their chemical properties of the DNA sequences or not, that code for a sequence of amino acids.\n\nIn this work, we try to make a quantitative investigation of the underlying biological facts in DNA sequential and primary protein level in mathematical paradigm. It may gives a holistic view of the interrelationships existing between hub genes and neighbour genes in few selective AD subnetworks. We define a mapping model from physicochemical properties of DNA sequence to chemical characterization of amino acid sequences. We use distribution of chemical groups present in a sequence after decoding into corresponding amino acids to investigate the fact that whether any hub genes are associated closely with its neighbour genes chemically in the subnetworks. Interestingly, our preliminary results confirm the fact the dependent genes that are coexpressed with its hub gene are also having proximity with respect to their amino acid chemical group distributions.\n\nCCS Concepts*Applied computing [->] Computational genomics;

synthetic biology

Effective Variations of related Physicochemical Properties of Nucleotides Leading to Amino Acids for Characterizing Genes and Proteins

The aim of this paper is to make quantitative analysis of the properties which is really being carried from DNA sequence and finally landing up to the properties of a protein structure through its primary protein sequence. Thus, the paper has a theory which is applicable for any arbitrary DNA sequence whether it is of various species or mutated data or a bunch of genes responsible for a function to be occurred. Irrespective to genes of any families, species, wild type or mutated, our paper here gives a standard model which defines a mapping between physicochemical properties of any arbitrary DNA sequence and physicochemical properties of its amino acid sequence. Experiments have been carried out with PPCA protein family and its four homologs PPC(B E) which establishes that DNA sequence keeps its signature even after its translation into the corresponding amino acid sequence.

systems biology

The variations of human miRNAs and Ising like base pairing models

miRNAs are small about 22-base pair long, RNA molecules are of extreme biological importance. Like other longer RNA molecules, messages in miRNAs are encoded by the permutations of only four nucleotide bases represented by A, U, C and G. However, just like words in any language, not all combination of these alphabets make a meaningful word. In fact, we find that the distributions of nucleotides bases in human miRNAs show significant deviation from randomness. First, a miRNA sequence containing four bases are mapped into a binary string with three kinds of classifications according to their chemical properties. Then, we propose a simple nearest neighbor model (Ising model) to understand the statistical variations in human miRNAs.

bioinformatics

Characterizing glutamate receptor genes of Rat having vertebrate nervous system and Arabidopsis thaliana a plant having equivalent nervous system based on chemical properties of amino acids and investigating evolutionary relationships between them.

iGluR gene family of a vertebrate, Rat and AtGLR gene family of a plant, Arabidopsis thaliana [4] perform some common functionalities in neuro-transmission, which have been compared quantitatively. Our attempt is based on the chemical properties of amino acids [6, 7, 8] comprising the primary protein sequences of the aforesaid genes. 19 AtGLR genes of length varying from 808 amino acid (aa) to 1039 aa and 16 iGluR genes length varying from 902aa to 1482 aa have been taken as data sets. Thus, we detected the commonalities (conserved elements) during the long evolution of plants and animals from a common ancestor [4]. Eight different conserved regions have been found based on individual amino acids. Two different conserved regions are also found, which are based on chemical groups of amino acids. We have tried too to find different possible patterns which are common throughout the data set taken. 9 such patterns have been found with size varying from 2 to 5 amino acids at different regions in each primary protein sequences. Phylogenetic trees of AtGLR and iGluR families have also been constructed. This approach is likely to shed light on the long course of evolution.

systems biology

Testing Equality of Curves After Covariate Adjustment

SO_SCPLOWUMMARYC_SCPLOWThis paper is concerned with providing simple methodological approaches for global and local tests of difference between the mean of treatment and control groups when the measured outcome is a function. The added complexity is that for every subject we have repeated samples for the same curve and additional covariates of interest. We propose a permutation based approach to test for a global difference between the averages of two functional processes after covariate adjustment. The within group averages are estimated by modeling the relationship of the functional outcome on the covariate using functional regression methods and then averaging with respect to the covariate distribution in each group. The test statistic is the L2 area under the squared difference curve. We also test for the localized differences between the two average curves using a nonparametric bootstrap of subjects to obtain the 95% pointwise and joint confidence intervals for the estimated covariate-adjusted difference curve. Extensive simulation studies illustrate that the proposed tests preserve the type one error and are highly sensitive to detecting departures from the null assumption. We illustrate our method by studying the differences in time varying oxygen consumption between the frail Interleukin 10tm1Cgn (IL10tm) mice and the wildtype mice after adjusting for body composition measures.

epidemiology

Distribution of Purines and Pyrimidines over miRNAs of Human, Gorilla and Chimpanzee

Meaningful words in English need vowels to break up the sounds that consonants make. The Nature has encoded her messages in RNA molecules using only four alphabets A, U, C and G in which the nine member double-ring bases (adenine (A) and Guanine (G)) are purines, while the six member single-ring bases (cytosine (C) and uracil (U)) are pyrimidines. Four bases A, U, C and G of RNA sequences are divided into three kinds of classifications according to their chemical properties. One of the three classifications, the purine-pyrimidine class is important. In understanding the distribution (organization) of purines and pyrimidines over some of the non-coding regions of RNA, all miRNAs from three species of Family Hominidae (namely human, gorilla and chimpanzee) are considered. The distribution of purines and pyrimidines over miRNA shows deviation from randomness. Based on the quantitative metrics (fractal dimension, Hurst exponent, Hamming distance, distance pattern of purine-pyrimidine, purine-pyrimidine frequency distribution and Shannon entropy) five different clusters have been made. It is identified that there exists only one miRNA in human hsa-miR-6124 which is purely made of purine bases only.\n\nAMS Subject Classification: 92B05 & 92B15

bioinformatics