bioRxiv Science⌕ Search

Biology subjects

Cruz, I. N.

Publications and source records attributed to Cruz, I. N..

3 recordsLinked to original sources

Secundary Structure of Physicochemical Clustered Proteins

Diverse methods have been proposed for protein secondary structure prediction. However, such task still presents a challenge in bioinformatics. In this article various of these methods are implemented and analysed. First, a baseline using Support Vector Machine. Then a convolutional neural network (CNN), a Long Short-Term Memory (LSTM) and a strategy of Ensembling both of these methods. Lastly, a novel technique Secundary Structure of Physicochemical Clustered Proteins (SSPCP) is proposed, which combines multiple CNNs trained accordingly to a protein feature clustering and combined using a neural network. The rationale behind SSPCP is that amino acids from proteins which have similar physicochemical characteristics should have the same secondary structure prediction for similar amino acids, but amino acids from differing proteins might have different structures. All of these methods use as features PSSM matrices extracted from PSIBLAST. For performance evaluation, 25pdb dataset was split into training and validation and the same subsets were used on all these methods achieving the Q3 score of CNN: 70.11%, LSTM: 69.25%, Ensemble: 70.71%, SSPCP: 70.91%. The experimental results show that the features extracted from clustering of physicochemical properties of proteins seem to improve the accuracy of highly specific CNN models for accurate protein secondary structure prediction.

bioinformatics↗

Computational Assessment on Catalytic Activity of PET Hydrolase

BackgroundPET hydrolase from Ideonella sakaiensis might provide a response for PET accumulation in the environment. In this project some previously studied mutations were implemented and their performance was evaluated via computational methods with tools such as Modeller, HADDOCK, PyMOL and Gromacs. One possible mutation that could lead to improved catalytic activity was proposed. ResultsPET hydrolase DM S209F W130H and I179 provide interesting binding results with studied ligands, however a solution that combines both mutations does not seem viable, since the binding cleft becomes occluded. Following the same rationale, the triple mutant S209F W130H I179Q is proposed but instead leaves space in the binding cleft for ligand to enter and might bond with the oxygen at the ester group. The experiments conducted with triple mutant S209F W130H I179Q failed to beat HADDOCK score for DM, however its experimental results could still increase PET degradation. Results from surface charge may indicate an increase in stability and binding affinity for the protein. ConclusionsAmong models implemented, DM S209F W130H seems the best model studied regarding BHET or PET binding. Despite Protein Engineering is a complex process, computational tools might provide a way of studying binding sites of hypothetical proteins. Supplementary informationSupplementary data is available in annexes.

bioinformatics↗

Gene Expression and Physiological traits in Mice

BackgroundGene expression regulates several complex traits observed. In this study, datasets comprising of transcriptome information and clinical traits regarding fat composition and vitals were analyzed via several statistical methods in order to find relations between genes and clinical outcomes. ResultsBiological big data is diverse and numerous, which makes for a complex case study and difficulties to stablish a metric. Histological data with semi-quantitative scores proved unreliable to correlate with other vitals, such as cholesterol composition, which complicates prediction of clinical outcomes. A composition of vitals, turned out to be a better variable for regression and factors for gene analysis. Several genes were found to be statistically significant after statistical analysis by ANOVA regarding the progressive categories of the preferred clinical variable. ConclusionsANOVA is proposed as a method for genetic information retrieval in order to extract biological meaning from RNA seq or microarray data, accounting for multiple classes of target variables. It Provides a reliable statistical method to associate genes or clusters of genes with particular traits. Supplementary informationSupplementary data are available in annexes.

genomics↗