bioRxiv ScienceSearch

Biology subjects

Obolenska, M.

Publications and source records attributed to Obolenska, M..

2 recordsLinked to original sources

Creation of gene expression database on preeclampsia-affected human placenta

Publication of gene expression raw data in open access at online resources like NCBI or ArrayExpress made it possible to use these data for cross-experiment integrative analysis and make new insights into biological phenomena. However, most popular of the present online resources are meant to be archives rather than ready for immediate access and interpretation databases. Data uploaded by independent contributors is not standardized and sometimes incomplete and needs further processing before it is ready for the analysis. Hence, the need for a specialized database appears.\n\nGiven in this article is the description of the database that was created after processing a collection of 33 relevant datasets on pre-eclampsia-affected human placenta. Data processing includes the choice of relevant experiments from ArrayExpress database, the experiment sample attributes standardization according to MeSH term dictionary and Experimental Factor Ontology and the completion of missing data using information from the corresponding articles and authors.\n\nA database of more than 1000 samples contains sufficient sample-wise metadata for them to be arranged into relevant case-control groups. Metadata includes information on biological specimen, donors diagnosis, gestational age, mode of delivery etc. The average size of these groups will be higher than it is in separate experiments. This will reduce experiment bias and enhance statistical accuracy of the subsequent analysis such as search for differentially expressed genes or inferring gene networks. The article concludes with the guidelines for the microarray experiment metadata uploading for future contributors.

bioinformatics

Designing the database for microarray experiments metadata

Advancements in both computer science and biotechnology opened way to unprecedented amount and variety of gene expression studies raw data in the open access. It is sometimes worth to rearrange and unite data from several similar gene expression studies into new case-control groups to test new hypothesis using available data. Unfortunately, most popular gene expression databases such as GEO and ArrayExpress were not designed to allow such cross-study procedures. In order to locate comparable samples in different studies numerous steps are required including gathering additional sample metadata and its standardization. Specialized databases are developed by investigators in their own fields of interest to reuse the processed data and create different case-control groups and test multiple hypothesis.\n\nHere we present detailed description of the specialized database creation along with its use case which is 32 gene expression cDNA microarray datasets on human placenta under conditions of pre-eclampsia containing expression data on more than 1000 biological samples. Samples contain sufficient metadata for them to be merged into relevant cross-experiment case-control groups for further integrative analysis.

bioinformatics