bioRxiv ScienceSearch

Biology subjects

Bongsong Kim

Publications and source records attributed to Bongsong Kim.

2 recordsLinked to original sources

Numericware i: Identical in state matrix calculator

Herein we introduce software, Numericware i to compute a matrix consisting of all pairwise identical in state (IIS) coefficients from genotypic data. Since the emergence of high throughput technology for genotyping, calculating an IIS matrix between many pairs of entities has required large computer memory and lengthy processing times. Numericware i addresses these limitations with two algorithmic methods: multithreading and forward chopping. The multithreading feature allows computational routines to concurrently run on multiple CPU processors. The forward chopping addresses memory limitations by dividing the genotypic data into appropriately sized subsets. Numericware i allows researchers who need to estimate an IIS matrix for big genotypes to use typical laptop/desktop computers. For comparison with different software, we calculated genetic relationship matrices using Numericware i, SPAGeDi and TASSEL with the same small-sized data set. Numericware i measured kinship coefficients between zero and two, while the matrices from SPAGeDi and TASSEL produced different ranges of values, including negative values. The Pearson correlation coefficient between the matrices from Numericware i and TASSEL was high at 0.993, while SPAGeDi rarely showed correlation with Numericware i (0.088) and TASSEL (0.087). To compare the capacity with high dimensional data, we applied the three software to a simulated data set consisted of 500 entities by 1,000,000 SNPs. Numericware i spent 71 minutes using seven CPU cores on a laptop (DELL LATITUDE E6540), while SPAGeDi and TASSEL failed to start. Numericware i is freely available for Windows and Linux under CC-BY license at https://figshare.com/s/f100f33a8857131eb2db.

Bioinformatics

Hierarchical Association Coefficient Algorithm

Suppose that members in a universal set categorized based on observations, and that categories can be stratified based on the average of observations within each category. Two sorting extremes can be obtained from the perspective of arbitrariness of an order of observations. The first sorting extreme is an increasing order of observations on ascendingly stratified categories. The second sorting extreme is a decreasing order of observations on ascendingly stratified categories. Hierarchical association coefficient (HA-coefficient) algorithm is based on a principle that any order of observations in stratified categorization can be placed between the two sorting extremes. The algorithm produces a proportion of how much an order of observations in stratified categorization is close to the first sorting extreme, or how much an order of categorized observations is distant from the second sorting extreme. This paper introduces a theory about the HA-coefficient algorithm, and shows its applications with example data. In addition, proving a reliability of the algorithm is shown through a simulation.

Bioinformatics