bioRxiv Science⌕ Search

Biology subjects

Cashman, M.

Publications and source records attributed to Cashman, M..

2 recordsLinked to original sources

Dissecting Complexity: The Hidden Impact of Application Parameters on Bioinformatics Research

Biology is a quest; an ongoing inquiry about the nature of life. How do the different forms of life interact? What makes up an ecosystem? How does a tiny bacterium work? To answer these questions biologists turn increasingly to sophisticated computational tools. Many of these tools are highly configurable, allowing customization in support of a wide range of uses. For example, algorithms can be tuned for precision, efficiency, type of inquiry, or for specific categories of organisms or their component subsystems. Ideally, configurability provides useful flexibility. However, the complex landscape of configurability may be fraught with pitfalls. This paper examines that landscape in bioinformatics tools. We propose a methodology, SOMATA, to facilitate systematic exploration of the vast choice of application parameters, and apply it to three different tools on a range of scientific inquires. We further argue that the tools themselves are complex ecosystems. If biologists explore these, ask questions, and experiment just as they do with their biological counterparts, they will benefit by both finding improved solutions to their problems as well as increasing repeatability and transparency. We end with a call to the community for an increase in shared responsibility and communication between tool developers and the biologists that use them in the context of complex system decomposition.

bioinformatics↗

Climatic clustering and longitudinal analysis with impacts on food, bioenergy, and pandemics

Predicted growth in world population will put unparalleled stress on the need for sustainable energy and global food production, as well as increase the likelihood of future pandemics. In this work, we identify high-resolution environmental zones in the context of a changing climate and predict longitudinal processes relevant to these challenges. We do this using exhaustive vector comparison methods that measure the climatic similarity between all locations on earth at high geospatial resolution. The results are captured as networks, in which edges between geolocations are defined if their historical climates exceed a similarity threshold. We then apply Markov clustering and our novel Correlation of Correlations method to the resulting climatic networks, which provides unprecedented agglomerative and longitudinal views of climatic relationships across the globe. The methods performed here resulted in the fastest (9.37 x 1018 operations/sec) and one of the largest (168.7 x 1021 operations) scientific computations ever performed, with more than 100 quadrillion edges considered for a single climatic network. Correlation and network analysis methods of this kind are widely applicable across computational and predictive biology domains, including systems biology, ecology, carbon cycles, biogeochemistry, and zoonosis research.

systems biology↗