bioRxiv Science⌕ Search

Biology subjects

Callahan, F. M.

Publications and source records attributed to Callahan, F. M..

3 recordsLinked to original sources

STEM-LM: Spatio-Temporal Ecological Modeling via Masked Language Model for Joint Species Distribution

Joint species distribution models (JSDMs) are central to biodiversity forecasting and conservation decision-making. As ecological datasets grow in size, dimensionality, and spatio-temporal resolution, there is a need for flexible yet scalable JSDMs tailored to large-scale species observation data. Recent advances in masked language modeling for text and genomics suggest a natural alternative: by treating each species presence or absence as a token, and a sites species assemblage together with its spatio-temporal and ecological covariates as a sentence, we can learn joint co-occurrence structure by reconstructing masked species from their neighboring sites. We propose STEM-LM1, a Transformer-based JSDM that frames joint species distribution modeling as masked language modeling. By varying the masking rate during training, a single trained model supports both purely spatiotemporal/ecological prediction and conditioning on arbitrary subsets of observed species for joint co-occurrence inference at a given site. On a North American butterfly and a global plant distribution dataset, STEM-LM performs better or on par with other statistical and deep-learning based methods in terms of discriminative ranking, while producing substantially better rank-calibrated occurrence probabilities. Utilizing partial species observations at the same site greatly enhances prediction performance.

ecology↗

Co-occurrence networks can preserve emergent properties of ecological communities

Interaction networks, in which nodes represent species and edges represent direct interactions between species, have a long and impactful history in community ecology. However, co-occurrence networks, where edges represent statistical relationships among species presences or abundances, are often easier to construct from lab and field data. It is clear that co-occurrence edges often do not represent direct interactions, but frameworks for the interpretation of co-occurrence networks have not kept pace with their generation. It is therefore unclear when and how these networks can be used to gain insight into community dynamics. Here, we use a Generalized Lotka-Volterra-based model to explore the contexts in which emergent properties of species interaction networks are identifiable in their resulting co-occurrence networks. We find that, in spite of many differences in direct edges, key features of the true interaction network, such as unipartite modularity, high-degree nodes (hubs), and bipartite modularity and nestedness, can be preserved in co-occurrence networks. In contrast, node degree distributions are not preserved even in the most idealized scenarios. We propose that networks derived from large co-occurrence datasets could therefore be used in future empirical work to test existing hypotheses of how emergent network structures drive ecological community dynamics.

ecology↗

Challenges in detecting ecological interactions using sedimentary ancient DNA data

With increasing availability of ancient and modern environmental DNA technology, whole-community species occurrence and abundance data over time and space is becoming more available. Sedimentary ancient DNA data can be used to infer associations between species, which can generate hypotheses about biotic interactions, a key part of ecosystem function and biodiversity science. Here, we have developed a realistic simulation to evaluate five common methods from different fields for this type of inference. We find that across all methods tested, false discovery rates of inter-species associations are high under simulation conditions where the assumptions of the methods are violated in a variety of ecologically realistic ways. Additionally, we find that for more realistic simulation scenarios, with sample sizes that are currently realistic for this type of data models are typically unable to detect interactions better than random assignment of associations. Different methods perform differentially well depending on the number of taxa in the dataset. Some methods (SPIEC-EASI, SparCC) assume that there are large numbers of taxa in the dataset, and we find that SPIEC-EASI is highly sensitive to this assumption while SparCC is not. Additionally, we find that for many methods, default calibration can result in high false discovery rates. We find that for small numbers of species, no method consistently outperforms logistic and linear regression, indicating a need for further testing and methods development.

ecology↗