Biological Cartography: Building and Benchmarking Representations of Life
1The continued scaling of genetic perturbation technologies combined with high-dimensional assays such as cellular microscopy and RNA-sequencing has enabled genome-scale reverse-genetics experiments that go beyond single-endpoint measurements of growth or lethality. Datasets emerging from these experiments can be combined to construct perturbative "maps of biology", in which readouts from various manipulations (e.g., CRISPR-Cas9 knockout, CRISPRi knockdown, compound treatment) are placed in unified, relatable embedding spaces allowing for the generation of genome-scale sets of pairwise comparisons. These maps of biology capture known biological relationships and uncover new associations which can be used for downstream discovery tasks. Construction of these maps involves many technical choices in both experimental and computational protocols, motivating the design of benchmark procedures to evaluate map quality in a systematic, unbiased manner. Here, we (1) establish a standardized terminology for the steps involved in perturbative map building, (2) introduce key classes of benchmarks to assess the quality of such maps, (3) construct maps from four genome-scale datasets employing different cell types, perturbation technologies, and data readout modalities, (4) generate benchmark metrics for the constructed maps and investigate the reasons for performance variations, and (5) demonstrate utility of these maps to discover new biology by suggesting roles for two largely uncharacterized genes. 2 Author SummaryWith the proliferation of genetic perturbation, laboratory robotics, computer vision and sequencing technologies, a growing number of researchers are producing datasets that capture digital readouts of cellular responses to genetic perturbations at the full-genome-scale. Since each of these efforts utilizes different cellular models, experimental approaches, terminology, code bases, analysis methods and quality metrics, it is exceptionally difficult to reason through the pros and cons of possible design choices or even discuss the primary considerations when embarking on such an endeavor. These datasets can be powerful discovery tools to look at known biological relationships and uncover new associations in an unbiased manner, but only when paired with a computational pipeline to assemble the data into a digestible format. Moreover, there is great promise in looking across these data to highlight commonalities and differences that may be attributed to experimental or analytical approaches or the biological context. Therefore, a unified framework is necessary to align this nascent field and speed progress in assessing technologies and methods. In this work we define a unified framework for building and benchmarking these perturbative maps, benchmark four different datasets assembled into 18 different maps, explore the impact of different design decisions and demonstrate how these maps can be used to elucidate gene functions. The framework we propose highlights the necessary steps for building any such map - embedding, filtering, aligning, aggregating and relating the data across perturbations. For benchmarking, we propose two main types of metrics and give examples which highlight the impact of different processing pipelines. Finally, we explore these maps to demonstrate their utility for confirming known biological relationships and nominating annotations for genes with unknown function. We expect that this work will positively impact the nascent field of perturbative map building by enabling easier comparisons within and between technologies and methods through a shared language. Additionally, the associated code base is openly available and flexible enough to be easily extended with new methods, so we hope that it will become a resource for future researchers working on developing both laboratory and computational methodology. While there are too many confounding variables to make recommendations on the strengths of different technologies and cellular models at this time, highlighting that fact may prompt studies designed with the goal of directly comparing methods while holding other confounding variables fixed. Moreover, as the number of perturbative maps grows, the field will naturally consider the advantages of combining maps across modalities and the framework provided here can also help guide the evaluation of those efforts.