Evaluating the Transferability of Pathology Foundation Models Across Cancer-related H&E Neurodegeneration-related Immunohistochemical Classification Tasks
Foundation models (FMs) have rapidly become dominant in artificial intelligence and are increasingly being adopted in computational pathology. Numerous pathology-specific FMs have been developed and evaluated for a variety of downstream tasks, most commonly using frozen image embeddings with linear probes. While a pathology FM, at least implicitly suggests broad reusability across tasks, the transferability of these models to neurodegenerative disease-related tasks remains largely unexplored. In this study, we evaluated fourteen frozen feature extractors, including general-vision and pathology FMs, across four pathology image classification datasets spanning two specialized neurodegenerative disease immuno-histochemistry (IHC) tasks (Tau neurofibrillary tangle (NFT) and amyloid-{beta} plaque classification) and two cancer-related hematoxylin and eosin (H&E) tasks (the breast tissue BACH dataset and the multi-tissue TIL dataset). The cancer-related H&E datasets represent tasks more closely aligned with the predominant pretraining domain of many current pathology FMs, whereas the neurodegenerative IHC datasets represent a more specialized domain with limited representation in existing FM pretraining cohorts. Frozen embedding linear probes were compared against a conventional supervised ResNet-50 convolutional neural network (CNN) trained directly on image tiles using identical train, validation, and test splits. Across both neurodegenerative IHC datasets, the supervised CNN substantially outperformed all frozen FM linear probes. In contrast, pathology FMs achieved performance comparable to, and in some cases exceeding, the supervised CNN across the two cancer-related H&E datasets, with UNI2-h achieving the highest performance on BACH and several pathology FMs performing on par with the CNN on the larger TIL dataset. Furthermore, a supervised CNN trained using only 1% of the Tau NFT training data (1,938 tiles) still exceeded the performance of the best frozen FM linear probe trained on the complete dataset. Together, these results suggest the transferability of frozen pathology FMs may depend strongly on how well the downstream task is represented by their pretraining domain. These findings demonstrate frozen pathology FMs transfer effectively to the cancer-related H&E tasks evaluated here but may be less effective for specialized neurodegenerative IHC tasks less well represented in current FM pretraining cohorts. Expanding pathology FM pretraining datasets to include a broader range of disease domains and staining modalities may therefore improve transferability to specialized pathology applications. Our findings also highlight the continued importance of conventional supervised learning and motivate future work investigating nonlinear probes and end-to-end FM fine-tuning.