Do Foundation Model Image Embeddings Capture Radiological Semantics in CT Lung Nodules?
Abstract
Foundation models have enabled pretrained image embeddings to serve as transferable latent representations for a wide range of downstream tasks. However, whether these representations preserve clinically meaningful semantic concepts in thoracic computed tomography (CT) remains unclear. In this work, we investigate the semantic quality of embeddings extracted from two pretrained vision encoders, MedSigLIP and RadImageNet, using the LIDC-IDRI lung nodule dataset. Specifically, we examine whether these embeddings capture radiologist-defined semantic attributes, including malignancy and spiculation. We evaluate the learned representations through PCA and UMAP visualization, feature-wise Spearman correlation analysis, and linear probing with multiple supervised classifiers. Across both embedding models, low-dimensional projections exhibit substantial overlap among semantic categories, feature-level correlations remain weak, and linear probes achieve near-random performance despite class balancing and multiple classifiers. These preliminary findings suggest that off-the-shelf pretrained image embeddings may not readily encode fine-grained CT lung nodule semantics in a linearly accessible manner. Our study provides an initial empirical assessment of foundation model image embeddings for capturing clinically meaningful radiological semantics and highlights the need for domain-adaptive representation learning for thoracic CT analysis.