The Embedder's Dilemma: LLMs Are Better, but at What Cost?
Abstract
Can a single prompted LLM, with no task-specific training, outscore every purpose-built text embedding model? We show that it can: in a controlled, cost-aware comparison of three frontier LLMs against 26 embedding models (118M–14B parameters) across 38 tasks spanning classification, STS, clustering, pair classification, and retrieval, the best LLM tops all embeddings in aggregate. But this edge is not statistically significant (p = 0.14), hinges on a single temporal-reasoning task, and costs 704× more (USD 177 vs. USD 0.25) at 3–41× lower throughput. Beneath the aggregate, the two paradigms are complementary: LLMs lead on retrieval requiring cross-document reasoning; embeddings win on classification and are competitive across all other categories. A "thinking token tax" underlies the cost asymmetry: chain-of-thought tokens consume 59–78% of LLM spend, yet reducing thinking by 63–97% improves five of six retrieval tasks. LLMs and embedding models are complements, not substitutes: default to embeddings for similarity and classification, reserve LLMs for reasoning-intensive retrieval. Code, data, and a leaderboard are available at an anonymous link.