TierCache: A Multi-Granularity Semantic Caching Framework for Structured Query Generation
Abstract
Semantic caching is a promising solution to reduce LLM serving costs by reusing cached responses. However, existing methods are predominantly designed for single-level, verbatim response reuse. This limits their effectiveness in structured query generation, where queries often share the same output structure. To address this, we propose TierCache, a multi-granularity semantic caching framework integrating two-stage retrieval and two-tier reuse. Tier 1 provides the fastest path via semantic retrieval and verbatim reuse, while Tier 2 enables the reuse of structurally similar queries through structure-aware retrieval and cost-efficient slot filling. We instantiate our approach on Text-to-SQL and construct SQLStream, a benchmark designed to evaluate structured query generation caching under realistic cache-reuse opportunities. Across all workloads, TierCache consistently dominates the quality-cost Pareto frontier. Under a strict quality-preserving constraint (>95% retention), it improves cache hit rates by 18.90%–47.91% and reduces end-to-end latency by 2.01×–2.89× over the state-of-the-art semantic caching method.