VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean
Abstract
Large language models have recently achieved striking results in interactive theorem proving, particularly in Lean. While most benchmarks for LLM-based proof automation are drawn from mathematics and involve the Mathlib ecosystem, many proofs in software verification and programming-languages semantics are developed inside sizable codebases that define their own core abstractions and involve substantial project-specific libraries. In this setting, progress often depends on locating and using a small set of relevant local definitions and lemmas amid a large amount of surrounding context. This paper introduces VeriSoftBench, a benchmark of 500 Lean 4 proof obligations drawn from open-source formal-methods developments and that are packaged to preserve project context and cross-file dependencies. We evaluate frontier LLMs and specialized provers on VeriSoftBench and find that success rates drop sharply as contextual dependency increases. This highlights the difficulty that current systems face when proofs are grounded in project-specific libraries rather than a single shared ecosystem like Mathlib.