Attr-Kit: An Efficient Toolkit for No-Decode Source Attribution
Sai Sundaresan ⋅ Archit Gupta ⋅ Debabrata Mahapatra
Abstract
Source attribution, the task of identifying which input passages support an LLM-generated response, is critical for trustworthy grounded generation. A naive solution is to make additional autoregressive LLM calls to generate citations; but the added cost and latency discourage practitioners from incorporating transparency features like attribution. We study efficient $\textbf{no-decode}$ alternatives that leverage the model's internal activations instead. Toward this goal, we introduce Attr-Kit, a modular framework that formalizes source attribution, unifies several prior approaches, and proposes a new activation signal called $\textbf{Value Flow}$ (VF). To establish the theoretical potential of activation-based attribution, we formulate a subset selection problem that evaluates the maximum achievable accuracy for any type of signal. The oracle analysis reveals that all activation signals contain sufficient information for near-perfect attribution, and that VF reaches this ceiling with the fewest layers computed. Comparing across multiple datasets and model families, VF emerges as the strongest practical signal overall. VF-based attribution produces zero decode tokens, yet achieves 38\% improvement in accuracy on average compared to autoregressive citation generation from the same model, while making significantly fewer LLM calls. In fact, our no-decode methods using mid-sized models (4-8B) close the accuracy gap to frontier systems such as GPT-5.4 and Claude Opus 4.6, even surpassing them in few settings.
Successful Page Load