Improved Robustness against Indirect Prompt Injection Can Be Built into LLMs
Abstract
LLM agents with tool-calling capabilities are vulnerable to indirect prompt injection, where adversarial instructions in retrieved data hijack the agent’s control flow. Existing defenses operate at the pipeline level and remain costly, attack-specific, or harmful to utility. We introduce ODILE, a representation-level defense that trains a lightweight LoRA adapter to disrupt harmful internal states before they produce dangerous tool calls. The core challenge is that agentic harm is context-dependent: the same tool call is benign or malicious depending on whether it was triggered by the user or an injection. We address this with paired execution traces that share identical structure, differing only in the model’s response to an injection, and apply loss over an early completion window where harmful and benign trajectories diverge. On AgentDojo (Llama-3.3-70B), ODILE reduces attack success rate from 56.3% to 1.4% while preserving 88% of benign capability, at standard inference cost with no external dependencies. We further evaluate generalization to additional attack families, robustness under train–eval mismatch between tool-calling trace formats, and transfer to Qwen 2.5 7B; supplementary tables collect training-setting aggregates and relative-reduction details.