Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
Abstract
Retrieval-Augmented Generation (RAG) systems, which combine large language models (LLMs) with external knowledge sources, have shown promise in improving the factual accuracy of LLM-generated content. However, these systems remain susceptible to intrinsic hallucinations, where the model generates unfaithful or fabricated information that is not supported by the retrieved evidence. We propose a novel framework to assess model robustness against this phenomenon using natural, semantically equivalent variations of a user query. We apply our framework, which enforces strict semantic equivalence constraints and an intrinsic hallucination objective, to a range of attack techniques across white-box, gray-box, and black-box adversarial settings. Evaluating these attacks on 5 open-source and 3 closed-source generator models across 3 datasets, we demonstrate that even state-of-the-art models are highly susceptible to meaning-preserving perturbations. These attacks significantly degrade contextual faithfulness (by up to 52.3\% for gpt-5-nano). Our findings highlight critical robustness gaps in current RAG systems and underscore the urgent need for more resilient architectures and training methodologies.