Extrinsic Hallucinations in LLMs
Lilian Weng published an article titled "Extrinsic Hallucinations in LLMs," narrowing the problem of hallucination to cases where model output is fabricated and not grounded by either the provided context or world knowledge.
The post distinguishes two types: in-context hallucination, where output should be consistent with source content in context, and extrinsic hallucination, where output should be grounded by the pre-training dataset. Weng notes that given the size of pre-training data, retrieving and identifying conflicts per generation is too expensive. Treating the pre-training corpus as a proxy for world knowledge, she writes that models need to be factual and to acknowledge not knowing an answer when applicable.
On causes, Weng points to pre-training data issues, including out-of-date, missing, or incorrect information crawled from the public internet that models may memorize by maximizing log-likelihood. She also cites Gekhman et al. 2024, which studied whether fine-tuning on new knowledge encourages hallucinations. That work found LLMs learn new-knowledge examples slower than examples consistent with pre-existing knowledge, and that once learned, such examples increase the model's tendency to hallucinate.
For detection, the post covers retrieval-augmented evaluation, citing FactualityPrompt (Lee et al. 2022), which uses Wikipedia documents from FEVER and reports hallucination named-entity errors and entailment ratios; FActScore (Min et al. 2023), which decomposes long-form generation into atomic facts validated against a knowledge base; SAFE (Wei et al. 2024), which uses a language model agent to issue Google Search queries and reason about support, reporting 72% agreement with humans and a 76% win rate over humans when they disagree; and FacTool (Chern et al. 2023), a fact-checking workflow for tasks including QA, code generation, math, and literature review. It also covers sampling-based detection via SelfCheckGPT (Manakul et al. 2023).
On calibration of unknown knowledge, the post cites TruthfulQA (Lin et al. 2021), comprising 817 questions across 38 topics, and SelfAware (Yin et al. 2023), with 1,032 unanswerable and 2,337 answerable questions. It also references calibration work by Kadavath et al. 2022 and Lin et al. 2022, and indirect query approaches by Agrawal et al. 2023 for hallucinated references.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Hallucination in large language models usually refers to the model generating unfaithful, fabricated, inconsistent, or nonsensical content. As a term, hallucination has been somewhat generalized to cases when the model makes mistakes. Here, I would like to narrow down the problem of hallucination to cases where the model output is fabricated and not grounded by either the provided context or world knowledge. There are two types of hallucination: In-context hallucination: The model output should be consistent with t