What is an AI hallucination?
An AI hallucination is output from a generative AI system that is stated confidently but is false, made up or unsupported by the source material. The U.S. standards agency NIST calls the phenomenon confabulation.
Also known as: confabulation, fabrication, LLM hallucination
Researched and fact-checked by AI, with no human review. 6 sources listed below. How we verify
Last updated
How it happens
The U.S. National Institute of Standards and Technology (NIST) uses the term confabulation. Its July 2024 generative AI risk profile describes a system producing wrong or false content and presenting it with confidence. That includes output that strays from the prompt or contradicts what the system said earlier. NIST treats "hallucination" and "fabrication" as colloquial names and notes that some commenters say they attribute human traits to software.
NIST calls confabulation a natural result of how generative models are designed. It says a large language model predicts the next token from statistical patterns in its training data, which can produce accurate or inaccurate text.
Researchers from OpenAI and Georgia Tech argued in a September 2025 preprint that evaluation practices keep the problem alive. Most benchmarks, they wrote, grade answers as right or wrong and give no credit for "I don't know," so a model that guesses when unsure outscores one that abstains. They propose rescoring mainstream benchmarks so expressing uncertainty is no longer penalized.
Why it matters
The risk, NIST says, arises when people believe and act on false content because it is delivered confidently. Its example is a made-up summary of patient records that leads a doctor to a wrong diagnosis. Models can also invent citations or reasoning that appear to justify an answer.
In law, researchers at Stanford's RegLab and Institute for Human-Centered AI reported in January 2024 hallucination rates of 69% to 88% on specific legal queries. The preprint study tested GPT 3.5, Llama 2 and PaLM 2. Their article recounts a Manhattan lawyer filing a brief written mostly by ChatGPT in May 2023.
A 2024 paper in Nature by University of Oxford researchers proposed a detection method. It samples several answers to the same question, groups them by meaning and measures how much they disagree. The authors said it does not address cases where a model is confidently wrong for systematic reasons.
Where things stand in 2026
Stanford's 2026 AI Index presents two benchmarks whose scales, it says, are not directly comparable. Vectara's HHEM leaderboard tests how often models introduce false information when summarizing news articles. Its top 15 models had hallucination rates between 1.8% and 5.4%, according to the report's responsible AI chapter. AA-Omniscience is a 6,000-question knowledge benchmark from Artificial Analysis that does not penalize declining to answer. Its rates across 26 models ranged from 22% to 94%.
Sources
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, National Institute of Standards and Technology (NIST AI 600-1)
- Why Language Models Hallucinate, arXiv preprint (Kalai, Nachum, Vempala and Zhang; OpenAI and Georgia Tech)
- Detecting hallucinations in large language models using semantic entropy, Nature
- Hallucinating Law: Legal Mistakes with Large Language Models are Pervasive, Stanford Institute for Human-Centered AI (HAI)
- Responsible AI | The 2026 AI Index Report, Stanford Institute for Human-Centered AI (HAI)
- AI Index Report 2026, Chapter 3: Responsible AI, Stanford Institute for Human-Centered AI (HAI)