What is a large language model (LLM)?
A large language model (LLM) is an AI system trained on very large amounts of text to predict what comes next in a sequence. Models of this kind are used to answer questions, translate text and help write computer code.
Also known as: LLM, LLMs, language model
Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify
Last updated
How it works
A 2023 explainer from Georgetown University's Center for Security and Emerging Technology (CSET) says the "large" in the name refers to the number of parameters. These are the internal values that govern how a model turns input into output. It said cutting-edge models might have thousands or even millions of times as many parameters as those trained ten years earlier. GPT-3, described in a 2020 paper, had 175 billion.
The transformer, an architecture built entirely on attention mechanisms, was introduced in the 2017 paper "Attention Is All You Need". In a March 2025 report, the U.S. National Institute of Standards and Technology (NIST) describes generative pre-trained transformers as the current predominant LLM architecture. These are first trained on large sets of unlabeled text. Such a model can then be fine-tuned for specific tasks, the report says. NIST's July 2024 generative AI profile describes LLMs as predicting the next token or word in a sentence.
Why it matters
The March 2025 NIST report says LLMs are increasingly embedded in software and internet infrastructure. It lists online search, coding help for developers and chatbots used by millions of people daily among their uses. It says LLMs are also being linked to corporate documents and set up as agents that can take actions such as browsing the web.
The July 2024 profile notes that predicting text statistically can also produce output that contains factual errors or contradicts itself.
Where things stand in 2026
The 2026 AI Index from Stanford's Institute for Human-Centered AI reports that industry produced more than 90% of notable AI models in 2025. It says the United States produced 59 and China 35. Parameter counts, it says, have held at about 1 trillion for three years. But it says such figures are no longer disclosed for several of the most resource-intensive systems, including ones from OpenAI, Anthropic and Google.
A European Commission Q&A cites the EU AI Act's description of large generative AI models as a typical example of a general-purpose AI model. Obligations for providers of such models have applied since August 2, 2025. They include technical documentation, a copyright policy and a published summary of training content. The Commission's AI Act page says its AI Office and member-state authorities have been responsible for enforcing the Act since August 2, 2026. It says the AI Office can fine providers of general-purpose models.
Sources
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, National Institute of Standards and Technology (NIST AI 100-2e2025)
- Attention Is All You Need, arXiv (Vaswani et al.)
- What Are Generative AI, Large Language Models, and Foundation Models?, Center for Security and Emerging Technology, Georgetown University
- Language Models are Few-Shot Learners, arXiv (Brown et al.)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, National Institute of Standards and Technology (NIST AI 600-1)
- Research and Development | The 2026 AI Index Report, Stanford Institute for Human-Centered AI (HAI)
- General-Purpose AI Models in the AI Act – Questions & Answers, European Commission
- AI Act | Shaping Europe's digital future, European Commission