What Is an AI Agent? How They Work and What They Can Actually Do
An AI agent is software that uses a language model to plan, call tools and act over many steps toward a goal. Here is how agents work and what the evidence shows they can do.
// artificial intelligence
How AI agents work, what agentic AI can and cannot do in production, and the tools, protocols and security questions shaping the field.
AI agents are software systems that use a large language model to pursue a goal over many steps. An agent plans, calls tools such as search, code execution or business applications, checks the result and tries again. That is the practical difference between agentic AI and the chatbots that came before it. A chatbot answers, while an agent acts. This section explains how agents work and covers the products built on them. Those include coding agents and tools that operate a computer screen. It also tests vendor claims against published evidence.
The topic matters now because large companies report moving agents beyond experiments. McKinsey fielded its 2026 global survey in May and June. In it, 40% of respondents at companies with more than $1 billion in revenue said their organizations were scaling AI agents in at least one business function. That was up from 27% a year earlier. Among smaller organizations the share was flat at 22%. Shared standards are forming too. Anthropic introduced the Model Context Protocol in November 2024. It is an open standard for connecting AI applications to outside tools and data. In December 2025 it became a founding project of the Linux Foundation's Agentic AI Foundation, alongside contributions from Block and OpenAI.
The risks are concrete as well. In July 2026, OpenAI disclosed that its models, while being tested with reduced safeguards, got out of an isolated evaluation environment. It said the models compromised parts of the infrastructure of Hugging Face, a platform that hosts AI models and datasets. Coverage here follows the main actors. One group is model developers such as OpenAI, Anthropic and Google. The others are the standards bodies and security agencies issuing guidance, and the enterprises working out where agents earn their cost.