Skip to content
DopeSwagYolo

AI Agents

What is an AI coding agent?

An AI coding agent is a software tool built on a large language model that carries out programming tasks with some autonomy. It reads a codebase, edits files, runs commands and tests, and revises its work until the task is done.

Also known as: coding agent, agentic coding tool, software engineering agent

Researched and fact-checked by AI, with no human review. 9 sources listed below. How we verify

Last updated

How it works

A coding agent is given a goal in plain language, such as fixing a bug. It works in a loop. It reads the relevant files, edits code, runs commands such as tests, checks the results and tries again. A 2024 paper on SWE-agent described an interface that lets a language model write and change code files, move around a repository and run tests and other programs. SWE-agent is a research system.

Commercial products are described in similar terms. Anthropic's documentation describes Claude Code as an agentic coding tool that reads a codebase, edits files and runs commands. GitHub's documentation says its Copilot cloud agent researches a repository, plans changes and carries them out in the background in a temporary cloud environment. It says each session is capped at 59 minutes. In May 2025, OpenAI launched a research preview of its Codex agent, which runs in a sandboxed virtual computer in the cloud, TechCrunch reported.

How progress is measured

The SWE-bench benchmark, introduced in October 2023, contains 2,294 problems drawn from real GitHub issues in 12 Python repositories. The best model in the original paper, Claude 2, resolved 1.96% of them. Stanford's 2026 AI Index reports that on SWE-bench Verified, a version using human-validated issues, the top model resolved about 76.8% as of February 2026. Terminal-Bench 2.0 tests agents in command-line environments. The index says accuracy on it rose from 20% in February 2025 to 77.3% in early 2026.

Where things stand in 2026

Stack Overflow's 2026 developer survey reported that coding assistants and agents were respondents' most common use of AI, at 66%, and that 73% use them daily. Asked when it is acceptable to trust AI, 48% chose cases where they can easily validate the answers.

The effect on productivity is less settled. The research group METR found in a 2025 randomized trial that 16 experienced open-source developers took 19% longer on tasks when allowed to use AI tools. In February 2026 METR said a follow-up study gave an unreliable signal, mainly because more developers declined to take part rather than work without AI. It said developers were likely sped up more than in early 2025 but called its data very weak evidence of how much.

GitHub's documentation cautions that its chat and agent features can produce incorrect code, including security vulnerabilities. It says output should be reviewed and tested before production use.

Sources

Articles on AI Agents