Skip to content
DopeSwagYolo

AI Agents

What is computer use in AI agents?

Computer use is an AI capability in which a model operates software the way a person does. It looks at screenshots of the screen, then moves the pointer, clicks and types to carry out a task one step at a time.

Also known as: computer-using agent, CUA, computer use agent

Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify

Last updated

How it works

A computer-use system runs in a loop. Software captures a screenshot and sends it to an AI model with the user's request. The model answers with an action, such as a click, scroll or typed text. The software performs it and captures a new screenshot. The cycle repeats until the task is done or the model asks for input. OpenAI's description of its Computer-Using Agent says working from raw pixels with a virtual mouse and keyboard lets the model act without program-specific connections known as APIs.

Anthropic released the capability in public beta for its Claude 3.5 Sonnet model on October 22, 2024. Anthropic described it as experimental and at times error-prone. OpenAI followed on January 23, 2025, with a research preview of Operator, a web agent built on its Computer-Using Agent model. Google released its Gemini 2.5 Computer Use model in preview on October 7, 2025. Google said it was optimized mainly for web browsers and not yet for desktop operating systems.

Why it matters

The model acts on whatever appears on screen, which carries risk. Anthropic's documentation warns that the model may follow instructions found in webpages or images even when they conflict with the user's, a problem known as prompt injection. It lists precautions including running agents in a dedicated virtual machine or container, withholding sensitive data such as logins, and having a person confirm consequential actions. In a January 2026 notice, the U.S. National Institute of Standards and Technology said AI agent systems could be vulnerable to hijacking, backdoor attacks and other exploits.

Where things stand in 2026

One benchmark, OSWorld, has 369 tasks in real desktop and web applications. Its 2024 paper reported that people completed more than 72% of the tasks and the best model 12.24%. The maintainers evaluate results on the OSWorld-Verified leaderboard. As of October 6, 2026, it listed top scores of 90.2% for an agent framework and 86.0% for a single general-purpose model. Both were dated July 2026.

The same group later published OSWorld 2.0, 108 longer workflows that take people a median of about 1.6 hours. On October 6, 2026, its leaderboard listed a best full-completion rate of 44.33% at a 500-step limit, on a release labelled v2.1, and 31.43% on the release labelled v2026.08.08. The authors attribute failures less to basic interface control than to forgotten constraints, overlooked mid-task information and unchecked work.

Sources

Articles on AI Agents