Skip to content
DopeSwagYolo

AI Policy & Safety

What is AI red teaming?

AI red teaming is adversarial testing in which people or automated tools deliberately attack an AI system. The goal is to uncover weaknesses, unsafe behavior and ways it could be misused, so that developers can address them.

Also known as: red-teaming, adversarial testing, AI red team

Researched and fact-checked by AI, with no human review. 5 sources listed below. How we verify

Last updated

How it works

In security practice, a red team plays the adversary and attacks an organization's own systems to expose weaknesses. The US National Institute of Standards and Technology (NIST) glossary covers the AI version. It describes an organized effort to probe an AI system for flaws and vulnerabilities. It says this is often done in a controlled setting and with the system's developers.

The International AI Safety Report 2026 says AI red teams often search for inputs that trigger undesirable behavior. One example is jailbreaking, in which prompts are crafted to get around a model's safety restrictions. Exercises can be run by domain experts on a specific risk or left open-ended. They can be carried out by a company's own staff, outside groups or automated systems. Unlike fixed benchmarks, a red team can tailor its attacks to the system being tested.

Why it matters, and its limits

According to the report, red teams can build custom inputs to look for worst-case behavior, openings for misuse and unexpected failures. It cautions that finding no problems does not show risk is low. Flaws are frequently missed, more so when red teams are short of access or resources. Outcomes can shift with the team's makeup and instructions, the number of attack rounds and the tools the model can use. It adds that researchers have questioned whether results are reliable and reproducible.

A March 2026 post from NIST's Center for AI Standards and Innovation described a public competition hosted by Gray Swan. It said more than 400 participants made over 250,000 attempts to hijack AI agents running on 13 frontier models. At least one attack succeeded against every model. As of October 2026, NIST's website labels the blog as coming from the Center for Advancing Innovation and Standards for Super Intelligence.

Where things stand in 2026

Adversarial testing is a legal duty for some developers. Article 55 of the EU AI Act covers providers of general-purpose AI models with systemic risk. They must conduct and document adversarial testing as part of model evaluation.

US agencies have described the practice in different terms. A November 2024 blog post came from two officials at the US Cybersecurity and Infrastructure Security Agency. They described AI red teaming as safety and security evaluation of AI systems by third parties. They argued that it belongs within established software testing, evaluation, verification and validation practice.

Sources

  1. artificial intelligence red-teaming - Glossary | CSRC, NIST Computer Security Resource Center
  2. International AI Safety Report 2026 (full text, arXiv:2602.21012), International AI Safety Report, hosted by arXiv
  3. Insights into AI Agent Security from a Large-Scale Red-Teaming Competition, National Institute of Standards and Technology (NIST)
  4. Article 55: Obligations of providers of general-purpose AI models with systemic risk, European Commission, AI Act Service Desk
  5. AI Red Teaming: Applying Software TEVV for AI Evaluations, Cybersecurity and Infrastructure Security Agency (CISA)

Articles on AI Policy & Safety