AI Governance

Glossary

AI Red-Teaming

Adversarial testing that tries to elicit harmful, insecure or policy-violating behavior from an AI system.

Last reviewed 2026-08-22

In plain language

Red-teaming, borrowed from security practice, is now common for generative models. It can be internal, contracted, or coordinated with governments. Limits include incomplete coverage, legal constraints on testers, and the gap between lab attacks and real-world misuse. Red-teaming is a method, not a certification.

Why it matters for AI governance

Policymakers increasingly mention red-teaming as a safety control. Buyers and regulators should ask who tested what, against which threat model, and what changed afterward.

Authoritative sources