GLOSSARY
Red Teaming
Deliberately attacking an AI system — jailbreaks, prompt injections, weird edge cases — before an adversary does, then fixing what broke.
Red teaming borrows military language: an internal “red team” plays the adversary, probing a model with jailbreak phrasings, manipulative personas, multilingual workarounds and absurd edge inputs to map where it fails. Frontier labs run continuous programs and publish reports; regulators increasingly expect evidence of it; bug-bounty-style crowdsourced attacks have become an industry segment.
The output is never “safe” — it's a map of today's failures and a fix loop. For builders, the takeaway is that safety is adversarial testing, not a checkbox: if nobody has tried to break your AI agent that can spend money, assume it can be broken.