Cybersecurity definition
What is AI Red Teaming?
By Cyber Tool Stack Editorial TeamUpdated
Definition
Structured adversarial testing of an AI model or application to uncover exploitable behavior, policy failures, data leakage, unsafe outputs, and weaknesses in surrounding controls.
Where AI Red Teaming fits in the security landscape
These beginner lessons use this term while explaining the surrounding security control.
Related cybersecurity terms
AI JailbreakA prompt or interaction pattern designed to make an AI system bypass its safety or policy restrictions and produce behavior that the system was intended to refuse.Prompt InjectionAn attack in which untrusted instructions supplied directly by a user or indirectly through retrieved content attempt to override an AI application's intended instructions, policy, or tool-use boundaries.AI SecurityThe discipline of identifying, testing, governing, and reducing security risks in machine-learning models, generative-AI applications, and autonomous agents across their development and operating lifecycle.