BREAKING
OpenAI Unveils GPT-Red
Self-Play Red-Teaming Loop
1
Attacker probes
↓
2
Find injection
↓
3
Adversarial train
↓
4
Harden model
84% vs 13% attack success
GPT-Red
84
Humans
13
0
%
direct injection fails
0
%
indirect resistance
0
x
fewer failures
Vendy Case Study
Attack
Andon Labs
●
Price manipulation
●
Unauthorized orders
●
Order cancellation
Response
●
Vulnerabilities disclosed
●
New safeguards in testing
Claims Await Verification
AI NEWS BLITZ
OpenAI has revealed GPT-Red, an internal AI that attacks its own models to harden them.