BREAKING
OpenAI Unveils GPT-Red
Self-Play Red-Teaming Loop
1Attacker probes
2Find injection
3Adversarial train
4Harden model
84% vs 13% attack success
GPT-Red84
Humans13
0%
direct injection fails
0%
indirect resistance
0x
fewer failures
Vendy Case Study
AttackAndon Labs
Price manipulation
Unauthorized orders
Order cancellation
Response
Vulnerabilities disclosed
New safeguards in testing
Claims Await Verification
AI NEWS BLITZ
OpenAI has revealed GPT-Red, an internal AI that attacks its own models to harden them.