OpenAI has disclosed an internal, attack-focused AI model called GPT-Red that automatically generates and executes prompt-injection attacks against its own systems, a technique the company says it used to sharply improve the security of its latest model. The tool was detailed in a company blog post published around July 15, 2026, and framed as a way to scale up red-teaming beyond what human testers can achieve.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.