ainewsblitz.com

Breaking

OpenAI Unveils GPT-Red, an Internal AI System That Attacks Its Own Models to Harden Them Against Prompt Injection

  • Security
  • Foundation Models
  • AI Agents

OpenAI has built an adversarial AI system called GPT-Red that automatically generates prompt injection attacks against its own models, using the results to make production systems markedly more resistant to manipulation. In one striking demonstration, the system tricked an office AI vending-machine agent into repricing a high-value item to its floor of \$0.50, setting up a \$100-plus product to sell for \$0.50, and cancelling another customer's order.

Continue reading

The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.

$20
Read this article
$29/month
Unlimited — all 5,756 articles, the full archive, and comprehension quizzes
Save 72%
$98/year
≈ $8.17/month
Unlimited, billed once a year