OpenAI has built an adversarial AI system called GPT-Red that automatically generates prompt injection attacks against its own models, using the results to make production systems markedly more resistant to manipulation. In one striking demonstration, the system tricked an office AI vending-machine agent into repricing a high-value item to its floor of \$0.50, setting up a \$100-plus product to sell for \$0.50, and cancelling another customer's order.
Continue reading
The rest of this article is for AI News Blitz readers. Choose an option below to keep reading.
Already purchased? Sign in✓ Signed in — this article isn’t included in your current plan.