- Self-play reinforcement learning trains one model as attacker and defender at the same time, so prompt injection and similar flaws are found automatically.
- Attack success rate: 84% for GPT-Red against 13% for the human red team.
- GPT-5.6 Sol, built on these findings, is described as six times more resistant to direct prompt injection than earlier models.
As more companies build AI agents into real workflows, automated adversarial testing becomes a security criterion you can apply to your own deployments.