OpenAI's GPT-Red AI Outperforms Human Hackers, Slashing GPT-5 Vulnerabilities by Over 70%
Summary
OpenAI's GPT-Red AI is revolutionizing cybersecurity by outperforming human hackers in red-teaming tests, slashing GPT-5 vulnerabilities by over 70% — but the powerful exploit-finding system will remain locked away from the public.
Key Points
- OpenAI has developed GPT-Red, an AI-powered red-teaming system that automates cybersecurity testing by simulating prompt injection attacks and other exploits against its own language models to identify and patch vulnerabilities before release.
- GPT-Red is trained using a self-play loop where it competes against other models in simulated real-world scenarios, and has already discovered a novel attack called a 'fake chain of thought,' which tricks models into acting on falsified reasoning steps.
- Testing shows GPT-Red outperforms human red-teamers in finding effective attacks, with over 90% of its strongest exploits working against GPT-5 but fewer than 23% succeeding against the newly released GPT-5.6, though OpenAI confirms it will not be releasing GPT-Red to the public.