Loading…

OpenAI's internal GPT-Red model finds successful attacks in 84 percent of test scenarios through self-play training. Human red teamers manage just 13 percent. The results feed directly into hardening models like GPT-5.6 Sol. The article OpenAI is now using AI to attack its own…
To respect copyright, we link to the source rather than republishing the full text. Read the complete article on The Decoder.