OpenAI’s GPT-5.6 Tests Show Prompt-Injection Gains and Agent Risks
OpenAI has added prompt-injection results to the GPT-5.6 system card, reporting a low failure rate for attacks delivered directly through chat but higher rates in tests involving AI agents and external content. In the Aug. 3 update, GPT-5.6 Sol failed on about 0.05% of direct attacks generated by GPT-Red, the company’s automated red-teaming model. Malicious instructions hidden in content processed by agents were more successful. Average attack success rates reached 3.77% for Sol, 3.32% for Terra and 2.94% for Luna in OpenAI’s indirect tests. Such instructions can arrive through emails, webpages, uploaded files, code repositories or tool responses. Direct attacks fall as agent tests remain harder The updated GPT-5.6 system card describes direct prompt injection as a user’s attempt to override higher-priority instructions. An indirect attack embeds malicious instructions in material supplied to the model through a tool. The percentages measure successful attack attempts across OpenAI’s evaluation environments, not the probability of a production breach. OpenAI trained GPT-Red through self-play, rewarding it for finding prompts that caused defender models to violate higher-priority instructions. The company …









