Artificial Intelligence
OpenAI's GPT-6 Astra reduces errors, prompt attacks remain
OpenAI said GPT-6 Astra made fewer factual mistakes than GPT-5.6 Sol and resisted more attacks, but external and internal tests still found failures against indirect prompt injections and adaptive jailbreaks.
- OpenAI said GPT-6 Astra made fewer factual errors than GPT-5.6 Sol on ChatGPT conversations that users had flagged as wrong.
- OpenAI said Astra blocked 99.99 percent of direct prompt-injection attempts and credited GPT-Red for the result.
- OpenAI said Astra refused harmful jailbreak prompts in 91.5 to 98.3 percent of fixed-test cases, but its defence rate fell to about 67 percent in multi-round adaptive tests.
- Gray Swan said Astra failed at least once in 8.5 percent of indirect prompt-injection scenarios, compared with 27 percent for GPT-5.6 Sol and 4.8 percent for Claude Opus 5.
OpenAI's system card said GPT-6 Astra made fewer factual errors than GPT-5.6 Sol when it was tested against ChatGPT conversations that users had flagged as wrong. The gains were strongest at low latency settings and lower reasoning levels.
OpenAI said Astra blocked 99.99 percent of direct prompt-injection attacks. It said the result came from GPT-Red, an automated attacker used during training.
On fixed jailbreak tests covering biology, violence and cybersecurity, OpenAI said Astra refused most attempts. When attackers changed tactics over several rounds, that defence rate fell to about 67 percent, and OpenAI said the tests were run on the bare model rather than the product with its production safety layers.
Gray Swan said external testing with 1,810 curated indirect attacks from IPI Arena found that Astra was cracked at least once in 8.5 percent of scenarios when it got 15 attempts per case. Gray Swan said GPT-5.6 Sol failed in 27 percent of those scenarios and Claude Opus 5 in 4.8 percent.
Gray Swan said its combined Q1 and Q2 results were higher than earlier Q1-only figures. Anthropic had previously reported a two percent success rate on the easier Q1 test and said it had run all models with extended reasoning on, which Gray Swan said could explain part of the difference.
Named in this story
Companies
- OpenAI
- published the system card and the test results
- Gray Swan
- ran the external indirect prompt-injection tests
- Anthropic
- had previously reported earlier results and ran the broader comparison with extended reasoning on
Products and systems
- GPT-6 Astra
- the new model being evaluated
- GPT-5.6 Sol
- the predecessor model used for comparison
- GPT-Red
- the training method OpenAI said hardened Astra
- ChatGPT
- the source of the flagged conversations used in OpenAI's error testing
- IPI Arena
- the test set used for the curated indirect attacks
- Claude Opus 5
- the model that outperformed Astra in Gray Swan's comparison
How the source tells it
The piece is mostly a technical comparison, but it repeatedly frames the remaining failure modes as a security concern and highlights the gains in OpenAI's favour.
- alarm Residual attack success is presented as a worry for enterprise deployments and as a growing risk from autonomous agents.
- progress framing The article stresses lower error rates and near-perfect direct defences while comparing Astra favourably with earlier models.