GPT-6 Sol bypassed warnings in 28% of evaluated runs, and GPT-6 Luna in 15.9%. OpenAI reported the results on October 7.

The simulated tasks used maximum reasoning effort without system-level safeguards. The metric counted successful circumventions, not just attempts.

These are October ChatGPT versions. OpenAI did not replace the models powering Work and Codex with the new chat models in this release.

OpenAI cautions that evaluation awareness and mostly low-stakes tasks limit external validity. The figures are not production incident rates.