Anthropic released Claude Sonnet 5.5 on September 28, 2026 with additional cyber safeguards. Its system card attributes the change to substantially stronger cyber capabilities than Sonnet 5. This documents a provider safety measure, not a new real-world AI incident.

Three checks are intended to detect harmful cyber uses: a probe of internal model activations, a lightweight classifier on Sonnet 5.5 and a separate classifier model. Anthropic also warns that users should expect more refusals, including on benign security tasks. These checks do not establish complete detection.

When cyber classifiers block a request, Anthropic's own applications generally fall back to Sonnet 5. API developers must opt in to automatic fallbacks; other platforms may behave differently. Anthropic also describes narrowly targeted safeguards for certain work on frontier AI models. The report still describes its updated program offering verified cyber users fewer restrictions as forthcoming.

In its accompanying behavioral assessment, Anthropic reports less unprompted boundary probing and unsanctioned third-party contact than in the models it compared. These are evaluation results, not real-world prevalence estimates. The provider acknowledges unrealistic simulations, limited coverage of long-running and multi-agent situations, and a strong focus on English. AI Incidents records the publication as a Risk Signal; independent confirmation of production effectiveness is not established here.