AI Incidents
Back to published register
Risk signalIndustry responseNot an incident

OpenAI launches a public framework for model misbehavior

OpenAI published a formal process through which employees can flag concerning model behavior and request public disclosure. At launch, the company released six specific training and evaluation cases and said alignment and monitoring are not sufficiently solved to keep scaling responsibly at maximum speed for much longer.

Published
Sep 16, 2026
Signal type
Industry response
Confidence
100%
Organization
OpenAI
Sources
3
Last reviewed
Sep 19, 2026

Assessment

What this signal changes

The framework defines reportable behavior, internal escalation routes, three investigation tracks, and minimum contents for public reports. It covers unauthorized action, model-to-model coordination, evasion of oversight, and failed safety assumptions. Publishing six detailed reports at the same time provides inspectable examples and makes the measure more substantive than a statement of intent alone.

Evidence boundary

What the evidence does not show

The process is internal, voluntary, and controlled by OpenAI. It does not yet establish binding industry-wide criteria, independent decision authority, or guaranteed disclosure timelines for complex cases. OpenAI says the six reports are an initial selection, not a complete account of known cases or investigations. The signal remains separate from the incident register.

Classification: This page records a sourced warning, research finding, policy action, or industry response. It is deliberately separate from the Incident register and does not prove that an unauthorized AI action occurred.