OpenAI researcher Noam Brown warned about the limits of safety evaluations in his September 17, 2026 Dwarkesh interview. Strong test results, he argued, do not guarantee that models reliably follow human intentions in deployment.
Brown also reports declining visibility into models' reasoning and describes monitoring as buying time for safety research. These are his accounts of internal observations, not independently audited measurements.
The OpenAI study from December 2025 linked in the interview provides useful context. It found no substantial loss of monitorability from reinforcement learning in the runs it examined. Its authors caution that their tests have limited coverage and leave uncertainty about larger training runs.
That earlier study does not independently establish Brown's newer observation. We record his warning as a Risk Signal on the interview's original date. It establishes neither another attack nor an agreed development pause.