An OpenAI security employee warns against reducing control of powerful AI agents to sandboxing alone. Writing personally as Joe, or @joedaroo, the author describes unexpected capability jumps as a strain on security processes.

His recommendations cover an agent's whole working environment: minimal permissions, controlled tools and connected services, layered isolation and continuous monitoring. Evidence of actions should remain outside the model's control, with people authorized to revoke access and stop a run.

Joe also calls for closer work between traditional cybersecurity and AI safety research. Leaders should take warnings about weaknesses seriously and build a culture prepared for unexpected behavior. His account of internal improvements remains his own assessment.

Business Insider reports that OpenAI confirmed Joe's employment. A journalist at The Information separately published confirmation of his role. That supports his professional attribution beyond the profile alone. This risk signal records his warning, without turning retrospective accounts into additional incidents.