Stability AI co-founder Emad Mostaque proposed checks on September 14 designed to show which conditions keep AI within its boundaries. His essay “Manners Maketh the AI” argues for looking beyond ordinary safety-test scores.
He would vary computational budget, unfamiliar tasks and perceived oversight separately. Competence and unauthorized attempts would receive separate measures, avoiding confusion between an incapable system and a safe one.
Mostaque calls for independent evaluation and access to training records when investigators examine causes. Behavior alone cannot establish his explanation. Whether his proposed tests offer better protection remains to be studied.
In his own reply to the relayed Selsam statement, he linked the essay as a different view of agent swarms. This is a September proposal, not a newly deployed safeguard.