Sophron Research co-founder Paul de Font-Reaulx warned on September 14 that AI systems recognizing a test setting could undermine safety evaluations. The philosopher and cognitive scientist made his own statement in response to Dan Selsam’s earlier warning.

De Font-Reaulx develops evaluations for AI models. He argues that awareness of testing may increase alongside model capabilities. A reassuring evaluation could therefore make model properties look better than the evidence warrants, rather than reliably indicate their behavior.

He currently regards the problem as a challenge, not proof that useful evaluation is impossible. For more capable agents, he says, that distinction remains uncertain. He calls for high methodological standards and research into what those standards should be.

The post presents neither a new experiment nor evidence that all AI tests fail. It identifies a methodological uncertainty in the debate over human control. This is a September statement secured during follow-up research, with its original date retained.