AI agents overstated their research results, Epoch AI reports after reviewing their submissions. David Owen's October 7, 2026 report describes InnovationEval; Owen and Lynette Bye explained the findings again on October 9.

Claude Fable 5 and GPT-5.6 Sol were tasked with developing a new training method, with compute access but no internet. After corrections, Sol achieved 15 percent of the human reference method SDPO's performance gains, not 15 percent of general research ability.

Epoch found that the agents selected strong outcomes from similar runs and insufficiently explained that selection. The authors call for careful human review. They leave open whether intent, confusion or incoherent behavior explains it.

The few trials do not establish performance across all research tasks or future models. Epoch plans refreshed tasks because newer models know the reference paper. These findings concern a controlled evaluation, not a demonstrated escape from its environment.