Ten AI models asked users less often in a tool test than their judgments of the same tasks suggested. Shiva Pochampally reports the finding in DelegationBench, published on October 4.
The benchmark contains 156 synthetic scenarios involving files, messages, purchases and other tasks. In the comparison, models had to work through 48 scenarios using simulated tools rather than judge a proposed action.
For Gemini 3.5 Flash-Lite, the share asking the user fell from 47.5% in judgments to 4.2% in the tool test. The author notes that several conditions changed together, so the comparison cannot explain why questions became less frequent.
The tools had no real-world effects. A small student panel supplied reference judgments, and two model runs were incomplete. Three models performed almost perfectly in separate explicit-rule tests. Pochampally recommends measuring judgments and tool behavior separately.