A new study reports data leakage even when an additional auditor checks AI agent actions. Shouju Wang and Haopeng Zhang published AgentPrivArena on arXiv on October 5.
The researchers tested five models on 389 synthetic tasks in self-hosted services. They report that contextual runtime auditing reduced average disclosure of protected information from 46.8% to 17.8% of tasks.
The auditor records information after reads and checks it before outbound writes. It nevertheless missed some protected facts the agents had already read. Repeated abstraction could also make individual answers unhelpful.
Language models judged the results, and human verification was unfinished. All records were synthetic and the services isolated from the open internet. The authors also warn that their audit allows actions when it fails, so it should not be treated as a dependable defense against adversarial attacks.