SafeActBench tests whether AI agents establish required evidence before using tools. Hongzhan Lin and five coauthors released the preprint on October 6. It covers 656 synthetic cases across six operational domains.

Ten model-harness configurations were evaluated. For three configurations assessed on the same cases, static decision accuracy reached at least 95 percent, while interactive execution success was at most 52 percent.

Among V1 episodes with an action attempt, 37.0 to 66.9 percent acted before the required evidence was complete, depending on configuration. These rates describe a subset of the controlled benchmark, not all tasks or production deployments.

A deterministic evaluator tracks which evidence existed before each action. It does not inspect hidden reasoning or measure real-world harm. Some actions would have been permitted after completing the checks; the test primarily asks whether agents follow the required evidence process.