Long AI-agent tests strain human oversight, the UK AI Security Institute warns. On October 7, the institute released Transect to help reviewers examine those runs.

A final score does not reveal discarded strategies or delegated work. Transect connects activities and automated judgments to the relevant transcript passages.

An accompanying preprint examines a research evaluation that generated about 13 million tokens. Its authors find little sustained hypothesis generation, while acknowledging possible automated classification errors.

Agreement between AI judges does not establish correctness. The descriptive case study confirms neither general research capability nor the causes of behavior. Expert review of the underlying evidence remains necessary.