Goodfire reports detection of roughly 93 percent of harmful test sessions with its new cyber monitors. At that operating point, 5.5 percent of benign sessions would be interrupted at least once. Its October 8 study describes monitors for Kimi K3 and GLM 5.3.

The method reads internal model activations and sends only suspicious steps to another language model for review. It aims to make checks before risky tool calls more practical.

Goodfire tested roughly 2,400 chat and agent sessions containing over 60,000 turns. Model judgments supplied reference labels; parts of the dataset used simulated users.

A static FAR.AI evaluation reported by Goodfire found no universal bypass but 18 successful individual attacks. The study therefore documents progress and remaining monitoring limits, not a guarantee against arbitrary attacks.