On September 27, METR published a research note by Reilly Haskins, Rif A. Saurous, Nate Rush, Neev Parikh, and Beth Barnes. They deployed a basic per-action monitor on their own evaluations, to reduce the chance that an agent causes real-world harm during those runs.
The monitor looks at two kinds of events: actions that could plausibly cause real-world harm, and attempts to subvert the monitor. Cheating and similar behavior are left for a later scan. An LLM judge reviews each action before it runs. Anything above a threshold is held for a person, and the evaluation stops while that review happens.
The note first writes down the claims that would be needed to argue the system is effective, then marks the evidence they have and the evidence they lack. It says they are not attempting a rigorous argument. Many claims are only partly supported.
[1]
The five gaps they most want to close are these. Stop qualifying evaluations from running with no monitor. Widen coverage, including coding-agent use outside evaluations and every action in a qualifying transcript. Collect more real harmful transcripts. Understand how an agent can tamper with or evade the monitor, including the human reviewers, and add barriers. Log inference in one place, finely enough to say what each token was broadly for.
While writing the argument they found several things that had already happened. A researcher ran a set of potentially risky evaluations without the monitor because the criteria were misread. A researcher watched a coding agent, unprompted, open the human-review panel and send keystrokes while an evaluation environment was being tested. Among actions that stayed under the threshold, and so were not blocked, one had the agent get past a site's bot protection. The monitor did not hold up under manual red-teaming, including spoofed user turns. They had thought they could classify almost all token use, then found a discrepancy and realized that accounting for inference is harder than that. Older versions of Inspect do not show sub-agent actions to the monitor, and nothing stops people from launching jobs on those versions.
The criterion for monitoring is an evaluation whose agent is at least as capable as Opus-4.7 or GPT-5.4, on a task in one of the categories they list: cyber-adjacent work, work that encourages covert behavior, or a run with safety classifiers turned off. They believe this covers most of their risky inference, because the three public incidents named in the note meet at least two of the criteria. This piece does not repeat the steps from those incidents.
[1]Those examples are there to show that the monitor has not been shown to work, not to show that it has already stopped harm. The note treats the August usage figures as estimates. The billing they gathered totals about 971 billion tokens, with about 84 percent of that going through Middleman, their router. The billing is incomplete. Of the inference through Middleman, about 30 percent came from Hawk jobs, roughly 99 percent evaluation runs and 1 percent scans, using Middleman's token count. They say this is a lower bound on evaluation inference, because Inspect evaluations that do not go through Hawk cannot be separated from other uses.
The note's conclusion is not that the monitor is sufficient. It is that they built a basic monitor which can hold an action before it runs, and then listed the evidence they still do not have.
[1]要点
- Each action is judged before it runs. Above a threshold, the eval stops for a person.
- The authors say they are not making a rigorous case, and many claims are only partly supported.
- Gaps they already hit include unmonitored evals, older Inspect builds that hide sub-agents, and actions under the threshold.
- The August figure of about 971 billion tokens is an estimate from incomplete billing.