TLDRocket
Sign in

Watcher Monitors and Secures AI Code Execution

Watcher by Apollo Research

Apollo Research launched Watcher, a security layer that watches AI coding agents like Claude Code and Codex in real time and blocks risky actions. Because a single bad agent action can already cost companies millions, and nobody's watching until it's too late.

Based on reporting by Watcher by Apollo Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Apollo Research, the same lab that's been picking apart frontier AI systems for safety research, has built a product out of that work: Watcher, a monitoring layer that sits on top of coding agents and stops them from doing something destructive before it happens. The pitch is blunt. AI breaches average $6 million, according to the company, and most teams running Claude Code or Codex right now have no record of what those agents actually did and no way to stop a bad action mid-flight. You find out after the damage is done. Watcher wants to change that timing problem.

The setup is refreshingly simple to try — one curl command spins it up in cloud mode, pointed at Apollo's own infrastructure in Western Europe. For companies that don't want their agent traffic touching someone else's servers, self-hosting is an option too, and Apollo says that's historically taken about an hour to get running on your own environment. Once it's live, Watcher blocks dangerous commands as they're attempted — the demo example is an agent trying to run terraform apply -auto-approve against production — and it logs every session so a security team can review what happened afterward, including flagging rule updates like preventing a production bypass.

What's more interesting than the product itself is the argument Apollo makes against just using the safety features Anthropic and OpenAI already ship. Claude Code has auto-mode, Codex has auto-review, and Apollo says plainly these are good and it's glad the labs are building safety in by default — the company even ran a red-teaming campaign with the Claude Code auto-mode team. But those tools are built for one developer working with one agent. Watcher is pitched as the enterprise layer on top: one set of rules defined once and enforced across every coding agent in the building, centralized visibility across every developer's sessions, and admin-locked policy instead of settings a developer can just toggle off.

The cross-model monitoring angle is the part worth paying attention to. Apollo's observation is that a model watching itself has blind spots baked in from the same training — and they claim GPT models tend to run overly suspicious while Claude models can be too trusting of their own reasoning. Watcher's answer is to mix vendors in the monitoring pipeline itself, using something like Gemini Flash for a fast first pass and Claude Sonnet further down the chain, so the reviewer doesn't inherit the same biases as the agent it's reviewing. Cursor support is still being built; Claude Code and Codex work today.

My take — AI-written commentary, not fact-checked reporting

Using a different model family to police an agent's own work is the smartest idea buried in this launch, because letting a lab's model grade its own homework was always going to have blind spots. Whether $6 million average breach figure holds up under scrutiny is a separate question worth watching closely — vendors selling the fix rarely undersell the fire. Still, the shift from developer-toggled safety to admin-enforced, cross-agent policy is the obvious next step as companies stop treating coding agents like toys and start treating them like staff with root access.

Read more about this at: Watcher by Apollo Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.