TLDRocket
Sign in

Datadog uses Codex for system-level code review

OpenAI

Datadog is now using OpenAI's Codex to review code across entire systems, not just single files. It's a sign AI code review is moving from nitpicking syntax to actually understanding how a whole codebase fits together.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Datadog, the company that built its name monitoring other people's infrastructure, just handed some of its own code review over to OpenAI's Codex. The pairing makes sense once you think about it: a company obsessed with visibility into complex systems teaming up with a coding agent built to reason across more than a single pull request.

Most AI code review tools today work like a very fast, very literal intern. They flag typos, catch obvious bugs, maybe suggest a cleaner loop. What Datadog is doing with Codex is different in scope. Instead of reviewing a diff in isolation, the tool is being used to understand how a change ripples through a larger system — dependencies, downstream services, the kind of stuff that usually only shows up when a senior engineer with years of tribal knowledge looks at the change.

That's a meaningfully harder problem than autocomplete-style suggestions. System-level review means the model has to hold a mental map of how services talk to each other, where the fragile parts of the architecture live, and what a seemingly small change might break three layers down. Datadog's own platform, built to trace exactly those kinds of cross-service failures, gives it an unusually good testing ground for whether Codex can actually do this at scale rather than just in a demo.

For OpenAI, this is another case study in its push to position Codex as infrastructure for real engineering organizations, not just a coding assistant for solo developers. Landing a company like Datadog, whose entire business depends on catching subtle system failures before customers do, is a useful proof point. It suggests Codex is being trusted with judgment calls, not just syntax fixes.

Whether this becomes standard practice across the industry depends on whether Codex's system-level suggestions hold up under real production pressure, not curated blog-post examples. Code review tools have a long history of looking impressive in announcements and then quietly getting turned off six months later when they generate too much noise. Datadog has more incentive than most to make sure that doesn't happen here, since its reputation is built on trustworthy signal, not false alarms.

My take — AI-written commentary, not fact-checked reporting

I'll believe 'system-level understanding' when I see Codex catch something a senior Datadog engineer would've missed, not just something a linter already would have. Coding agents are genuinely useful, but every vendor partnership announcement like this doubles as a stress test for OpenAI's credibility with enterprise buyers, and Datadog's brand is basically built on not tolerating noisy false positives. If this quietly disappears from their workflow in a year, that tells us more than the launch post ever will.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.