TLDRocket
Sign in

Building self-improving tax agents with Codex

OpenAI

OpenAI teamed with Thrive and Crete to build a tax-filing agent using Codex that learns from its own mistakes. It's a real-world test of AI handling messy, high-stakes paperwork most people dread.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Tax software has never been glamorous, but it's exactly the kind of grinding, rule-heavy work that AI agents are supposed to be good at. OpenAI's latest case study, built with fintech partners Thrive and Crete, shows what happens when you point Codex at that problem directly instead of just automating the easy parts.

The setup is straightforward on paper: an agent reads tax documents, drafts filings, and checks its own output against expected results. What's different here is the feedback loop. Instead of a static script that does the same thing every time, the system uses Codex to write and rewrite pieces of its own logic when it spots recurring errors or edge cases it handled badly. Engineers at Thrive and Crete fed it real filing scenarios, let it fail in controlled ways, and then had it patch its own code to close the gaps.

That self-correcting loop is the actual story, more than the tax use case itself. Most AI deployments in finance still rely on humans catching mistakes after the fact. This one tries to shrink that human review burden by having the agent notice its own blind spots and generate fixes before a person ever has to intervene. OpenAI frames it as evidence that Codex can be trusted with iterative, semi-autonomous coding tasks in a domain where a wrong number has real consequences, not just a rendering bug on a webpage.

Accuracy gains and faster turnaround are the headline results, though the more interesting detail is how much of the improvement came from the agent's own revisions rather than manual tuning by Thrive's or Crete's engineers. It suggests a workflow where developers spend less time writing every rule by hand and more time deciding what counts as a good outcome, then letting the agent iterate toward it.

Whether this generalizes beyond tax filing is the open question. Tax rules are messy but at least they're written down somewhere. Plenty of other back-office work is far less structured, and that's where a self-improving agent would really earn its keep.

My take — AI-written commentary, not fact-checked reporting

Self-improving agents that patch their own logic sound impressive right up until you ask who's liable when the patch is wrong on a tax return, and OpenAI's writeup conveniently skims past that part. I'd rather see this kind of autonomy pointed at low-stakes drudgery first, with humans in the loop on anything touching real money, before anyone starts calling it a solved problem.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.