TLDRocket
Sign in

OpenAI o3 and o4-mini System Card

OpenAI Covered by 2 sources

OpenAI dropped a system card for o3 and o4-mini, its newest reasoning models that can now use every tool in the ChatGPT box at once.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI has published the system card for o3 and o4-mini, and the headline isn't really the reasoning gains—it's the plumbing underneath them. These two models are the first from OpenAI built to reason and use tools in the same breath. Web browsing, Python execution, image and file analysis, image generation, the canvas editor, automations, file search, and memory are all available to the model simultaneously, not bolted on as separate modes you have to switch between.

That matters because most of the friction in earlier assistants came from the handoffs. A model would think, then stop, then call a tool, then wait, then resume with whatever the tool spat back. o3 and o4-mini are designed to fold that back-and-forth into the reasoning process itself, so the model can decide mid-thought that it needs to run a script or pull up a document rather than guessing or bluffing its way through. OpenAI frames this as a step toward models that behave less like chatbots answering questions and more like agents completing multi-step jobs.

o4-mini is the interesting one here, honestly. It's positioned as the smaller, cheaper sibling, but with the same tool access as o3, which suggests OpenAI wants tool-using reasoning to be the default rather than a premium feature reserved for the flagship model. That's a meaningful shift in how these systems get deployed—cost-sensitive applications don't have to sacrifice the ability to browse or run code just because they're using the lighter model.

System cards exist for a reason, and this one presumably covers the usual ground: capability evaluations, safety testing, red-teaming results, and the guardrails OpenAI put around giving a model this much operational reach. Handing a reasoning model direct access to Python execution and live web browsing is not a small decision, and the fact that it's shipping across two models at once—one large, one small—signals OpenAI is committing to this tool-integrated approach as the new baseline going forward, not a one-off experiment.

My take — AI-written commentary, not fact-checked reporting

I'll believe the 'agent' framing when I see one of these models complete a genuinely messy multi-day task without a human quietly fixing its homework in between. Giving o4-mini the same tool access as o3 is the smart move here—cheap models that can actually browse and code are more useful to more people than another marginal reasoning benchmark win, and OpenAI clearly knows it.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.