TLDRocket
Sign in

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

TechCrunch Russell Brandom Covered by 3 sources

OpenAI showed off Jalapeño, its chip for AI inference, and early tests beat today’s top gear on speed and power. It’s still years from broad use, so the real race may be how fast everyone else moves.

Based on reporting by TechCrunch, Russell Brandom — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI gave Jalapeño its first public benchmark lap at Hot Chips on Tuesday, and the early numbers are not shy. On Semianalysis’s InferenceX test, the chip posted more tokens per user and more throughput per kilowatt than the state-of-the-art inference processors available right now.

Richard Ho, OpenAI’s head of hardware, called the result a “very, very significant performance advance” and said the chip can do more AI work for each unit of power while also answering faster. That’s the key appeal here: high volume, low latency, and less wasted energy when serving a lot of users at once.

The comparison is against an Nvidia Blackwell system, which makes the claim punchy but also time-sensitive. Jalapeño is not rolling out tomorrow. Ho said it should arrive at the end of 2026 in “very small volumes,” with wider deployment coming in 2027.

First announced last October, Jalapeño was built with Broadcom and with OpenAI’s own models helping during development. OpenAI says it wants the chip to be part of a multigenerational platform, with AI products, models, chips and memory all developed together instead of as separate pieces.

That full-stack setup is really the story. OpenAI says it used it to reduce friction in the parts of inference that slow things down, especially prefill and communication. The company says it kept model state, including the KV cache, local when needed and tuned compute, memory and networking around each phase. The promise is simple: fewer delays, less movement, more useful work per watt.

The chip race is turning into a power-efficiency race, which is far less glamorous than a giant model demo but probably more important. If OpenAI can keep hardware, models and memory under one roof, that is not just engineering discipline; it is a very expensive way of saying it trusts no one else to build the plumbing properly.

My take — AI-written commentary, not fact-checked reporting

OpenAI is doing the obvious frontier-company thing: turning inference into a private hardware project and calling it architecture. That usually means two things at once — better performance for the people inside the moat, and fewer reasons to believe the rest of the market gets a fair shake. The cloud era loved abstraction; this era seems to love vertical integration with a straight face.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.