TLDRocket
Sign in

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

Latent Space

OpenAI says its Jalapeño chip beats NVIDIA GB200/GB300 on real inference work. That’s the bigger story: the AI chip race is now about watts and latency, not just raw speed.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

The headline from Hot Chips was OpenAI’s Jalapeño. Not a side project, not a lab curiosity, but a custom inference chip the company says is already good enough to start showing up inside its own infrastructure by year-end. And the numbers it published are the sort that make Nvidia people sit up straighter: 1.5–1.9x more work per watt at peak throughput, 1.7–3.6x lower end-to-end latency, and 2.1–4.1x higher performance on highly interactive workloads against GB200/GB300 systems on real model runs.

The interesting part is what OpenAI seems to be optimizing for now. This isn’t just about brute-force throughput. The whole point is the ugly tradeoff between serving lots of traffic and keeping responses snappy. Jalapeño is rated at 700W, but OpenAI says the tested runs stayed at or below 550W, which is a nice reminder that chip specs and operational reality are not the same thing.

There’s also a broader inference-stack message hiding in the benchmarks. Several reactions noted that Jalapeño reportedly held up even without some of the usual tricks, like aggressive prefill/decode disaggregation or speculative decoding in certain setups, while still beating systems that were using them. That’s the kind of result that suggests OpenAI is trying to redesign the serving stack, not just bolt on a faster part.

The other detail that matters is how much of this seems to be getting co-written by models. OpenAI said GPT-Astra and Codex helped write and optimize low-level kernels, and that three open-weight models reached high performance on Jalapeño in about two months. For selected attention and MoE blocks, those kernels were reportedly 1.5–1.8x faster than existing human-written code. That is a very modern kind of infra team.

Taken together, Jalapeño points to a shift that goes beyond one company’s chip. Frontier labs are no longer behaving like permanent tenants of Nvidia’s inference economics. Packaging and foundry capacity are still real constraints, but the old assumption — that the smartest model shop just rents the whole stack — looks a little less solid today.

My take — AI-written commentary, not fact-checked reporting

This is the part where the industry pretends custom silicon was always inevitable, which is adorable. The more useful reading is simpler: if OpenAI can use its own models to tune its own kernels and squeeze better latency per watt, the moat is starting to look very physical. Nvidia still owns the default, but the rent is getting challenged by people who’d rather own the plumbing and complain about it later.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.