TLDRocket
Sign in

Inference is giving AI chip startups a second chance to make their mark

The Register Covered by 3 sources

As AI adoption shifts from model training to inference, specialized chip startups are gaining opportunities to compete against Nvidia by handling specific workload phases, with companies like Groq, Cerebras, and SambaNova winning design partnerships with major cloud providers. Nvidia's acquisition of Groq for $20 billion and subsequent partnerships from AWS and Intel demonstrate that inference workloads are being split between different processors—GPUs handling compute-heavy prefill operations while specialized chips accelerate the bandwidth-constrained decode phase. This disaggregated approach, along with emerging technologies like Lumai's optical accelerators targeting exaOPS performance by 2029, is reshaping how inference infrastructure is built, though some startups like Tenstorrent are pursuing unified platforms as alternatives to the multi-chip model.

Why it matters

In a disaggregated AI world, Nvidia can be both a friend and an enemy

Also covered by

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.