TLDRocket
Sign in

Tenstorrent’s Galaxy Blackhole AI servers escape the event horizon

The Register

Tenstorrent's Galaxy Blackhole AI servers are now shipping, packing 32 chips into a $110k box. That's a third to a fifth the price of Nvidia's DGX rigs, and it scales up to 4,000+ chips.

Based on reporting by The Register — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Tenstorrent just flipped the switch on general availability for its Galaxy Blackhole platform, and the pitch is refreshingly blunt: cheaper hardware that still gets the job done. Each 6U chassis crams in 32 of the Blackhole accelerators the company showed off last fall, wired together with 100 Tbps of aggregate Ethernet bandwidth. Add it all up and you get 1 TB of GDDR6, 16 TB/s of memory bandwidth, and 23 petaFLOPS of dense FP8 compute for $110,000 — a system that undercuts Nvidia's eight-way DGX boxes by three to five times, even though those boxes still win on raw speed and capacity.

What makes this more than a single-box story is the mesh. Tenstorrent's networking approach mirrors what Google does with TPU pods and Amazon does with Trainium2: stack nodes together, tune the ratio of tensor and pipeline parallelism, and scale out to bigger models or more users. A base Galaxy Supercluster with four nodes runs $440,000, and the architecture theoretically stretches to 144 nodes and more than 4,000 chips — enough to make you wonder how many companies will actually need that much, but the option is there.

The more interesting shift is on the software side. When The Register first tested Blackhole hardware last year, model support was thin and whatever did run hadn't been tuned properly, so performance scaling looked rough. Senior fellow Jasmina Vasiljevic says that's changed, with real optimization work happening even as the company quietly dialed back the chip's peak performance specs a few months back. For DeepSeek V3 specifically, Tenstorrent claims a four-node Supercluster can chew through a 100,000-token prompt — about 166 pages — in under four seconds, and push out 300 tokens per second per user, with 350 targeted through further software tweaks.

Those numbers come with an asterisk, though. Tenstorrent hasn't published the batch sizes behind these figures, and batch size is the detail that actually tells you whether a system scales in production or just looks good running one query at a time. The company says it scales cleanly from batch eight to batch 64 depending on whether you're optimizing for throughput or interactivity, which is a reasonable answer but not quite the same as showing the work.

Beyond text generation, Tenstorrent is also pitching Galaxy Blackhole for video, claiming faster-than-real-time 720p generation on a four-node cluster. Vasiljevic says Moonshot AI's Kimi K2 and other frontier models are being ported over, aided by a new Python-based kernel interface, and the company is boasting that 90 percent of Hugging Face models already run on its silicon — a claim worth testing rather than taking at face value. Early customers include Cirrascale, Equinix, and Japan's ai&, giving buyers a rental option before committing six figures. Tenstorrent says there's more to come at its TT-Deploy event on May 1.

My take — AI-written commentary, not fact-checked reporting

Cheap AI hardware that actually runs real models is exactly what this industry needs more of, and I like that Tenstorrent is chasing price instead of just chasing benchmarks nobody can replicate. But withholding batch size on the throughput numbers is a classic trick to make single-user performance look like production-ready scaling, and I won't stop side-eyeing that until they publish it. If the 90-percent Hugging Face compatibility claim actually holds up under independent testing, this becomes a genuinely interesting Nvidia alternative rather than just a nice discount.

Read more about this at: The Register

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.