TLDRocket
Sign in

Hardware & Infrastructure

171 summarised stories in Hardware & Infrastructure, each linking back to the original source. Browse all topics →

Monday, 6 July 2026

📈 Data to start your week

Exponential View 2 weeks ago

Nvidia's Grace-Blackwell GPU deployment remains slow with over 95% of units not yet deployed since December 2024. An AI model with 35 billion parameters now matches trillion-parameter models on long-horizon benchmarks using a different training approach. These advances suggest smaller, more efficient models may compete with larger systems across different tasks.

jamesob's guide to running SOTA LLMs locally

TLDR Dev 2 weeks ago

James Obermayer published a detailed technical guide for building local GPU systems to run state-of-the-art language models, ranging from $2,000 setups with RTX 3090s running Qwen models to $40,000+ configurations with 4× RTX 6000 Pro cards (384GB VRAM) running models approaching Claude Opus quality. The guide includes specific hardware bill-of-materials, BIOS configuration steps, kernel parameters, and Docker-based serving configurations, with the $40k system achieving 27.5 GB/s unidirectional GPU peer-to-peer bandwidth through custom PCIe Gen4 switches. Users can now deploy high-performance local inference with detailed hardware recommendations and ready-to-run software configurations instead of relying on cloud APIs.

Performance per dollar is getting faster and cheaper

TLDR Dev 2 weeks ago

AMD's MI355X GPU achieves comparable inference performance to NVIDIA's Blackwell at roughly 2.75x lower cost per unit, though it historically lagged due to software support issues. Wafer demonstrated 2626 tokens per second aggregate throughput on the MI355X versus 3192 on Blackwell, and 213 tokens per second on GLM5.2, by optimizing quantization, speculative decoding, and kernel selection. As software tooling and agent-driven optimization improve, AMD's cost advantage is becoming increasingly accessible without requiring custom kernel development.

Nvidia taps AI cloud providers to expand compute access for startups

TLDR 2 weeks ago

Nvidia launched a partnership program connecting AI startups with cloud service providers to access computing infrastructure powered by Nvidia chips, with revenue sharing between the parties. The company named Sharon AI and Firmus Technologies as initial partners, with the latter building a data center expected to reach 170,000 Nvidia GPUs. The move addresses startup demand for scarce GPU capacity as the sector faces liquidity constraints and compute availability issues.

Amazon Designing Custom AI Chips for Echo and Fire TV

The Neuron 2 weeks ago

Amazon is designing custom semiconductors for its Echo and Fire TV devices to run AI models locally rather than relying on cloud processing. The company unveiled its AZ3 and AZ3 Pro chips in October, which handle on-device AI inference for improved speed and security. Amazon plans to expand this approach across additional consumer devices and is developing portable AI gadgets that will sync contextual data across its ecosystem of products.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.