TLDRocket
Sign in

Hardware & Infrastructure

152 summarised stories in Hardware & Infrastructure, each linking back to the original source. Browse all topics →

Monday, 20 July 2026

Google just bet its inference future on a chip built for one model

The New Stack 3 hours ago 2 sources

Google is developing a specialized chip called Frozen v2 designed specifically for its Gemini AI model, which would hardwire parts of Gemini's architecture while keeping weights updatable. The chip is projected to deliver six to ten times more tokens per watt compared to Google's current AI chips. If successful, this approach could significantly reduce inference costs for developers using Gemini while establishing a trend toward model-specific silicon rather than general-purpose accelerators.

Google is working on a new AI chip designed to make Gemini more efficient

TechCrunch AI 4 hours ago 2 sources

Google is developing a custom AI chip called Frozen v2 to run its Gemini models more efficiently, following a strategy shared by other AI companies seeking independence from Nvidia. The chip could deliver 6 to 10 times better efficiency per unit of power compared to Google's current AI chips, with a planned release in 2028. The news boosted investor confidence in Google's massive AI spending plans, sending the stock up 3% as the company aims to prove its $180–190 billion investment strategy will generate returns.

Introducing Cosmos 3 Edge

Hugging Face Blog 9 hours ago 3 sources

NVIDIA released Cosmos 3 Edge, a 4-billion-parameter open-source world model designed to run on edge devices like Jetson modules and RTX GPUs, enabling robots and vision AI systems to understand scenes, predict outcomes, and generate actions in real time. The model achieves real-time inference at 15 Hz on NVIDIA Jetson Thor while generating 32 actions per inference, and ranks first among similar-sized models on VANTAGE-Bench for vision analytics. The release includes post-trained checkpoints, training recipes, and a robot manipulation policy variant, allowing developers to fine-tune the model for specific applications before deploying to edge hardware.

In-House LLM Serving at Netflix

TLDR Dev 14 hours ago

Netflix built an in-house system to serve large language models using vLLM and NVIDIA Triton, integrating it into their existing JVM-based serving infrastructure rather than using external APIs. The platform supports both gRPC and OpenAI-compatible HTTP endpoints, with deployment strategies including red-black and versioned rollouts to handle model updates without dropping requests. The system implements constrained decoding via vLLM's logits processor interface to generate compliant outputs by construction, though this required optimization across vLLM versions to handle batching efficiently at scale.

Bristol Myers Squibb Building Life Science Industry’s Most Advanced AI Factory on NVIDIA Vera Rubin

NVIDIA 14 hours ago

Bristol Myers Squibb deployed its second NVIDIA DGX SuperPOD with eight Vera Rubin NVL72 systems, doubling down on AI infrastructure for drug discovery. The new system delivers 10x the performance per megawatt compared to its predecessor and will be accessible to all BMS scientists globally through a unified AI platform including NVIDIA BioNeMo Agent Toolkit. This removes computational bottlenecks and enables faster drug discovery cycles by democratizing access to supercomputing resources across the organization's research pipeline.

Best Local LLMs You Can Run on a Single 24GB GPU in 2026: Qwen, Gemma, Mistral, DeepSeek Compared

MarkTechPost 1 day ago

A guide compares six open-weight language models optimized for running on a single 24GB GPU, including Qwen3.6-27B, Gemma 4 26B, Mistral Small 3.2 24B, and DeepSeek-R1-Distill-Qwen-32B. These models range from 20B to 35B parameters and use Q4_K_M quantization to fit within memory constraints while leaving room for context and inference overhead. The strategy shifts from squeezing the largest 70B models onto a card to running right-sized 20B–35B dense or efficient mixture-of-experts models that decode faster and leave 1–6GB of headroom for context and serving stack overhead.

Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community

Together AI 1 day ago

Together AI and Y Combinator partnered to provide YC portfolio startups with dedicated GPU cluster access for training and inference workloads. The cluster is fully utilized today and allows startups to reserve compute capacity for short-term sprints at long-term rates without multi-year commitments. YC founders can now provision GPUs in minutes through a self-service portal, eliminating the need to raise funding solely for expensive compute contracts.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.