TLDRocket
Sign in

Hardware & Infrastructure

448 summarised stories in Hardware & Infrastructure, each linking back to the original source. Browse all topics →

Thursday, 27 August 2026

Anthropic's new hardware standard lets AI agents control the physical world

Ars Technica 1 week ago 29 4 sources

Anthropic introduced the Model Hardware Standard (MHS) to let agentic AI systems interface with and control physical devices via standardized drivers. Anthropic says MHS research preview integration could cut experimental setup from weeks or months to hours or minutes. This shifts AI agents beyond computer-only actions by providing a common “translation” layer for data sharing and device coordination across networks.

The AI storage stack gets an inference-era rethink

SiliconANGLE 1 week ago 19 2 sources

DataDirect Networks Inc. unveiled DDN Enterprise AI HyperPOD with Super Micro Computer and Solidigm to simplify enterprise AI inference storage and data management based on Nvidia’s AI Data Platform. The solution supports scaling from a single rack and expanding by “pop[ping] in another rack,” rather than committing upfront to large numbers of GPUs. The shift moves storage toward a capacity-focused role for LLM workloads and aims to reduce customer integration complexity while enabling on-prem and hybrid sovereign AI deployments with multi-tenancy.

Previewing the Model Hardware Standard

Anthropic 37 3 sources

Anthropic and HHMI Janelia Research Campus launched a research preview of the Model Hardware Standard (MHS), a shared specification for AI agents to safely operate multiple physical instruments in parallel. The preview is being shared with a first group of scientific research labs and advanced manufacturers, with partners including Genentech and HHMI Janelia, and it targets reducing hardware integration from weeks or months to hours or minutes. MHS will let devices connect through a standardized driver and common control protocols, enabling autonomous lab workflows with real-time updates and fewer bespoke integrations, with plans to open-source it later after safety evaluations.

Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics

Amazon Web Services 1 week ago 40

Deepgram added Enhanced Metrics to its Deepgram speech-to-text and text-to-speech deployments on Amazon SageMaker AI to close gaps in billing and performance observability. The billing transparency stream reports ConsumedUnits, using the same metered inference-unit values that drive AWS Marketplace metered billing. The result is CloudWatch metrics for per-request billing/feature usage plus engine and per-GPU visibility via Prometheus and OpenTelemetry within SageMaker’s detailed observability, without deploying an agent, sidecar, or collector.

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

Amazon Web Services 1 week ago 4

Heidi Health, with AWS and NVIDIA, deployed CUDA Multi-Process Service (MPS) and NVIDIA Triton Inference Server on Amazon EC2 to serve its fine-tuned Parakeet TDT ASR model more efficiently. Using MPS reduced the number of GPU instances needed from 16 to 4 while keeping sub-second transcription latency at 92.1 requests per second per GPU. The setup changes production operations by partitioning one L40S GPU into concurrent MPS execution contexts and batching/scheduling requests through Triton, cutting infrastructure cost without increasing latency.

Nvidia CFO says about half of its data center business comes from customers beyond hyperscalers

Fortune 30 3 sources

Nvidia reported fiscal second-quarter results and CFO Colette Kress said about half of its data center revenue comes from customers beyond major hyperscalers. Data Center revenue was $89.0 billion, and Kress said non-hyperscaler growth covers roughly half of the business. The company’s outlook for AI infrastructure demand is framed as increasingly diversified across enterprise and regional/sovereign and edge deployments rather than tied mainly to a few big cloud buyers.

Deep Learning Weekly: Issue 470

Deep Learning Weekly 1 week ago 45 2 sources

Z.ai open-sourced GLM-5.3-Flash, a stealth “Ox Alpha” mixture-of-experts model with hybrid sparse-plus-linear attention and a 1M-token multimodal context, and multiple other deep-learning updates were shared alongside it. The issue’s FreeToken paper reports edge-native MoE serving that supports running a 753B GLM-5.2 on a single workstation GPU. Benchmarks and tool writeups in the issue shift focus toward more practical local/edge deployment and evaluating real-world LLM/agent behavior rather than only final answers.

AI’s memory crunch is coming for Android apps

TechCrunch 1 week ago 13

Google updated Android app quality requirements to address memory chip shortages linked to the AI data center boom. The new thresholds must be met by February 2027. Android apps and games will have to limit dynamic memory and bitmap usage, meet code optimization targets, and use tools like alerts plus a Memory Limiter to prevent exceeding device memory limits.

How much of a problem is AI’s water use?

Ars Technica 1 week ago 10 2 sources

Locals in Texas, New Mexico, and Arizona have increasingly protested U.S. data center expansion tied to AI capacity, citing water use, energy demand, and pollution concerns. Data center equipment can run at internal temperatures up to 176° Fahrenheit, driving heat removal that can involve water. The protests and public debate intensify while online claims about water use vary widely and available company consumption data remains incomplete or inconsistent.

GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

NVIDIA 1 week ago 17 2 sources

NVIDIA announced multiple updates to GeForce NOW, adding new DLSS 4.5 tuning controls and expanding device and platform support across Steam, GOG, Firefox, and more Fire TV options. Ultimate members can redeem a limited-time offer to get CONTROL Resonant at launch with a 12-month GeForce NOW Ultimate membership, running Tuesday Aug. 25 through Sunday Sept. 27. The service will broaden where games can be streamed and how they can be configured, while new cloud-launch titles like CONTROL Resonant and STAR WARS Zero Company arrive on the platform.

Delivering Vera: NVIDIA’s First CPU Built for Agents Is Shipping Now

NVIDIA Blog 1 week ago 38 3 sources

NVIDIA is delivering its Vera CPU systems to cloud providers and AI labs as Vera begins shipping at scale. The systems being handed off include AWS’s first NVIDIA Vera CPU server and NVIDIA Vera Rubin GPU delivered in Seattle on Aug. 27, 2026. This expands Vera’s rollout across major customers, including AWS and prior deliveries to Oracle Cloud Infrastructure, Anthropic, OpenAI, and SpaceXAI.

Jensen Huang says Nvidia achieved AGI, again — not that it matters

The Verge 1 week ago 28

Nvidia CEO Jensen Huang said on the company’s earnings call that Nvidia has achieved AGI again and then dismissed the milestone as senseless. He made the claim on Wednesday. The practical result is that the statement doesn’t change industry uncertainty because there’s still no clear definition or verification method for AGI.

The Sequence Opinion #921: AI’s Sixth Layer Is Finance

TheSequence 1 week ago 12

The Sequence Opinion argues that AI’s sixth layer is finance, sitting beneath the commonly described layers that run from energy to applications. It points to “#921” and says Jensen Huang’s five-layer model is incomplete because the “cake” also exists on balance sheets and through financial constraints. As a result, the article reframes AI bottlenecks and growth as depending on power contracts, depreciation, debt covenants, and teams managing expensive hardware uptime rather than only on model quality.

Nvidia Posts Another Blockbuster Quarter, But Debt is Rising in AI Frenzy

Trending Topics 1 week ago 1 3 sources

Nvidia reported fiscal 2027 Q2 results that beat expectations, with its data center business driving the quarter while investors focused on guidance. It projected $108 billion in current-quarter revenue, slightly above an estimate of about $104.2 billion. Nvidia’s stock rose on the outlook and it also highlighted expanded AI infrastructure partnerships, while noting rising debt and margin pressure.

India’s data center boom is leaving the people it displaces with nothing

Rest of World 1 week ago 22

Digital Empowerment Foundation founder Osama Manzar says India’s data-center boom is displacing communities without consultation while also tying local impacts to AI-era demand. He cites data centers on track to consume water equivalent to the annual domestic needs of 1.3 billion people and electricity matching the demand of 650 million people by 2030. Protests and demands for community consent grow, with arguments for local benefits and more decentralized siting rather than treating the data-center presence as non-negotiable.

Nvidia gave its first-ever year-ahead forecast—a 70% growth bombshell meant to silence AI bubble critics and ‘circular financing’ doomsayers

Fortune 21 4 sources

Nvidia issued its first year-ahead forecast, projecting a 70% revenue increase next fiscal year, after reporting results that beat expectations and citing accelerating demand for its AI chips. The company forecast fiscal 2028 revenue growth of 70% and said customers’ forecasts indicate growth doubling next year, while also noting supply constraints limit delivery. Nvidia’s new guidance resets market expectations for next-year sales and sets updated supply, margin, and financing-related outlooks for investors to follow.

[AINews] Hot Chips: OpenAI’s Jalapeño, Cerebras CS-5, Groq 3 LPX, Apple M6

Latent Space 1 week ago 39 3 sources

OpenAI presented details of its Jalapeño custom inference chip at Hot Chips, shifting the focus to performance per watt and claiming improved efficiency and latency on real model workloads. The chip is rated at 700W but reportedly stayed at or below 550W in tested runs. OpenAI says deployment into its own infrastructure will start by year-end, with additional Gen 2 and Gen 3 work already underway, reinforcing an inference-architecture and kernel-optimization loop rather than an application-only approach.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.