Nvidia said its DLSS 5 AI rendering that was initially planned for this evening on only one game and only RTX 50 GPUs will also be brought to older RTX 40-series GPUs. Nvidia confirmed the RTX 40-series expansion in a spokesperson statement to The Verge. As a result, DLSS 5 support widens beyond RTX 50, while gamers still won’t get full control over how it’s used.
Australia’s data centres have expanded rapidly and local residents are increasingly pushing back over noise, water use, and environmental impacts as more capacity is planned to support AI services. The article cites potential water use of up to 25% of Sydney’s drinking water by 2035 and notes data-centre energy demands could triple by 2030 nationwide. A new 2027 legal requirement will force large-scale data centres to underwrite power supplies, limit water use, and fund extra water infrastructure, while some community groups want a pause on further development.
Amazon’s EKS Auto Mode and related platform components were instrumented end-to-end for GPU inference pod startup, revealing six sequential bottlenecks that add up to an 8-minute time-to-first-token response on a 70B-class model. For the 203 GB model, downloading weights from S3 took 423 seconds (about 92% of startup time), partly due to a pattern that left 98% of bandwidth idle. Configuration changes and platform features cut cold-start latency from 8–15 minutes to under 1 minute (with warm-node restarts dropping to under 30 seconds), mainly by fixing S3 weight loading, CUDA kernel compilation caching, and some node/image startup steps.
Amazon Bedrock AgentCore was used to migrate a LangGraph customer-support agent from running on a user-managed container with local session state to a hosted runtime with gateway-published tools and durable memory. In the walkthrough’s Stage 1, 45 lines inside the agent changed, alongside 22 lines of new supporting code and 85 lines imported unchanged. As a result, compute/OS patching and session isolation move to AgentCore, tool authorization is centralized at the gateway, and conversation state is stored in AgentCore memory while inference call behavior stays the same.
NVIDIA, Microsoft, and partners announced tools and devices at IFA 2026 to make running local AI agents faster and easier on NVIDIA hardware. Local inference throughput is claimed to be up to 1.9x higher on a GeForce RTX 5090 via updated llama.cpp optimizations. The update adds simplified local setup apps, a personal AI router that parallelizes workloads across idle PCs, and new RTX Spark Windows PCs launching in October 2026.
Nvidia launched Nvidia Personal AI Router (PAIR), an open source software router that lets idle home Macs and PCs run small local models for agent workflows using subagents. In Nvidia’s example, two RTX 5090 PCs with 32 GB RAM each sped up work with five subagents by about 1.6x. It changes agent execution by routing requests to eligible local machines behind a single interface rather than pooling GPUs or splitting one inference across computers.
Nvidia launched Personal AI Router (PAIR), an open-source software tool that links compatible home PCs for running local AI inference with setups like Ollama and LM Studio. It supports Nvidia GeForce RTX 20-series GPUs and newer (plus RTX Pro GPUs and DGX Spark systems). Users can now pool their idle machines into a single personal “AI data center” for agentic workflows without adding hardware.
Nvidia will ship RTX Spark laptops using its Arm-based N1X chip in two configurations starting in October. The top configuration includes a 6144-core Blackwell RTX GPU, paired with a 20-core Grace CPU and 24GB to 128GB of unified memory. This adds an RTX Spark N1X option for both high-end laptops and small-form-factor PCs, while the N1 chip is not arriving this year.
Perplexity open sourced Lily, a local single-process Rust + Metal inference engine for running Qwen3.6-35B-A3B on Apple silicon via an OpenAI-compatible chat-completions API. It reported averaging 4,156 prefill tokens/s versus 3,388 for MLX-LM (1.23x) and 170.0 decode tokens/s versus 126.4 (1.35x) on a 40-core, 128 GB M5 Max. This shifts execution away from PyTorch/MLX into a Metal-kernel runtime specialized to one model and hardware family, with standalone demos and an inference server codepath for deployment.
Every AI story that matters,
in your inbox by 8am.
TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the
day in two minutes. Follow companies and topics for alerts, or get the
briefing in Slack. Free, no spam, unsubscribe anytime.
Reading TLDRocket needs no cookies, and the readership counts we rely on come from
our own cookieless analytics. Google Analytics is the exception: it sets cookies and
reports to Google, so it stays switched off until you allow it. You can change your
mind any time from “Cookie settings” in the footer.
Strictly necessary
Session security and form protection (tldrocket-session,
XSRF-TOKEN, 2 hours). The site cannot work without them,
so they need no consent.
Always on
Google Analytics 4 (_ga,
_ga_<id>, up to 2 years). Measures which
stories and sections readers use. Google acts as a third-party processor and may
store the data outside the EU. No advertising, no profiling, no data sold.