TechCrunch AI
·
3 days ago
India's smartphone market experienced a 10% shipment decline in Q2 as memory chip manufacturers shifted production toward AI accelerators, driving up costs for consumer devices. The sub-₹15,000 segment saw shipments fall 45% year-over-year, with overall smartphone prices rising between 4% and 68% depending on the model. Consumers are delaying upgrades to approximately four-year cycles, Chinese smartphone brands are retreating from unprofitable markets, and memory shortages are expected to persist until at least the end of 2027.
Sakana AI
Sakana AI and NVIDIA developed new GPU kernels and data formats to accelerate sparse transformer language models by reshaping sparsity patterns to match hardware capabilities rather than forcing hardware adaptation. The hybrid sparsity format (TwELL) achieved over 20% speedups and significant memory and energy savings in billion-parameter scale models. This enables more efficient inference and training of large language models by better exploiting the natural sparsity that emerges in transformer feedforward layers.
NVIDIA
·
4 days ago
NVIDIA introduced the Vera Rubin platform designed to optimize post-training workloads for agentic AI models that continuously adapt and learn from production environments. The Nemotron 3 Ultra model achieved 71.7% on SWE-bench by fixing real software bugs, with Vera Rubin reducing GPU requirements by 75% compared to the previous Blackwell generation for the same training tasks. This shift makes continuous post-training economically viable, allowing AI systems to maintain and improve intelligence throughout their operational lifetime rather than as a one-time process.
The New Stack
·
4 days ago
● 2 sources
Arm and Google announced infrastructure options for running agentic AI workloads, leveraging Google's custom Axion processors alongside accelerators to optimize cost and security. Google's Kubernetes Engine Agent Sandbox running on Axion N4A instances delivers up to 30% better price performance than competing hyperscale cloud providers for orchestrating untrusted AI-generated code. Organizations can now route heavy computational tasks to specialized accelerators while using Axion CPUs for orchestration and management, reducing overall infrastructure costs and enabling safer autonomous agent deployment.
TechCrunch AI
·
4 days ago
General Compute, an AI inference cloud startup, secured a $400 million loan from Upper90 backed by SambaNova inference chips, marking the first time inference-specific hardware was used as collateral. The SN50 chips are designed to provide 16 times faster inference than GPU-based clouds while reducing costs through power efficiency and eliminating the need for expensive water-cooling systems. This financing signals a broader shift toward cheaper inference infrastructure using open-source models as alternatives to expensive frontier AI systems become more cost-competitive.
Simon Willison
·
4 days ago
An article humorously suggests that hyperscalers like Google could offset their data center water consumption by purchasing golf courses and converting them to public parks, noting that Google used 10.9 billion gallons in 2025 while the Coachella Valley's 120 golf courses collectively use 750,000 gallons daily. The proposal calculates that acquiring approximately 40 of the region's golf courses could theoretically match Google's annual water usage. The suggestion is trivial satire with no serious policy implications or concrete action.