Ben's Bites
·
3 weeks ago
● 3 sources
Etched, a hardware startup focused on AI inference, raised $800M and secured $1B+ in backlog orders while achieving first-silicon success on TSMC 4nm in under three years. The company built inference chips and clusters through vertical integration with a team of 400+ engineers from NVIDIA, Google, and other major chip programs. OpenAI released GPT-5.6 with limited access to select partners, with plans for broader availability pending government approval.
NVIDIA
·
3 weeks ago
NVIDIA's inference software stack has reduced token costs for DeepSeek V4 by up to 5x on the Blackwell platform within one month through optimizations across production operations, application acceleration, and infrastructure access layers. Companies like Baseten, Cognition, and Deep Infra are using NVIDIA's TensorRT-LLM and Dynamo frameworks to achieve throughput gains ranging from 30% to 50% improvements in token generation speed. The full-stack approach compounds individual optimizations to increase Blackwell token throughput per GPU by up to 20x, enabling lower cost-per-token for production AI inference workloads.
OpenAI Blog
·
3 weeks ago
OpenAI engineers analyzed large-scale core dumps to identify the root causes of rare infrastructure crashes in their systems. They discovered an 18-year-old software bug alongside a hardware fault, using the crash data to trace issues that had persisted undetected for nearly two decades. The findings enabled them to fix both the legacy code defect and address the hardware problem, improving system reliability.