TLDR
·
3 weeks ago
● 2 sources
OpenAI and Broadcom unveiled Jalapeño, a custom AI chip designed for inference workloads, marking OpenAI's first entry into silicon manufacturing. The chip was designed in nine months and will be deployed starting late 2026, with significant scaling expected in 2027 and early 2028. OpenAI aims to reduce dependence on Nvidia GPUs and build a complete technology stack to serve AI models more efficiently and affordably.
Google Research
·
3 weeks ago
Google researchers developed linear elastic caching that dynamically adjusts cache size using lightweight machine learning to optimize the trade-off between memory costs and cache misses. In production testing on Spanner, the approach reduced memory usage by 15.5% and total cost of ownership by approximately 5% while increasing cache misses by only 5.5%. The system frames cache eviction as a ski rental problem where data can be kept in expensive RAM or evicted to slower storage, with a shallow decision tree predicting optimal retention times for each data page.
IBM Research
·
3 weeks ago
● 2 sources
IBM announced 0.7 nanometer transistor chips, the smallest in the world, using a new three-dimensional nanostack architecture with innovations in wafer bonding and memory scaling. The chips are 70% more efficient than IBM's previous 2 nanometer chips from 2021, and could theoretically enable AI accelerators to deliver 9,000 TOPS compared to current accelerators' 1,500 TOPS, potentially reducing training time for large language models from three months to two weeks. The nanostack design could support a decade of further chip innovations by stacking transistors vertically rather than only shrinking them horizontally.