OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI 2 months ago 36 ● 3 sources
OpenAI and Broadcom have jointly developed Jalapeño, a custom chip designed to optimize inference workloads for large language models. The chip targets improvements in performance and energy efficiency compared to existing inference hardware, though specific benchmark numbers were not disclosed. This development could reduce OpenAI's dependence on third-party inference accelerators and lower operational costs for running LLM services at scale.