WebLLM released as an in-browser LLM inference engine that runs locally in browsers using WebGPU
Open source release Provisional 64% confidence first seen
Two pieces of coverage describe the release of WebLLM, an in-browser language model inference engine that uses WebGPU to run selected models locally without server support. The articles also note an approach to execute models in the browser (including streaming chat support) and provide guidance for developers integrating it into web applications.