TLDRocket
Sign in

Bringing serverless GPU inference to Hugging Face users

Hugging Face Blog

Hugging Face and Cloudflare launched an integration enabling developers to run open-source AI models as serverless APIs on Cloudflare's GPU infrastructure without managing their own servers. An example RAG application handling 1,000 requests daily with Llama 2 7B would cost approximately $1 per day under the pay-per-request pricing model. Developers can now deploy popular models like Llama, Gemma, and Mistral directly from Hugging Face's Hub using either Cloudflare's REST API or AI SDK. Note: The article's November 2024 update states this integration is no longer available and directs users to alternative deployment options.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.