TLDRocket
Sign in

Baseten on Hugging Face Inference Providers 🔥

Hugging Face

Hugging Face just added Baseten as an inference provider on the Hub. You can now run models like DeepSeek V4 Flash and GLM-5.2 straight from HF's SDKs or website using Baseten's infra.

Hugging Face keeps stacking inference providers onto its Hub, and the newest addition is Baseten, an AI infrastructure company that handles serverless inference and training for a growing list of frontier models. The integration slots Baseten right into the model pages, the JS and Python SDKs, and the usual routing setup Hugging Face has been building out for months.

For now, Baseten's slice of the Hub covers conversational and text-generation tasks, which means access to open-weight LLMs such as Kimi K3, the latest DeepSeek V4 Flash, and GLM-5.2. Hugging Face says more task types are coming, and Baseten's own model catalog stretches well beyond text into things like text-to-speech, so there's room for this partnership to widen fast.

The mechanics are the same two-lane system Hugging Face uses with every provider it onboards. Bring your own Baseten API key and requests go direct, with Baseten billing you on your own account. Skip that and let Hugging Face route the call, and you're billed through your HF account at the provider's standard rate, no markup added, at least for now, since Hugging Face has left the door open for revenue-sharing deals down the line.

Developers calling DeepSeek V4 Flash through Baseten just point the OpenAI-style client at Hugging Face's router endpoint and tag the model with a `:baseten` suffix. It's a small detail, but it's the kind of thing that makes swapping providers feel like changing a URL rather than rewriting an integration. And because Hugging Face's Inference Providers now plug into agent harnesses like Pi, OpenCode, and OpenClaw, Baseten-hosted models can drop into existing agent workflows without extra glue code.

PRO subscribers get $2 of monthly inference credit usable across any provider on the Hub, plus perks like ZeroGPU and higher rate limits, and free-tier users still get a small quota to test things out.

My take

Every provider Hugging Face bolts onto its router makes the Hub look less like a model zoo and more like a genuine utility layer for AI, and that's the smart long game here. The real story isn't Baseten specifically, it's that Hugging Face is quietly becoming the default routing layer for open-weight models the same way cloud brokers did for compute, and providers that don't plug in risk becoming invisible to a huge chunk of developers who now default to the Hub first.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.