TLDRocket
Sign in

Hugging Face and FriendliAI partner to supercharge model deployment on the Hub

Hugging Face

Hugging Face just added FriendliAI as a one-click deploy option on the Hub. Now you can spin up models on H100 GPUs straight from a model card, no separate setup.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Hugging Face has struck a deal with FriendliAI, folding the inference company's infrastructure directly into the "Deploy this model" button on the Hub. That means anyone browsing a model card can now push it live on FriendliAI's servers without leaving the page, using an existing Friendli Suite account.

The two companies aren't strangers. FriendliAI built a Hugging Face integration last year that let its customers pull in thousands of open-source models from the Hub. This new move flips that relationship around, putting FriendliAI's deployment tools inside Hugging Face's own interface instead of the other way around. Artificial Analysis has ranked FriendliAI as the fastest GPU-based inference provider around, largely thanks to continuous batching, native quantization, and autoscaling that adjusts to demand on the fly.

Practically, there are two paths once you hit deploy. Friendli Dedicated Endpoints puts your model on NVIDIA H100 GPUs as a managed service, aimed at production workloads where FriendliAI claims its optimizations cut the number of GPUs needed to hit a given performance target — which matters a lot given how pricey H100 capacity still is. The other option, Friendli Serverless Endpoints, is built for people who just want to test-drive open-source models through simple APIs, and the deployment page even lets you chat with a model while your own instance is still spinning up.

None of this is revolutionary on its own. Model hosting marketplaces have existed for years, and Hugging Face already partners with several inference providers. But shaving away the friction between picking a model and running it on solid hardware is exactly the kind of unglamorous plumbing that decides whether smaller teams actually ship things or just keep bookmarking cool repos.

My take — AI-written commentary, not fact-checked reporting

This is Hugging Face doing what it does best: quietly becoming the default control panel for open-source AI by stitching in whoever's fastest at the moment, rather than building its own inference stack from scratch. I like that the incentive here still runs toward open models getting cheaper to run, not proprietary APIs getting more locked down — that's the trend worth watching, not the H100 speed claims, which every provider makes until someone else claims to be faster next quarter.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.