TLDRocket
Sign in

Making thousands of open LLMs bloom in the Vertex AI Model Garden

Hugging Face

Hugging Face and Google Cloud just made it stupidly easy to deploy open LLMs on Vertex AI or GKE. Thousands of models, one-click deploy, no infra headaches — that's the pitch.

Based on reporting by Hugging Face — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Hugging Face and Google Cloud are teaming up again, and this time the result is something called Deploy on Google Cloud, a button that turns a fairly annoying engineering task into a couple of clicks. Starting today, developers can take any Hugging Face model tagged with text-generation-inference and push it straight into Vertex AI or Google Kubernetes Engine, either from the Hugging Face model card or from inside Google's own Vertex Model Garden.

The pain point here is a familiar one. Getting an open model from a repo into a production endpoint usually means wrangling GPU quotas, containers, serving frameworks and a dozen configuration files before anything actually answers a prompt. Hugging Face's Text Generation Inference engine handles the serving side, and now Google Cloud handles the infrastructure side, so the whole thing collapses into picking a model and hitting deploy. Zephyr Gemma is the example Hugging Face walks through, and the process really is just: open the Deploy menu on the model card, choose Google Cloud, land in the Vertex AI console, click deploy.

The more interesting half of this launch lives inside Vertex Model Garden itself, where Google has added a new "Deploy From Hugging Face" option. Type in a model ID, and Vertex AI auto-fills tested hardware configs for hundreds of the most-used open LLMs on the Hub. Gated models still work too, you just hand over a Hugging Face access token so the download gets authorized. Wenming Ye, a product manager at Google, framed it as removing the friction of switching between the Hub and the Google Cloud Console, letting developers start wherever is convenient and end up in the same place.

This builds on a partnership the two companies announced earlier this year, and it fits a pattern both sides have been chasing: making open models feel as frictionless to deploy as closed, API-gated ones. Hugging Face gets its catalog embedded in one of the three big clouds, Google gets an instant, curated shortcut to thousands of community models without building that curation itself. Neither side pretends this is the finish line — Hugging Face explicitly says more integrations are coming.

My take — AI-written commentary, not fact-checked reporting

This is the boring-but-important kind of AI news: infrastructure plumbing, not a new model beating a benchmark. But plumbing is exactly what determines whether open models actually get used in production instead of just downloaded and abandoned in a notebook. Every cloud provider chasing this kind of one-click deploy for open weights is a quiet win against the idea that only closed APIs are viable at scale, and I'll take that trade every time.

Read more about this at: Hugging Face

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.