TLDRocket
Sign in

Run a vLLM Server on HF Jobs in One Command

Hugging Face Blog

Hugging Face Jobs now allows users to deploy a vLLM server with a single command that creates an OpenAI-compatible endpoint on HF infrastructure without manual server provisioning. The service costs $1.50 per hour for an a10g-large GPU flavor and bills per second of usage. Users can query the endpoint from any location using curl or the OpenAI Python client with token-based authentication, making it suitable for testing, evaluations, and batch generation workloads.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.