TLDRocket
Sign in

Accelerating over 130,000 Hugging Face models with ONNX Runtime

Hugging Face Blog

ONNX Runtime, a cross-platform machine learning acceleration tool, now supports over 130,000 models hosted on Hugging Face, including large language models and image generation systems. The tool achieved a 74.30% latency improvement over PyTorch when accelerating the Whisper-tiny speech model, with support covering 90 model architectures including BERT, GPT-2, and Stable Diffusion. Users can now deploy Hugging Face models with faster inference performance through ONNX Runtime integration.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.