ONNX Runtime
Model ● Covered in 4 stories + Follow
ONNX Runtime is a cross-platform machine learning acceleration tool that optimizes inference performance for models deployed across different hardware backends. Recent developments show it now supports over 130,000 Hugging Face models including large language models and image generation systems, with reported latency improvements ranging from 74% for speech models to 229% for image generation models compared to PyTorch, and integration with tools like Hugging Face Optimum enabling faster inference for production workloads.
Updated 8 August 2026
Specifications
No specifications recorded yet.
Latest developments
June 2026
January 2024
October 2023
May 2022
Relationships
Products & technology
- Integrated with Hugging Face · 1 source
- Integrated with SD Turbo · 1 source
- Integrated with SDXL Turbo · 1 source
- Optimum integrated with this model · 1 source