TLDRocket
Sign in

Tools & Coding

975 summarised stories in Tools & Coding, each linking back to the original source. Browse all topics →

Tuesday, 10 May 2022

Accelerated Inference with Optimum and Transformers Pipelines

Hugging Face 4 years ago 26

Hugging Face released Optimum 1.2 with inference support for Transformers pipelines using ONNX Runtime, allowing users to accelerate model inference with a compatible API. In a tutorial example converting RoBERTa for question-answering, quantization reduced model size from 473 MB to 291 MB while maintaining prediction accuracy. Users can now replace standard Transformers model classes with optimized ORTModel equivalents to run faster inference on production workloads.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.