TLDRocket
Sign in

Tools & Coding

975 summarised stories in Tools & Coding, each linking back to the original source. Browse all topics →

Monday, 18 March 2024

Quanto: a PyTorch quantization backend for Optimum

Hugging Face 2 years ago 28

Hugging Face released Quanto, a PyTorch quantization backend for Optimum that reduces model size and computational costs by converting weights and activations to lower-precision data types like int8 or float8. The tool supports int2, int4, int8, and float8 weights across any model architecture and device (CPU, GPU, Apple Silicon), with accelerated int8-int8 and mixed-precision matrix multiplications on CUDA hardware. Quanto integrates directly into the transformers library, allowing developers to quantize models in a few lines of code without restricting themselves to specific model configurations or device types.

Easily Train Models with H100 GPUs on NVIDIA DGX Cloud

Hugging Face 2 years ago 37

Hugging Face launched Train on DGX Cloud, a service allowing Enterprise Hub organizations to fine-tune AI models using NVIDIA H100 GPUs through a no-code interface integrated into the Hugging Face Hub. The service charges $8.25 per GPU hour for H100 instances, with an example showing that fine-tuning Mistral 7B on 1,500 samples costs approximately $0.45. Users can now access GPU compute on-demand without writing training scripts, with fine-tuned models automatically saved to private repositories. (Note: The service was deprecated as of April 10, 2025.)

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.