TLDRocket
Sign in

Model Compression

20 summarised stories about Model Compression, each linking back to the original source. Browse all topics →

+ Follow this topic

Thursday, 5 June 2025

Model-Preserving Adaptive Rounding with YAQA

Together AI 1 year ago 1

Researchers introduced YAQA, a weight-only quantization method for large language models that directly minimizes KL divergence to the original model by using a Kronecker-factored Hessian approximation instead of layer-wise activation error minimization. YAQA reduced KL divergence by more than 30% across multiple models and quantizers compared to existing methods like LDLQ and GPTQ. The approach enables better model compression without requiring retraining and works with any hardware quantizer or memory-bound quantization method.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.