TLDRocket
Sign in

Model-Preserving Adaptive Rounding with YAQA

Together AI

Researchers introduced YAQA, a weight-only quantization method for large language models that directly minimizes KL divergence to the original model by using a Kronecker-factored Hessian approximation instead of layer-wise activation error minimization. YAQA reduced KL divergence by more than 30% across multiple models and quantizers compared to existing methods like LDLQ and GPTQ. The approach enables better model compression without requiring retraining and works with any hardware quantizer or memory-bound quantization method.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.