TLDRocket
Sign in

Nyströmformer: Approximating self-attention in linear time and memory via the Nyström method

Hugging Face Blog

Nyströmformer approximates the quadratic-complexity self-attention mechanism in standard Transformers by applying the Nyström matrix approximation method to queries and keys rather than directly to the attention matrix. The approach reduces complexity from O(n²) to O(n) while maintaining competitive performance with just 32 or 64 landmarks across sequences of 4,096 to 8,192 tokens. Four pre-trained checkpoints are available on HuggingFace for masked language modeling at different sequence lengths, allowing practitioners to trade off between model capacity and computational efficiency.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.