TLDRocket
Sign in

The Transformer Family

Lilian Weng

An article explains the Transformer architecture and describes improvements to the vanilla Transformer model for longer attention spans, reduced memory consumption, and better performance on various tasks. The vanilla Transformer uses multi-head self-attention with fixed segment lengths of L tokens and 6 stacked encoder-decoder layers. Transformer-XL and other enhanced variants address limitations by reusing hidden states between segments and adopting new positional encoding methods.

Why it matters

[Updated on 2023-01-27: After almost three years, I did a big refactoring update of this post to incorporate a bunch of new Transformer models since 2020. The enhanced version of this post is here: The Transformer Family Version 2.0. Please refer to that post on this topic.]

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.