The Transformer Family
Lilian Weng
An article explains the Transformer architecture and describes improvements to the vanilla Transformer model for longer attention spans, reduced memory consumption, and better performance on various tasks. The vanilla Transformer uses multi-head self-attention with fixed segment lengths of L tokens and 6 stacked encoder-decoder layers. Transformer-XL and other enhanced variants address limitations by reusing hidden states between segments and adopting new positional encoding methods.
Why it matters
[Updated on 2023-01-27: After almost three years, I did a big refactoring update of this post to incorporate a bunch of new Transformer models since 2020. The enhanced version of this post is here: The Transformer Family Version 2.0. Please refer to that post on this topic.]