Transformer
Model ● Covered in 6 stories + Follow
The Transformer is a model architecture introduced in 2017 that uses multi-head self-attention mechanisms for natural language processing tasks. Recent research has focused on improving its efficiency and capabilities through variants like Parcae (which reuses layers for parameter efficiency), Mamba (which replaces attention with state space models for faster inference), and Nyströmformer (which approximates attention in linear time), while unsupervised pre-training combined with Transformers has demonstrated strong performance across language understanding benchmarks.
Updated 4 August 2026
Specifications
No specifications recorded yet.
Latest developments
April 2026
March 2024
August 2022
August 2020
April 2020
June 2018
Relationships
Products & technology
- Nyströmformer derived from this model · 1 source