Transformer
Model ● Covered in 6 stories + Follow
The Transformer is a model architecture introduced in 2017 that uses multi-head self-attention mechanisms for natural language processing tasks. Recent research has focused on improving its efficiency and capabilities through variants like Parcae (which reuses layers for parameter efficiency), Mamba (which replaces attention with state space models for faster inference), and Nyströmformer (which approximates attention in linear time), while unsupervised pre-training combined with Transformers has demonstrated strong performance across language understanding benchmarks.
Updated 4 August 2026
Specifications
No specifications recorded yet.
Latest developments
2026
2024
2022
2020
2018
Relationships
Products & technology
- Nyströmformer derived from this model · 1 source