Transformer
Model ● Covered in 6 stories + Follow
The Transformer is a model architecture introduced in 2017 that uses multi-head self-attention mechanisms for natural language processing tasks. Recent research has focused on improving its efficiency and capabilities through variants like Parcae (which reuses layers for parameter efficiency), Mamba (which replaces attention with state space models for faster inference), and Nyströmformer (which approximates attention in linear time), while unsupervised pre-training combined with Transformers has demonstrated strong performance across language understanding benchmarks.
Updated 4 August 2026
Specifications
No specifications recorded yet.
Latest developments
Q2 2026
Q1 2024
Q3 2022
Q3 2020
Q2 2020
Q2 2018
Relationships
Products & technology
- Nyströmformer derived from this model · 1 source