DeepSeek V3
Model ● Covered in 4 stories + Follow
DeepSeek V3 is a 671-billion parameter large language model that uses a Mixture-of-Experts architecture to activate only 37 billion parameters during inference. Recently, DeepSeek released an updated version of the V3 base model under an MIT license with improvements to instruction following and performance in code and math tasks, now available through multiple inference providers including Fireworks and Hyperbolic.
Updated 8 August 2026
Specifications
No specifications recorded yet.
Latest developments
Q2 2026
Q3 2025
Q1 2025
Alibaba releases Qwen2.5 model family including vision-language, extended-context, and mixture-of-experts variants Model release