DeepSeek V3
Model ● Covered in 4 stories + Follow
DeepSeek V3 is a 671-billion parameter large language model that uses a Mixture-of-Experts architecture to activate only 37 billion parameters during inference. Recently, DeepSeek released an updated version of the V3 base model under an MIT license with improvements to instruction following and performance in code and math tasks, now available through multiple inference providers including Fireworks and Hyperbolic.
Updated 8 August 2026
Specifications
No specifications recorded yet.
Latest developments
April 2026
July 2025
March 2025
January 2025
Alibaba releases Qwen2.5 model family including vision-language, extended-context, and mixture-of-experts variants Model release