Welcome Mixtral - a SOTA Mixture of Experts on Hugging Face
Hugging Face Blog ● Covered by 4 sources
Mistral released Mixtral 8x7B, a mixture-of-experts language model that outperforms GPT-3.5 on most benchmarks and is now integrated across Hugging Face's ecosystem including Transformers, Inference Endpoints, and Text Generation Inference. The model contains 45 billion effective parameters across 8 specialized experts, decodes at the speed of a 12-billion parameter model, and achieves 40.2% accuracy on HumanEval coding tasks. Users can now run inference, fine-tune on single GPUs, and deploy Mixtral through multiple Hugging Face tools with support for quantization to reduce memory requirements from 90GB to 23GB.