Mixtral 8x7B
Open Weights
Mistral AI
Mistral AI's groundbreaking Sparse Mixture of Experts model with 46.7B total parameters. Released December 2023, uses 13B active parameters per token while achieving performance of 6x larger models. Outperforms Llama 2 70B on most benchmarks with 6x faster inference. Matches or exceeds GPT-3.5 across all evaluated benchmarks. Features 32K context window with sliding window attention enabling theoretical 128K token span. Best open-weights chatbot model as of December 2023 per MT-Bench. Supports 5 languages (English, French, German, Spanish, Italian). Apache 2.0 license enables unrestricted commercial use.
Strengths
- Outperforms Llama 2 70B with 6x faster inference via sparse MoE architecture
- Matches or exceeds GPT-3.5 across all benchmarks evaluated
- 46.7B total parameters but only 13B active per token (cost-effective)
- Apache 2.0 license enables unrestricted commercial use and modification
- 32K context with sliding window attention (theoretical 128K span)
- Best open-weights chatbot model (MT-Bench, December 2023)
- Supports 5 major languages including English, French, German
Caveats
- 46.7B size requires significant GPU memory despite sparse activation
- Grouped Query Attention adds complexity vs standard attention
- 32K practical context smaller than competitors (Gemini: 1M-2M)
- Text-only (no vision or multimodal capabilities)
- Superseded by newer Mistral models for cutting-edge performance
- Vendor
- Mistral AI
- Context Window
- 32,000 tokens
- Max Output
- 4,096 tokens
- License
- Apache-2.0
- Release Date
- 2023-12-11
- Modalities
- text
| Benchmark | Score | Max |
|---|---|---|
| livebench_language | 13.8 | 100.0 |
| livebench_coding | 13.3 | 100.0 |
| livebench_instruction_following | 57.9 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Pricing
Open source (Apache 2.0) - free. API providers typically $0.45-$0.70 per 1M input tokens