Skip to main content
llm.info

Mixtral 8x7B

Open Weights

Mistral AI

Mistral AI's groundbreaking Sparse Mixture of Experts model with 46.7B total parameters. Released December 2023, uses 13B active parameters per token while achieving performance of 6x larger models. Outperforms Llama 2 70B on most benchmarks with 6x faster inference. Matches or exceeds GPT-3.5 across all evaluated benchmarks. Features 32K context window with sliding window attention enabling theoretical 128K token span. Best open-weights chatbot model as of December 2023 per MT-Bench. Supports 5 languages (English, French, German, Spanish, Italian). Apache 2.0 license enables unrestricted commercial use.

Strengths

  • Outperforms Llama 2 70B with 6x faster inference via sparse MoE architecture
  • Matches or exceeds GPT-3.5 across all benchmarks evaluated
  • 46.7B total parameters but only 13B active per token (cost-effective)
  • Apache 2.0 license enables unrestricted commercial use and modification
  • 32K context with sliding window attention (theoretical 128K span)
  • Best open-weights chatbot model (MT-Bench, December 2023)
  • Supports 5 major languages including English, French, German

Caveats

  • 46.7B size requires significant GPU memory despite sparse activation
  • Grouped Query Attention adds complexity vs standard attention
  • 32K practical context smaller than competitors (Gemini: 1M-2M)
  • Text-only (no vision or multimodal capabilities)
  • Superseded by newer Mistral models for cutting-edge performance
Vendor
Mistral AI
Context Window
32,000 tokens
Max Output
4,096 tokens
License
Apache-2.0
Release Date
2023-12-11
Modalities
text
BenchmarkScoreMax
livebench_language13.8100.0
livebench_coding13.3100.0
livebench_instruction_following57.9100.0

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open source (Apache 2.0) - free. API providers typically $0.45-$0.70 per 1M input tokens

Reviews

Comments