Mistral 7B
Open Weights
Mistral AI
Mistral AI's compact but powerful 7.3B parameter model released September 2023. First model from Mistral AI, setting new standards for efficiency and performance in its size class. Outperforms Llama 2 13B on all benchmarks and Llama 1 34B on many benchmarks despite being significantly smaller. Approaches CodeLlama 7B code performance while maintaining strong English capabilities. Features Grouped-Query Attention for faster inference and Sliding Window Attention (4096 tokens) for efficient long sequences. Apache 2.0 license. Deployable anywhere including locally, cloud (AWS/GCP/Azure), available on HuggingFace.
Strengths
- Outperforms Llama 2 13B (nearly 2x size) on all benchmarks
- Outperforms Llama 1 34B (nearly 5x size) on many benchmarks
- Approaches CodeLlama 7B code performance while excelling at English tasks
- Apache 2.0 license enables unrestricted use and modification
- Grouped-Query Attention enables faster inference than standard attention
- Sliding Window Attention efficiently handles sequences at lower cost
- Deployable anywhere - local, cloud, edge devices
Caveats
- 7.3B size limits capabilities compared to larger models (70B+)
- 8K context window smaller than modern models (Llama 3.1: 128K)
- Text-only (no vision or multimodal)
- Superseded by newer Mistral models for cutting-edge performance
- 4K max output tokens standard for size class
- Vendor
- Mistral AI
- Context Window
- 8,192 tokens
- Max Output
- 4,096 tokens
- License
- Apache-2.0
- Release Date
- 2023-09-27
- Modalities
- text
| Benchmark | Score | Max |
|---|---|---|
| livebench_coding | 13.9 | 100.0 |
| livebench_language | 11.9 | 100.0 |
| livebench_instruction_following | 62.8 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Pricing
Open source (Apache 2.0) - free. Runs locally. API providers typically $0.20-$0.35 per 1M input tokens