Gemma 2 9B
Google's efficient mid-size open-weights model with 9B parameters. Released June-July 2024, trained on 8 trillion tokens primarily from web documents, code, and mathematics. Delivers class-leading performance, outperforming Llama 3 8B and other open models in its size category. Features Grouped-Query Attention, interleaved sliding window (4096 tokens) and global attention (8192 tokens). Trained with knowledge distillation from larger 27B model for improved quality. 8K context length. Optimized to run at incredible speed on range of hardware - gaming laptops, high-end desktops, cloud setups. Commercially-friendly license. Available via Google AI Studio, Kaggle, Hugging Face.
Strengths
- Outperforms Llama 3 8B and other models in 9B size category
- Runs efficiently on consumer hardware - gaming laptops and desktops
- Knowledge distillation from 27B model improves quality beyond typical 9B
- 8K context window (2x Gemma 2 27B's 4K)
- Interleaved attention optimizes context understanding vs computational cost
- Trained on 8T tokens for comprehensive coverage
- Commercially-friendly license enables unrestricted deployment
Caveats
- 8K context smaller than competitors (GPT-4o: 128K, Llama 3.1: 128K)
- Lower capabilities than 27B or 70B models for complex reasoning
- 4K max output tokens standard for size class
- 9B size still requires 18GB+ VRAM for optimal performance
- Primarily English - limited multilingual capabilities
- Vendor
- Context Window
- 8,192 tokens
- Max Output
- 4,096 tokens
- License
- Gemma License (commercially-friendly)
- Release Date
- 2024-06-27
- Modalities
- text
| Benchmark | Score | Max |
|---|---|---|
| livebench_instruction_following | 58.7 | 100.0 |
| livebench_language | 26.6 | 100.0 |
| livebench_coding | 22.5 | 100.0 |
Capabilities
Pricing
Open weights - free. Runs on gaming laptops, desktops. API providers typically $0.10-$0.30 per 1M input tokens