Gemma 2 27B
Open Weights
Google's largest open-weights model in Gemma 2 family with 27B parameters. Released June-July 2024, trained on 13 trillion tokens of web documents, code, and mathematics. Features interleaved local/global attention, logit soft-capping for stability, Grouped Query Attention for efficiency. Outperforms Qwen1.5 32B baseline and performs "only a few percent below" Llama 3 70B despite being significantly smaller. Optimized to run full precision on single Google Cloud TPU host, NVIDIA A100 80GB, or H100. 256K vocabulary with 4K context using RoPE. Commercially-friendly license. Available via Google AI Studio, Kaggle, Hugging Face.
Strengths
- Outperforms Qwen1.5 32B and approaches Llama 3 70B despite being smaller
- Optimized for single-node inference - runs on A100 80GB or H100
- Commercially-friendly open license enables unrestricted use
- Trained on 13T tokens for comprehensive knowledge
- Interleaved attention and logit soft-capping for improved stability
- Grouped Query Attention enhances inference efficiency
- Available via multiple platforms - Google AI Studio, Kaggle, Hugging Face
Caveats
- 4K context window smaller than competitors (GPT-4o: 128K, Gemini 1.5 Pro: 2M)
- 4K max output tokens lower than flagship models
- 27B size still requires GPU resources (80GB VRAM minimum)
- Primarily English - limited multilingual compared to some competitors
- Smaller than Llama 3 70B - lower ceiling for complex tasks
- Vendor
- Context Window
- 4,096 tokens
- Max Output
- 4,096 tokens
- License
- Gemma License (commercially-friendly)
- Release Date
- 2024-06-27
- Modalities
- text
| Benchmark | Score | Max |
|---|---|---|
| livebench_instruction_following | 71.3 | 100.0 |
| livebench_language | 33.8 | 100.0 |
| livebench_coding | 35.9 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Pricing
Open weights - free. Runs on single A100 80GB or H100. API providers typically $0.30-$0.80 per 1M input tokens