Skip to main content
llm.info

Gemma 2 27B

Open Weights

Google

Google's largest open-weights model in Gemma 2 family with 27B parameters. Released June-July 2024, trained on 13 trillion tokens of web documents, code, and mathematics. Features interleaved local/global attention, logit soft-capping for stability, Grouped Query Attention for efficiency. Outperforms Qwen1.5 32B baseline and performs "only a few percent below" Llama 3 70B despite being significantly smaller. Optimized to run full precision on single Google Cloud TPU host, NVIDIA A100 80GB, or H100. 256K vocabulary with 4K context using RoPE. Commercially-friendly license. Available via Google AI Studio, Kaggle, Hugging Face.

Strengths

  • Outperforms Qwen1.5 32B and approaches Llama 3 70B despite being smaller
  • Optimized for single-node inference - runs on A100 80GB or H100
  • Commercially-friendly open license enables unrestricted use
  • Trained on 13T tokens for comprehensive knowledge
  • Interleaved attention and logit soft-capping for improved stability
  • Grouped Query Attention enhances inference efficiency
  • Available via multiple platforms - Google AI Studio, Kaggle, Hugging Face

Caveats

  • 4K context window smaller than competitors (GPT-4o: 128K, Gemini 1.5 Pro: 2M)
  • 4K max output tokens lower than flagship models
  • 27B size still requires GPU resources (80GB VRAM minimum)
  • Primarily English - limited multilingual compared to some competitors
  • Smaller than Llama 3 70B - lower ceiling for complex tasks
Vendor
Google
Context Window
4,096 tokens
Max Output
4,096 tokens
License
Gemma License (commercially-friendly)
Release Date
2024-06-27
Modalities
text
BenchmarkScoreMax
livebench_instruction_following71.3100.0
livebench_language33.8100.0
livebench_coding35.9100.0

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open weights - free. Runs on single A100 80GB or H100. API providers typically $0.30-$0.80 per 1M input tokens

Reviews

Comments