Gemini 1.5 Pro
Google's flagship model with industry-leading 2 million token context window - the longest available. Sparse mixture-of-experts (MoE) Transformer architecture enables processing massive amounts of data: 2 hours of video, 19 hours of audio, codebases with 60,000 lines, or 2,000 pages of text. Matches or outperforms Gemini 1.0 Ultra on standard benchmarks while using significantly less training compute. Near-perfect retrieval (>99%) up to 10 million tokens in research tests. Major price reductions in October 2024.
Strengths
- Industry-leading 2 million token context window - longest available commercially
- Multimodal capabilities - processes text, images, audio, and video natively
- Near-perfect >99% retrieval on 1M token 'Needle in a Haystack' evaluation
- Can process 2 hours video, 19 hours audio, or 60K lines of code in one prompt
- 64% price reduction in October 2024 makes it cost-competitive
- Outperforms Gemini 1.0 Ultra on half of benchmarks with less training compute
- Research tested successfully up to 10 million tokens
Caveats
- Higher pricing above 128K tokens ($2.50 vs $1.25 per 1M input)
- Max output tokens (8K) lower than some competitors (GPT-5.1: 64K)
- Processing very large contexts can increase latency significantly
- Best performance requires careful prompt engineering for massive contexts
- Multimodal processing adds complexity compared to text-only workflows
- Vendor
- Context Window
- 2,000,000 tokens
- Max Output
- 8,192 tokens
- License
- Proprietary
- Release Date
- 2024-05-14
- Modalities
- textimageaudiovideo
| Benchmark | Score | Max |
|---|---|---|
| livebench_instruction_following | 78.2 | 100.0 |
| livebench_language | 47.7 | 100.0 |
| livebench_coding | 41.4 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Resources
Pricing
$1.25 per 1M input tokens (<128K), $2.50 per 1M (>128K). 64% price reduction Oct 2024