Skip to main content
llm.info

Gemini 1.5 Pro

Google

Google's flagship model with industry-leading 2 million token context window - the longest available. Sparse mixture-of-experts (MoE) Transformer architecture enables processing massive amounts of data: 2 hours of video, 19 hours of audio, codebases with 60,000 lines, or 2,000 pages of text. Matches or outperforms Gemini 1.0 Ultra on standard benchmarks while using significantly less training compute. Near-perfect retrieval (>99%) up to 10 million tokens in research tests. Major price reductions in October 2024.

Strengths

  • Industry-leading 2 million token context window - longest available commercially
  • Multimodal capabilities - processes text, images, audio, and video natively
  • Near-perfect >99% retrieval on 1M token 'Needle in a Haystack' evaluation
  • Can process 2 hours video, 19 hours audio, or 60K lines of code in one prompt
  • 64% price reduction in October 2024 makes it cost-competitive
  • Outperforms Gemini 1.0 Ultra on half of benchmarks with less training compute
  • Research tested successfully up to 10 million tokens

Caveats

  • Higher pricing above 128K tokens ($2.50 vs $1.25 per 1M input)
  • Max output tokens (8K) lower than some competitors (GPT-5.1: 64K)
  • Processing very large contexts can increase latency significantly
  • Best performance requires careful prompt engineering for massive contexts
  • Multimodal processing adds complexity compared to text-only workflows
Vendor
Google
Context Window
2,000,000 tokens
Max Output
8,192 tokens
License
Proprietary
Release Date
2024-05-14
Modalities
text
image
audio
video
BenchmarkScoreMax
livebench_instruction_following78.2100.0
livebench_language47.7100.0
livebench_coding41.4100.0

Capabilities

Vision
Audio
Video
Tool Use

Pricing

$1.25 per 1M input tokens (<128K), $2.50 per 1M (>128K). 64% price reduction Oct 2024

Reviews

Comments