Gemini 2.0 Flash
Google's next-generation model designed for the agentic era with superior speed and enhanced capabilities. Released December 2024, generally available February 2025. Outperforms Gemini 1.5 Pro on key benchmarks at 2x the speed. Features 1M token context window, multimodal inputs (text, image, audio, video), and groundbreaking multimodal output including natively generated images mixed with text and steerable text-to-speech in multiple languages. Built-in tool use, function calling, and advanced reasoning. Beats 1.5 Flash on MMLU Pro (77.6% vs 67.3%) and Bird SQL (57.4% vs 45.6%).
Strengths
- Outperforms Gemini 1.5 Pro across multiple benchmarks while being 2x faster
- Multimodal output - natively generates images mixed with text and steerable TTS audio
- 1 million token context window for large document processing
- Multimodal input support - text, images, audio, video in single model
- Significant performance gains: +10% MMLU Pro, +11.8% Bird SQL vs 1.5 Flash
- Built-in tool use and superior agentic capabilities
- Simplified pricing structure with higher rate limits
Caveats
- 8K max output tokens lower than some competitors (GPT-5.1: 64K)
- Pricing details not fully specified at general availability
- Agentic features may add complexity for simple use cases
- Multimodal generation is cutting-edge but may have occasional artifacts
- Relatively new model (Feb 2025) - ecosystem still maturing
- Vendor
- Context Window
- 1,000,000 tokens
- Max Output
- 8,192 tokens
- License
- Proprietary
- Release Date
- 2025-02-05
- Modalities
- textimageaudiovideo
| Benchmark | Score | Max |
|---|---|---|
| MMLU Pro | 77.6 | - |
| Bird SQL | 57.4 | - |
| livebench_language | 52.1 | 100.0 |
| livebench_coding | 54.4 | 100.0 |
| livebench_instruction_following | 80.5 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Resources
Pricing
Simplified pricing with higher rate limits (specific pricing TBD)