Gemini 1.5 Flash
Google's fast, cost-efficient model created through knowledge distillation from Gemini 1.5 Pro. Designed for high-volume, latency-sensitive applications while maintaining strong capabilities. Features 1 million token context window and multimodal support (text, image, audio, video). Excels at summarization, chat, image/video captioning, and data extraction from long documents. Massive price reductions throughout 2024 (78% input, 71% output) make it extremely cost-effective. Improved performance: 7% MMLU-Pro gain, 20% math boost.
Strengths
- Extremely cost-effective - $0.075 per 1M input (78% cheaper than original pricing)
- 1 million token context window for large document processing
- Multimodal capabilities - text, images, audio, video in single model
- Fast and efficient - optimized for high-volume, low-latency applications
- 7% improvement in MMLU-Pro, 20% boost in math benchmarks
- Knowledge distilled from 1.5 Pro maintains high quality at lower cost
- Excellent for summarization, chat, captioning, and data extraction
Caveats
- Lower capabilities than Gemini 1.5 Pro for complex reasoning tasks
- 8K max output tokens lower than some competitors
- Distillation process may lose some nuanced capabilities from Pro
- Best suited for high-volume tasks rather than maximum quality
- Performance on edge cases may not match full Pro model
- Vendor
- Context Window
- 1,000,000 tokens
- Max Output
- 8,192 tokens
- License
- Proprietary
- Release Date
- 2024-05-14
- Modalities
- textimageaudiovideo
| Benchmark | Score | Max |
|---|---|---|
| MMLU | 78.9 | - |
| MMMU | 56.1 | - |
| livebench_language | 30.3 | 100.0 |
| livebench_coding | 40.4 | 100.0 |
| livebench_instruction_following | 80.0 | 100.0 |
Capabilities
Vision
Audio
Video
Tool Use
Resources
Pricing
$0.075 per 1M input tokens (<128K), $0.30 per 1M output. 78% input price reduction Aug 2024