Gemini 1.0 Pro
Google's first-generation multimodal model optimized for wide range of tasks. Released December 2023 as part of three-model Gemini 1.0 family (Nano/Pro/Ultra). Natively multimodal - understands, operates across, and combines text, code, audio, images, and video for superior comprehension and reasoning. Integrated into Bard (now Gemini) in December 2023, delivering single biggest quality improvement to Bard since launch. Surpasses GPT-3.5 on 6 of 8 benchmarks including MMLU and GSM8K. Available to developers via Gemini API in Google AI Studio and Google Cloud Vertex AI. Superseded by Gemini 1.5 Pro (Feb 2024) and Gemini 2.0 (Dec 2024).
Strengths
- Natively multimodal - text, code, audio, images, video in single model
- Surpasses GPT-3.5 on 6 of 8 benchmarks (MMLU, GSM8K)
- Delivered biggest quality improvement to Bard/Gemini at launch
- Native understanding across modalities enables superior reasoning
- Available via Gemini API and Google Cloud Vertex AI
- Represents Google's first production multimodal model
- Positioned for wide range of tasks vs specialized models
Caveats
- Superseded by Gemini 1.5 Pro (Feb 2024) and 2.0 (Dec 2024)
- 32K context window small vs 1.5 Pro (2M) and 2.0 Flash (1M)
- 2K max output tokens extremely limited
- Consider migrating to Gemini 1.5 Pro or 2.0 Flash for new projects
- First-generation model with less refinement than successors
- Vendor
- Context Window
- 32,000 tokens
- Max Output
- 2,048 tokens
- License
- Proprietary
- Release Date
- 2023-12-06
- Modalities
- textimageaudiovideo
Capabilities
Vision
Audio
Video
Tool Use
Resources
Pricing
Superseded by Gemini 1.5 Pro and 2.0. Pricing varies by platform