Llama 3.2 90B Vision
Meta
Meta's first large-scale open-weights vision model with 90B parameters. Released September 2024 as part of Llama 3.2 family. Built on Llama 3.1 with trained vision adapter using cross-attention layers integrating pre-trained image encoder. Processes high-resolution images up to 1120x1120 pixels. Features 128K context window for extensive multimodal conversations. Excels at visual recognition, image reasoning, captioning, document understanding (charts, graphs), and visual grounding. Outperforms many closed models like Claude 3 Haiku on image understanding tasks. Supports 8 languages for text, English-only for vision tasks.
Strengths
- First large-scale open-weights vision model from Meta
- Processes high-resolution images up to 1120x1120 pixels
- Outperforms closed models like Claude 3 Haiku on image understanding
- 128K context window for extended multimodal conversations
- Document-level understanding - excels at charts, graphs, complex diagrams
- Permissive license enables commercial vision applications
- Visual grounding and image captioning capabilities
Caveats
- Vision tasks require English prompts only (text supports 8 languages)
- Significant GPU memory requirements for vision processing
- 4K max output tokens lower than text-only flagship models
- More complex deployment than text-only models
- API pricing varies and tends to be higher for vision capabilities
- Vendor
- Meta
- Context Window
- 128,000 tokens
- Max Output
- 4,096 tokens
- License
- Llama 3.2 Community License
- Release Date
- 2024-09-25
- Modalities
- textimage
Capabilities
Resources
Pricing
Open weights - free to download. Requires significant GPU resources for vision. API providers typically $0.60-$1.20 per 1M input tokens