Skip to main content
llm.info

Llama 3.2 90B Vision

Open Weights

Meta

Meta's first large-scale open-weights vision model with 90B parameters. Released September 2024 as part of Llama 3.2 family. Built on Llama 3.1 with trained vision adapter using cross-attention layers integrating pre-trained image encoder. Processes high-resolution images up to 1120x1120 pixels. Features 128K context window for extensive multimodal conversations. Excels at visual recognition, image reasoning, captioning, document understanding (charts, graphs), and visual grounding. Outperforms many closed models like Claude 3 Haiku on image understanding tasks. Supports 8 languages for text, English-only for vision tasks.

Strengths

  • First large-scale open-weights vision model from Meta
  • Processes high-resolution images up to 1120x1120 pixels
  • Outperforms closed models like Claude 3 Haiku on image understanding
  • 128K context window for extended multimodal conversations
  • Document-level understanding - excels at charts, graphs, complex diagrams
  • Permissive license enables commercial vision applications
  • Visual grounding and image captioning capabilities

Caveats

  • Vision tasks require English prompts only (text supports 8 languages)
  • Significant GPU memory requirements for vision processing
  • 4K max output tokens lower than text-only flagship models
  • More complex deployment than text-only models
  • API pricing varies and tends to be higher for vision capabilities
Vendor
Meta
Context Window
128,000 tokens
Max Output
4,096 tokens
License
Llama 3.2 Community License
Release Date
2024-09-25
Modalities
text
image

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open weights - free to download. Requires significant GPU resources for vision. API providers typically $0.60-$1.20 per 1M input tokens

Reviews

Comments