Skip to main content
llm.info

Llama 3.2 11B Vision

Open Weights

Meta

Meta's lightweight multimodal model balancing vision capabilities with efficiency. Released September 2024 alongside 90B variant as first Llama vision models. Features 11B parameters with 128K context window. Built on Llama 3.1 with vision adapter enabling image understanding up to 1120x1120 resolution. Optimized for visual recognition, image reasoning, captioning, document understanding, and visual grounding. More accessible hardware requirements than 90B while maintaining strong vision capabilities. Supports grouped-query attention for enhanced inference speed. Ideal for edge deployment and resource-constrained multimodal applications.

Strengths

  • Lightweight multimodal model - balances capability with efficiency
  • Processes high-resolution images up to 1120x1120 pixels
  • 128K context window for extended conversations with images
  • More accessible hardware requirements than 90B variant
  • Grouped-query attention enhances inference speed and efficiency
  • Suitable for edge deployment and resource-constrained environments
  • Permissive license enables commercial multimodal applications

Caveats

  • Lower vision capabilities than 90B variant for complex visual reasoning
  • Vision tasks require English prompts only (text supports 8 languages)
  • 4K max output tokens standard for the series
  • May struggle with highly complex or detailed visual analysis
  • Vision adapter adds complexity vs text-only models
Vendor
Meta
Context Window
128,000 tokens
Max Output
4,096 tokens
License
Llama 3.2 Community License
Release Date
2024-09-25
Modalities
text
image

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open weights - free to download. More accessible GPU requirements. API providers typically $0.10-$0.40 per 1M input tokens

Reviews

Comments