Skip to main content
llm.info

DeepSeek-V4-Pro

Open Weights

DeepSeek

DeepSeek's frontier open-weights MoE model (April 2026): 1.6T total / 49B active parameters with a 1M-token context window and a hybrid attention mechanism (Compressed Sparse Attention + Heavily Compressed Attention) for efficient long context. MIT-licensed and competitive with top closed models on math, STEM, and coding.

Strengths

  • Rivals top closed models on Math/STEM/Coding while fully open-weight
  • 1M-token context with highly efficient hybrid attention (CSA + HCA)
  • Permissive MIT license - free for commercial self-hosting
  • Far lower long-context inference cost than DeepSeek-V3.2

Caveats

  • 1.6T total parameters require a large multi-GPU deployment to self-host
  • Preview release - behavior may change before stable
Vendor
DeepSeek
Context Window
1,000,000 tokens
License
MIT
Release Date
2026-04-24
Modalities
text
image

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open weights under the MIT License - free to self-host; managed API/pricing via DeepSeek and other providers.

Reviews

Comments