DeepSeek-V4-Pro
Open Weights
DeepSeek
DeepSeek's frontier open-weights MoE model (April 2026): 1.6T total / 49B active parameters with a 1M-token context window and a hybrid attention mechanism (Compressed Sparse Attention + Heavily Compressed Attention) for efficient long context. MIT-licensed and competitive with top closed models on math, STEM, and coding.
Strengths
- Rivals top closed models on Math/STEM/Coding while fully open-weight
- 1M-token context with highly efficient hybrid attention (CSA + HCA)
- Permissive MIT license - free for commercial self-hosting
- Far lower long-context inference cost than DeepSeek-V3.2
Caveats
- 1.6T total parameters require a large multi-GPU deployment to self-host
- Preview release - behavior may change before stable
- Vendor
- DeepSeek
- Context Window
- 1,000,000 tokens
- License
- MIT
- Release Date
- 2026-04-24
- Modalities
- textimage
Capabilities
Vision
Audio
Video
Tool Use
Resources
Pricing
Open weights under the MIT License - free to self-host; managed API/pricing via DeepSeek and other providers.