DeepSeek-V3
DeepSeek
DeepSeek's groundbreaking 671B parameter Mixture-of-Experts model with 37B activated per token. Released December 2024 with MIT license enabling unrestricted commercial use. Outperforms open-source models and achieves performance comparable to leading closed-source models (GPT-4, Claude) on most benchmarks. Excels particularly on math and code tasks. Uses Multi-head Latent Attention (MLA) and DeepSeekMoE architectures for efficient inference. Pretrained on 14.8T diverse tokens with only 2.788M H800 GPU hours - breakthrough training efficiency. Pioneers auxiliary-loss-free load balancing and multi-token prediction objectives.
Strengths
- MIT license enables unrestricted commercial use - most permissive for model of this caliber
- Matches GPT-4 and Claude performance while being fully open
- Exceptional math and code performance - best in class for open models
- Breakthrough training efficiency - only 2.788M H800 GPU hours for 671B model
- MoE architecture activates only 37B of 671B params per token (cost-effective inference)
- Multi-head Latent Attention and DeepSeekMoE for efficient inference
- Pretrained on 14.8T tokens for comprehensive knowledge coverage
Caveats
- 671B total parameters require massive GPU infrastructure for self-hosting
- Text-only (no vision or multimodal capabilities)
- 4K max output tokens lower than newer flagship models
- Relatively new (Dec 2024) - ecosystem still developing
- API availability varies across providers
- Vendor
- DeepSeek
- Context Window
- 128,000 tokens
- Max Output
- 4,096 tokens
- License
- MIT
- Release Date
- 2024-12-25
- Modalities
- text
Capabilities
Pricing
Open source (MIT) - free. Requires significant GPU resources. API providers typically $0.27-$0.60 per 1M input tokens