Skip to main content
llm.info

DeepSeek-V3

Open Weights

DeepSeek

DeepSeek's groundbreaking 671B parameter Mixture-of-Experts model with 37B activated per token. Released December 2024 with MIT license enabling unrestricted commercial use. Outperforms open-source models and achieves performance comparable to leading closed-source models (GPT-4, Claude) on most benchmarks. Excels particularly on math and code tasks. Uses Multi-head Latent Attention (MLA) and DeepSeekMoE architectures for efficient inference. Pretrained on 14.8T diverse tokens with only 2.788M H800 GPU hours - breakthrough training efficiency. Pioneers auxiliary-loss-free load balancing and multi-token prediction objectives.

Strengths

  • MIT license enables unrestricted commercial use - most permissive for model of this caliber
  • Matches GPT-4 and Claude performance while being fully open
  • Exceptional math and code performance - best in class for open models
  • Breakthrough training efficiency - only 2.788M H800 GPU hours for 671B model
  • MoE architecture activates only 37B of 671B params per token (cost-effective inference)
  • Multi-head Latent Attention and DeepSeekMoE for efficient inference
  • Pretrained on 14.8T tokens for comprehensive knowledge coverage

Caveats

  • 671B total parameters require massive GPU infrastructure for self-hosting
  • Text-only (no vision or multimodal capabilities)
  • 4K max output tokens lower than newer flagship models
  • Relatively new (Dec 2024) - ecosystem still developing
  • API availability varies across providers
Vendor
DeepSeek
Context Window
128,000 tokens
Max Output
4,096 tokens
License
MIT
Release Date
2024-12-25
Modalities
text

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open source (MIT) - free. Requires significant GPU resources. API providers typically $0.27-$0.60 per 1M input tokens

Reviews

Comments