DeepSeek-R1
DeepSeek
DeepSeek's flagship reasoning model released January 2025 with 671B parameters using Mixture of Experts architecture (37B activated per token). Trained via large-scale reinforcement learning without supervised fine-tuning as preliminary step. Demonstrates self-verification, reflection, and long chain-of-thought reasoning capabilities comparable to OpenAI o1 across math, code, and reasoning tasks. Achieves 77.9% on AIME 2024 and 97.3% on MATH-500. Includes distilled models from 1.5B to 70B based on Qwen2.5 and Llama3. Released under MIT license making reasoning capabilities widely accessible. Maximum output length 20K tokens. Represents breakthrough in open-source reasoning models.
Strengths
- Performance on par with OpenAI o1 across math, code, reasoning tasks
- 77.9% on AIME 2024 and 97.3% on MATH-500 demonstrate exceptional math reasoning
- 2,029 Elo on Codeforces-like challenges shows strong coding ability
- Trained via RL without SFT - demonstrates self-verification and reflection
- 20K max output tokens enables detailed chain-of-thought explanations
- MIT license most permissive for open reasoning model
- Distilled variants (1.5B-70B) enable accessible deployment
Caveats
- 671B total parameters require significant GPU resources (37B active)
- Released Jan 2025 - early in reasoning model evolution
- Lower performance on some tasks vs specialized models
- 64K context smaller than GPT-4o (400K) and Gemini (2M)
- Distilled models have reduced capabilities vs full model
- Vendor
- DeepSeek
- Context Window
- 64,000 tokens
- Max Output
- 20,000 tokens
- License
- MIT
- Release Date
- 2025-01-20
- Modalities
- text
| Benchmark | Score | Max |
|---|---|---|
| AIME 2024 | 77.9 | - |
| MATH-500 | 97.3 | - |
| Codeforces | 2029.0 | - |
| livebench_language | 52.1 | 100.0 |
| livebench_coding | 67.4 | 100.0 |
| livebench_instruction_following | 80.6 | 100.0 |
Capabilities
Pricing
Open source (MIT) - free. Requires significant GPU. API providers typically $0.50-$1.50 per 1M input tokens