Falcon 180B
Technology Innovation Institute
Technology Innovation Institute's massive 180B parameter open-access model trained on 3.5T tokens. Causal decoder-only architecture with 80 layers, hidden dimension 14,848, vocabulary size 65,024. Trained on up to 4,096 A100 GPUs using Amazon SageMaker for ~7M GPU hours. Dataset consists of 85% RefinedWeb plus curated conversations, technical papers, and code (~3%). Achieved 68.74 on Hugging Face Open LLM Leaderboard - highest among open models at release. Surpassed Meta's LLaMA 2 and ranks near GPT-4 and PaLM 2. 2.5x larger than Llama 2 with 4x more compute. Released under Falcon 180B TII License (Apache 2.0 based) for research and commercial use.
Strengths
- Highest Open LLM Leaderboard score (68.74) among open models at release
- Massive 180B parameters - 2.5x larger than Llama 2 70B
- Trained on 3.5T tokens with 4x more compute than Llama 2
- Surpasses Meta's LLaMA 2 and ranks near GPT-4 and PaLM 2
- Trained on massive scale - up to 4,096 A100 GPUs, 7M GPU hours
- Apache 2.0-based license enables commercial use with permission
- Represents UAE's contribution to open-source AI leadership
Caveats
- Released Sep 2023 - superseded by newer models (Llama 3.1, Qwen, etc.)
- 2K context window extremely limited vs modern models (128K-2M)
- Requires massive resources - 640GB RAM (FP16) or 320GB (int4)
- 180B size makes deployment impractical for most use cases
- Consider newer 70B models (Llama 3.1, Qwen 2.5) for better efficiency
- Vendor
- Technology Innovation Institute
- Context Window
- 2,048 tokens
- Max Output
- 2,048 tokens
- License
- Falcon 180B TII License (Apache 2.0 based)
- Release Date
- 2023-09-06
- Modalities
- text
| Benchmark | Score |
|---|---|
| Hugging Face Open LLM Leaderboard | 68.7 |
Capabilities
Resources
Pricing
Open access (Falcon 180B TII License) - free. Requires 640GB RAM (FP16) or 320GB (int4). API providers typically $1.50-$3.00 per 1M input tokens