Skip to main content
llm.info

Falcon 180B

Open Weights

Technology Innovation Institute

Technology Innovation Institute's massive 180B parameter open-access model trained on 3.5T tokens. Causal decoder-only architecture with 80 layers, hidden dimension 14,848, vocabulary size 65,024. Trained on up to 4,096 A100 GPUs using Amazon SageMaker for ~7M GPU hours. Dataset consists of 85% RefinedWeb plus curated conversations, technical papers, and code (~3%). Achieved 68.74 on Hugging Face Open LLM Leaderboard - highest among open models at release. Surpassed Meta's LLaMA 2 and ranks near GPT-4 and PaLM 2. 2.5x larger than Llama 2 with 4x more compute. Released under Falcon 180B TII License (Apache 2.0 based) for research and commercial use.

Strengths

  • Highest Open LLM Leaderboard score (68.74) among open models at release
  • Massive 180B parameters - 2.5x larger than Llama 2 70B
  • Trained on 3.5T tokens with 4x more compute than Llama 2
  • Surpasses Meta's LLaMA 2 and ranks near GPT-4 and PaLM 2
  • Trained on massive scale - up to 4,096 A100 GPUs, 7M GPU hours
  • Apache 2.0-based license enables commercial use with permission
  • Represents UAE's contribution to open-source AI leadership

Caveats

  • Released Sep 2023 - superseded by newer models (Llama 3.1, Qwen, etc.)
  • 2K context window extremely limited vs modern models (128K-2M)
  • Requires massive resources - 640GB RAM (FP16) or 320GB (int4)
  • 180B size makes deployment impractical for most use cases
  • Consider newer 70B models (Llama 3.1, Qwen 2.5) for better efficiency
Vendor
Technology Innovation Institute
Context Window
2,048 tokens
Max Output
2,048 tokens
License
Falcon 180B TII License (Apache 2.0 based)
Release Date
2023-09-06
Modalities
text

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open access (Falcon 180B TII License) - free. Requires 640GB RAM (FP16) or 320GB (int4). API providers typically $1.50-$3.00 per 1M input tokens

Reviews

Comments