Skip to main content
llm.info

Swallow 70B

Open Weights

Tokyo Institute of Technology

Tokyo Institute of Technology's Japanese-enhanced Llama model with 70B parameters. Based on Llama 3.3, continually pre-trained on ~200B Japanese tokens from Swallow Corpus v2, Japanese/English Wikipedia, and math/code content. Features expanded vocabulary with Japanese characters and subwords for efficient tokenization and notably faster inference. Evaluated on 10 Japanese benchmarks (JCommonsenseQA, JEMHopQA, NIILC, JSQuAD) and 10 English benchmarks (OpenBookQA, TriviaQA, SQuAD 2.0, XWINO, HellaSwag). Achieves best 70B-class performance for Japanese as of Dec 2023. Also available in Llama 3.1 variants. Represents state-of-the-art Japanese language AI.

Strengths

  • Best 70B-parameter performance for Japanese language (as of Dec 2023)
  • Efficiently represents Japanese with expanded vocabulary - faster inference
  • Trained on 200B curated Japanese tokens from Swallow Corpus v2
  • Maintains strong English capability from Llama 3.3 base
  • 128K context window handles long Japanese documents
  • Evaluated on comprehensive 20-benchmark suite (10 Japanese + 10 English)
  • Open weights enable research and commercial use

Caveats

  • Primarily optimized for Japanese - less competitive on other languages
  • 70B size requires significant GPU memory
  • Training cutoff tied to Llama 3.3 base (mid-2024)
  • No multimodal capabilities (text-only)
  • Consider newer models for multilingual tasks beyond Japanese/English
Vendor
Tokyo Institute of Technology
Context Window
128,000 tokens
Max Output
4,096 tokens
License
Llama 3.3 Community License
Release Date
2024-12-01
Modalities
text

Capabilities

Vision
Audio
Video
Tool Use

Pricing

Open source - free. Requires GPU. API providers typically $0.80-$1.50 per 1M input tokens

Reviews

Comments