Swallow 70B
Tokyo Institute of Technology
Tokyo Institute of Technology's Japanese-enhanced Llama model with 70B parameters. Based on Llama 3.3, continually pre-trained on ~200B Japanese tokens from Swallow Corpus v2, Japanese/English Wikipedia, and math/code content. Features expanded vocabulary with Japanese characters and subwords for efficient tokenization and notably faster inference. Evaluated on 10 Japanese benchmarks (JCommonsenseQA, JEMHopQA, NIILC, JSQuAD) and 10 English benchmarks (OpenBookQA, TriviaQA, SQuAD 2.0, XWINO, HellaSwag). Achieves best 70B-class performance for Japanese as of Dec 2023. Also available in Llama 3.1 variants. Represents state-of-the-art Japanese language AI.
Strengths
- Best 70B-parameter performance for Japanese language (as of Dec 2023)
- Efficiently represents Japanese with expanded vocabulary - faster inference
- Trained on 200B curated Japanese tokens from Swallow Corpus v2
- Maintains strong English capability from Llama 3.3 base
- 128K context window handles long Japanese documents
- Evaluated on comprehensive 20-benchmark suite (10 Japanese + 10 English)
- Open weights enable research and commercial use
Caveats
- Primarily optimized for Japanese - less competitive on other languages
- 70B size requires significant GPU memory
- Training cutoff tied to Llama 3.3 base (mid-2024)
- No multimodal capabilities (text-only)
- Consider newer models for multilingual tasks beyond Japanese/English
- Vendor
- Tokyo Institute of Technology
- Context Window
- 128,000 tokens
- Max Output
- 4,096 tokens
- License
- Llama 3.3 Community License
- Release Date
- 2024-12-01
- Modalities
- text
Capabilities
Resources
Pricing
Open source - free. Requires GPU. API providers typically $0.80-$1.50 per 1M input tokens