LLM Models
Explore and compare the latest large language models from leading AI companies.
179 LLM Models found
Baichuan 2 13B
Baichuan
Baichuan's bilingual Chinese-English model with 13B parameters trained on 2.6T high-quality tokens. Supports both Chinese and English with best-in-class performance for its size across general, legal, medical, mathematics, code, and multilingual domains. Features ALiBi position encoding and 4K context window. Achieves 31.6% faster inference than LLaMA-13B. Evaluated on C-Eval, MMLU, and CMMLU benchmarks. Available in Base and Chat variants plus 4-bit quantized version. Outperforms comparable models on authoritative Chinese, English, and multilingual benchmarks. Free for commercial use with email application. Represents strong Chinese language capabilities at accessible model size.
Swallow 70B
Tokyo Institute of Technology
Tokyo Institute of Technology's Japanese-enhanced Llama model with 70B parameters. Based on Llama 3.3, continually pre-trained on ~200B Japanese tokens from Swallow Corpus v2, Japanese/English Wikipedia, and math/code content. Features expanded vocabulary with Japanese characters and subwords for efficient tokenization and notably faster inference. Evaluated on 10 Japanese benchmarks (JCommonsenseQA, JEMHopQA, NIILC, JSQuAD) and 10 English benchmarks (OpenBookQA, TriviaQA, SQuAD 2.0, XWINO, HellaSwag). Achieves best 70B-class performance for Japanese as of Dec 2023. Also available in Llama 3.1 variants. Represents state-of-the-art Japanese language AI.
Jais 30B
Core42
Core42's bilingual Arabic-English model with 30B parameters setting new benchmarks for Arabic language AI. Trained on 1.6T tokens including 126B Arabic, 251B English, and 50B code tokens. Features 48 transformer blocks, 56 attention heads, embedding dimension 7168. Trained on Condor Galaxy-1 (CG-1) AI supercomputer with 4 exaFLOPS of training compute. Offers 160% longer Arabic answers and 233% longer English answers vs predecessor. Arabic performance now on par with monolingual English models. Outperforms Jais 13B 96% in Arabic and 97% in English. Available on Azure AI and Hugging Face. Represents major advance in Arabic NLP capabilities.
DeepSeek-R1
DeepSeek
DeepSeek's flagship reasoning model released January 2025 with 671B parameters using Mixture of Experts architecture (37B activated per token). Trained via large-scale reinforcement learning without supervised fine-tuning as preliminary step. Demonstrates self-verification, reflection, and long chain-of-thought reasoning capabilities comparable to OpenAI o1 across math, code, and reasoning tasks. Achieves 77.9% on AIME 2024 and 97.3% on MATH-500. Includes distilled models from 1.5B to 70B based on Qwen2.5 and Llama3. Released under MIT license making reasoning capabilities widely accessible. Maximum output length 20K tokens. Represents breakthrough in open-source reasoning models.
Gemini 1.0 Pro
Google's first-generation multimodal model optimized for wide range of tasks. Released December 2023 as part of three-model Gemini 1.0 family (Nano/Pro/Ultra). Natively multimodal - understands, operates across, and combines text, code, audio, images, and video for superior comprehension and reasoning. Integrated into Bard (now Gemini) in December 2023, delivering single biggest quality improvement to Bard since launch. Surpasses GPT-3.5 on 6 of 8 benchmarks including MMLU and GSM8K. Available to developers via Gemini API in Google AI Studio and Google Cloud Vertex AI. Superseded by Gemini 1.5 Pro (Feb 2024) and Gemini 2.0 (Dec 2024).
BLOOM 176B
BigScience
Pioneering open-access multilingual model with 176B parameters designed and built through collaboration of 1,000+ researchers. Released July 2022 as first multilingual LLM trained in complete transparency. Decoder-only Transformer trained on ROOTS corpus covering 46 natural languages and 13 programming languages (59 total). Trained for 117 days (nearly 4 months) on 416 A100 80GB GPUs via Jean Zay supercomputer (French government). Training used 1.6TB multilingual dataset (350B tokens) achieving 150 TFLOPs throughput. Largest publicly available open multilingual model. Released under Responsible AI License. Represents unprecedented international research collaboration in AI.
Phi-4-mini-instruct
Microsoft
The compact 3.8B-parameter member of Microsoft's Phi-4 family (2025, MIT license), an instruction-tuned small model with a 128K-token context window designed for efficient on-device and edge deployment while retaining strong reasoning for its size.
Phi-4-reasoning
Microsoft
Microsoft's open-weight 14B reasoning model (April 2025, MIT license), fine-tuned from Phi-4 for chain-of-thought math and scientific reasoning. It scores 75.3% on AIME 2024 (81.3% for the reasoning-plus variant), beating much larger distilled models, with a 32K context window stable to 64K.
Phi-4
Microsoft
Microsoft's 14.7B-parameter small language model (December 2024), open-sourced on Hugging Face under the MIT license. Trained heavily on synthetic data, Phi-4 matches models 5x its size on math and reasoning, outperforming GPT-4o by ~6 points on GPQA and math while running on consumer hardware.
Phi-3 Medium
Microsoft
Microsoft's efficient Small Language Model (SLM) with 14B parameters. Released May 2024 as part of Phi-3 family alongside Mini (3.8B) and Small (7B). Available in two context lengths: 128K and 4K. Outperforms Gemini 1.0 Pro and GPT-3.5 Turbo despite significantly smaller size. Demonstrates strong reasoning and logic capabilities making it suitable for analytical tasks beyond traditional language tasks. Optimized variants available with ONNX Runtime and DirectML supporting wide range of devices including mobile and web deployments. Available as NVIDIA NIM inference microservices with standard API. Represents Microsoft's 'small but mighty' SLM approach.
Reka Flash 3
Reka AI
Reka AI's open-weight 21B general-purpose reasoning model (March 2025), designed for low-latency chat, coding, and function calling with a 128K context window. Released with publicly available weights under Apache 2.0, it is competitive with larger models for its size.
MiniMax M2
MiniMax
MiniMax's open-weight model built for maximum coding and agentic workflows, a 230B-parameter mixture-of-experts that activates only 10B parameters per token for efficient, sophisticated end-to-end tool use. It serves a 128K-token context window.