Skip to main content
llm.info

LLM Models

Explore and compare the latest large language models from leading AI companies.

Promoted

179 LLM Models found

Granite 3.3 8B Instruct

IBM

Open Weights

IBM's open 8B enterprise language model (2025), trained for general language, math, coding, and instruction following with a 128K context window. Released under Apache 2.0, the Granite family targets governed, on-prem enterprise deployments with strong tool use and RAG.

Context: 131,072 tokens
Tools

Jamba 1.5 Large

AI21 Labs

Open Weights

AI21's open hybrid Mamba-Transformer (SSM) MoE model (August 2024), with 398B total parameters and 94B active. Its hybrid architecture delivers up to 2.5x faster long-context inference than comparable models, with a 256K context window suited to agentic and document-heavy workloads.

Context: 256,000 tokens
Tools

Hunyuan-Large

Tencent

Open Weights

Tencent's open-source Transformer MoE (November 2024), among the largest open MoE models at 389B total parameters with 52B active, and a 256K context window. It matches or exceeds Llama 3.1 on many benchmarks and is released for research and commercial use.

Context: 256,000 tokens
Tools

ERNIE 4.5 300B

Baidu

Open Weights

Baidu's open-weight flagship text MoE (July 2025), part of the ERNIE 4.5 family open-sourced under Apache 2.0. The 300B-A47B model activates 47B parameters and delivers strong reasoning and Chinese/English performance, alongside multimodal and smaller variants from 0.3B to 424B.

Context: 131,072 tokens
Tools

Hermes 3 405B

Nous Research

Open Weights

Nous Research's flagship open model (August 2024), the first full-parameter fine-tune of Llama 3.1 405B. Hermes 3 is a steerable generalist with strong agentic capabilities, roleplay, reasoning, and long-context coherence, competitive with or exceeding Llama 3.1 Instruct.

Context: 131,072 tokens
Tools

Jamba 1.5 Mini

AI21 Labs

Open Weights

AI21's small hybrid Mamba-Transformer (SSM) model (August 2024), with 52B total parameters and 12B active, delivering long-context efficiency at small scale. It offers a 256K context window (up to ~140K on a single GPU) at very low cost, ideal for document-heavy and agentic workloads.

Context: 256,000 tokens
Tools

Pixtral 12B

Mistral AI

Open Weights

Mistral's first multimodal model (September 2024), a 12B open-weight model with a dedicated vision encoder that understands images alongside text at native resolution. Released under Apache 2.0, it brought strong open multimodal performance to consumer-scale hardware.

Context: 128,000 tokens
Vision
Tools

Mistral NeMo

Mistral AI

Open Weights

Mistral and NVIDIA's jointly released 12B open model (July 2024), a drop-in upgrade to Mistral 7B with a 128K context window. Trained on multilingual and code data with a new Tekken tokenizer, it ships under Apache 2.0 in base and instruct versions plus an FP8 variant.

Context: 128,000 tokens
Tools

Aya Expanse 32B

Cohere

Open Weights

Cohere Labs' open multilingual model (December 2024) mastering 101 languages through data arbitrage, preference tuning, and model merging. Aya Expanse 32B outperforms Gemma 2, Qwen 2.5, and Llama 3.1 in its class on multilingual Arena-Hard-Auto, advancing the multilingual open frontier.

Context: 128,000 tokens
Tools

Command R

Cohere

Open Weights

Cohere's open-weight 35B model for enterprise retrieval-augmented generation and tool use, with a 128K context window and strong multilingual support. Command R pairs with Cohere Rerank for high-precision RAG at a lower cost than the larger Command R+.

Context: 128,000 tokens
Tools

Command A

Cohere

Open Weights

Cohere's enterprise flagship (March 2025), a 111B open-weight model with a 256K context window tuned for agentic tasks, RAG, tool use, and multilingual business workflows. It runs on just two GPUs and is ~150% more efficient at inference than Command R+.

Context: 256,000 tokens
Tools

Command R7B

Cohere

Open Weights

Cohere's smallest and fastest enterprise model (December 2024), a 7B open-weight model optimized for retrieval-augmented generation, tool use, and agentic workflows. It runs on commodity and edge hardware with a 128K context window and multilingual support across 23 languages.

Context: 131,072 tokens
Tools