Skip to main content
llm.info

LLM Models

Explore and compare the latest large language models from leading AI companies.

Promoted

179 LLM Models found

Ministral 8B

Mistral AI

Open Weights

Mistral's edge-focused 8B model (October 2024), part of les Ministraux, built for on-device and low-latency inference with a 128K context window. It offers strong reasoning and instruction following for its size, suited to local assistants, function calling, and privacy-sensitive deployments.

Context: 131,072 tokens
Tools

Mistral Small 3.1

Mistral AI

Open Weights

Mistral's open-weight 24B multimodal model (March 2025, Apache 2.0), adding image understanding and a 128K context window to Mistral Small 3. It delivers strong text and vision performance at a size that runs on a single high-end GPU, competitive with larger proprietary small models.

Context: 131,072 tokens
Vision
Tools

Mistral Large 2

Mistral AI

Mistral AI's flagship model with 123B parameters designed for single-node inference at large throughput. Released July 2024 with breakthrough code generation - achieves state-of-the-art 92% on HumanEval (matched only by Claude 3.5 Sonnet). Beats many closed models on MATH problem-solving (71.5%, surpassing Gemini 1.5 Pro, Gemini 1.0 Ultra, GPT-4, Claude 3 Opus). Supports dozens of languages and 80+ coding languages. Features 128K context window for long-context applications. Performs on par with GPT-4o, Claude 3 Opus, and Llama 3 405B while enabling efficient single-node deployment.

Context: 128,000 tokens
Tools

Mistral Small

Mistral AI

Mistral AI's optimized model balancing performance and cost-effectiveness. Designed for latency-sensitive applications requiring reliable performance without flagship pricing. Outperforms Mixtral 8x7B while offering lower latency. Features 32K context window (some sources indicate 131K support) and 4K max output. Supports function calling and JSON formatting for structured outputs. Excellent for data analysis, research tasks, and everyday AI applications. Positioned as refined intermediary between open-weights and flagship tiers.

Context: 32,768 tokens
Tools

DeepSeek-Coder-V2

DeepSeek

Open Weights

DeepSeek's open-source code MoE (June 2024), further pretrained from DeepSeek-V2 on an additional 6T tokens. It expands programming-language support from 86 to 338 and context from 16K to 128K, reaching performance comparable to GPT-4 Turbo on code-specific tasks. Available in 236B (21B active) and 16B (2.4B active) sizes.

Context: 128,000 tokens
Tools

DeepSeek-V3.2

DeepSeek

Open Weights

DeepSeek's December 2025 open-weight successor to V3.2-Exp, a reasoning-first model built for long-context tasks, agent workflows, and tool use. It introduces DeepSeek Sparse Attention (DSA) to cut long-context inference cost while keeping output quality on par with V3.1, with a 131K context window and 64K max output.

Context: 131,072 tokens
Tools

DeepSeek-V3.1

DeepSeek

Open Weights

DeepSeek's August 2025 hybrid model, a 685B-parameter mixture-of-experts (37B active) that unifies chat, reasoning, and coding in a single checkpoint with switchable thinking. Open-weight and inexpensive to serve, it offers a 128K context window at a fraction of proprietary frontier pricing.

Context: 128,000 tokens
Tools

DeepSeek-V3

DeepSeek

Open Weights

DeepSeek's groundbreaking 671B parameter Mixture-of-Experts model with 37B activated per token. Released December 2024 with MIT license enabling unrestricted commercial use. Outperforms open-source models and achieves performance comparable to leading closed-source models (GPT-4, Claude) on most benchmarks. Excels particularly on math and code tasks. Uses Multi-head Latent Attention (MLA) and DeepSeekMoE architectures for efficient inference. Pretrained on 14.8T diverse tokens with only 2.788M H800 GPU hours - breakthrough training efficiency. Pioneers auxiliary-loss-free load balancing and multi-token prediction objectives.

Context: 128,000 tokens
Tools

Qwen 2.5 Coder 32B

Alibaba Cloud

Open Weights

Alibaba's state-of-the-art open-source code model achieving GPT-4o-level coding performance. Released November 2024 with 32B parameters - small enough to run on consumer hardware (32GB+ RAM). Best open-source model on multiple code generation benchmarks (EvalPlus, LiveCodeBench, BigCodeBench). Scores 73.7 on Aider code repair (comparable to GPT-4o). Excels across 40+ programming languages with 65.9 on McEval multi-language benchmark. Achieves 75.2 on MdEval code repair (first among open-source). Available in sizes from 0.5B to 32B. Represents breakthrough in accessible, GPT-4o-class coding assistance.

Context: 128,000 tokens
Tools

Grok-2

xAI

xAI's state-of-the-art AI assistant with advanced text and vision understanding, integrating real-time information from the 𝕏 (Twitter) platform. Released August 2024 for 𝕏 Premium and Premium+ users. Achieves competitive performance with frontier models on graduate-level science (GPQA), general knowledge (MMLU, MMLU-Pro), and math (MATH). Excels at vision tasks - reaches 93.6% on DocVQA document Q&A (outperforming GPT-4 Turbo 87.2% and Claude 3.5 Sonnet 89.3%). Introduced as 'sus-column-r' on LMArena, outperforming Claude and GPT-4 on LMSYS leaderboard. Includes Grok-2 mini variant balancing speed and quality.

Context: 128,000 tokens
Vision
Tools

Llama 3.2 11B Vision

Meta

Open Weights

Meta's lightweight multimodal model balancing vision capabilities with efficiency. Released September 2024 alongside 90B variant as first Llama vision models. Features 11B parameters with 128K context window. Built on Llama 3.1 with vision adapter enabling image understanding up to 1120x1120 resolution. Optimized for visual recognition, image reasoning, captioning, document understanding, and visual grounding. More accessible hardware requirements than 90B while maintaining strong vision capabilities. Supports grouped-query attention for enhanced inference speed. Ideal for edge deployment and resource-constrained multimodal applications.

Context: 128,000 tokens
Vision
Tools

Llama 3.2 90B Vision

Meta

Open Weights

Meta's first large-scale open-weights vision model with 90B parameters. Released September 2024 as part of Llama 3.2 family. Built on Llama 3.1 with trained vision adapter using cross-attention layers integrating pre-trained image encoder. Processes high-resolution images up to 1120x1120 pixels. Features 128K context window for extensive multimodal conversations. Excels at visual recognition, image reasoning, captioning, document understanding (charts, graphs), and visual grounding. Outperforms many closed models like Claude 3 Haiku on image understanding tasks. Supports 8 languages for text, English-only for vision tasks.

Context: 128,000 tokens
Vision
Tools