LLM Models
Explore and compare the latest large language models from leading AI companies.
179 LLM Models found
Ministral 8B
Mistral AI
Mistral's edge-focused 8B model (October 2024), part of les Ministraux, built for on-device and low-latency inference with a 128K context window. It offers strong reasoning and instruction following for its size, suited to local assistants, function calling, and privacy-sensitive deployments.
Mistral Small 3.1
Mistral AI
Mistral's open-weight 24B multimodal model (March 2025, Apache 2.0), adding image understanding and a 128K context window to Mistral Small 3. It delivers strong text and vision performance at a size that runs on a single high-end GPU, competitive with larger proprietary small models.
Mistral Large 2
Mistral AI
Mistral AI's flagship model with 123B parameters designed for single-node inference at large throughput. Released July 2024 with breakthrough code generation - achieves state-of-the-art 92% on HumanEval (matched only by Claude 3.5 Sonnet). Beats many closed models on MATH problem-solving (71.5%, surpassing Gemini 1.5 Pro, Gemini 1.0 Ultra, GPT-4, Claude 3 Opus). Supports dozens of languages and 80+ coding languages. Features 128K context window for long-context applications. Performs on par with GPT-4o, Claude 3 Opus, and Llama 3 405B while enabling efficient single-node deployment.
Mistral Small
Mistral AI
Mistral AI's optimized model balancing performance and cost-effectiveness. Designed for latency-sensitive applications requiring reliable performance without flagship pricing. Outperforms Mixtral 8x7B while offering lower latency. Features 32K context window (some sources indicate 131K support) and 4K max output. Supports function calling and JSON formatting for structured outputs. Excellent for data analysis, research tasks, and everyday AI applications. Positioned as refined intermediary between open-weights and flagship tiers.
DeepSeek-Coder-V2
DeepSeek
DeepSeek's open-source code MoE (June 2024), further pretrained from DeepSeek-V2 on an additional 6T tokens. It expands programming-language support from 86 to 338 and context from 16K to 128K, reaching performance comparable to GPT-4 Turbo on code-specific tasks. Available in 236B (21B active) and 16B (2.4B active) sizes.
DeepSeek-V3.2
DeepSeek
DeepSeek's December 2025 open-weight successor to V3.2-Exp, a reasoning-first model built for long-context tasks, agent workflows, and tool use. It introduces DeepSeek Sparse Attention (DSA) to cut long-context inference cost while keeping output quality on par with V3.1, with a 131K context window and 64K max output.
DeepSeek-V3.1
DeepSeek
DeepSeek's August 2025 hybrid model, a 685B-parameter mixture-of-experts (37B active) that unifies chat, reasoning, and coding in a single checkpoint with switchable thinking. Open-weight and inexpensive to serve, it offers a 128K context window at a fraction of proprietary frontier pricing.
DeepSeek-V3
DeepSeek
DeepSeek's groundbreaking 671B parameter Mixture-of-Experts model with 37B activated per token. Released December 2024 with MIT license enabling unrestricted commercial use. Outperforms open-source models and achieves performance comparable to leading closed-source models (GPT-4, Claude) on most benchmarks. Excels particularly on math and code tasks. Uses Multi-head Latent Attention (MLA) and DeepSeekMoE architectures for efficient inference. Pretrained on 14.8T diverse tokens with only 2.788M H800 GPU hours - breakthrough training efficiency. Pioneers auxiliary-loss-free load balancing and multi-token prediction objectives.
Qwen 2.5 Coder 32B
Alibaba Cloud
Alibaba's state-of-the-art open-source code model achieving GPT-4o-level coding performance. Released November 2024 with 32B parameters - small enough to run on consumer hardware (32GB+ RAM). Best open-source model on multiple code generation benchmarks (EvalPlus, LiveCodeBench, BigCodeBench). Scores 73.7 on Aider code repair (comparable to GPT-4o). Excels across 40+ programming languages with 65.9 on McEval multi-language benchmark. Achieves 75.2 on MdEval code repair (first among open-source). Available in sizes from 0.5B to 32B. Represents breakthrough in accessible, GPT-4o-class coding assistance.
Grok-2
xAI
xAI's state-of-the-art AI assistant with advanced text and vision understanding, integrating real-time information from the 𝕏 (Twitter) platform. Released August 2024 for 𝕏 Premium and Premium+ users. Achieves competitive performance with frontier models on graduate-level science (GPQA), general knowledge (MMLU, MMLU-Pro), and math (MATH). Excels at vision tasks - reaches 93.6% on DocVQA document Q&A (outperforming GPT-4 Turbo 87.2% and Claude 3.5 Sonnet 89.3%). Introduced as 'sus-column-r' on LMArena, outperforming Claude and GPT-4 on LMSYS leaderboard. Includes Grok-2 mini variant balancing speed and quality.
Llama 3.2 11B Vision
Meta
Meta's lightweight multimodal model balancing vision capabilities with efficiency. Released September 2024 alongside 90B variant as first Llama vision models. Features 11B parameters with 128K context window. Built on Llama 3.1 with vision adapter enabling image understanding up to 1120x1120 resolution. Optimized for visual recognition, image reasoning, captioning, document understanding, and visual grounding. More accessible hardware requirements than 90B while maintaining strong vision capabilities. Supports grouped-query attention for enhanced inference speed. Ideal for edge deployment and resource-constrained multimodal applications.
Llama 3.2 90B Vision
Meta
Meta's first large-scale open-weights vision model with 90B parameters. Released September 2024 as part of Llama 3.2 family. Built on Llama 3.1 with trained vision adapter using cross-attention layers integrating pre-trained image encoder. Processes high-resolution images up to 1120x1120 pixels. Features 128K context window for extensive multimodal conversations. Excels at visual recognition, image reasoning, captioning, document understanding (charts, graphs), and visual grounding. Outperforms many closed models like Claude 3 Haiku on image understanding tasks. Supports 8 languages for text, English-only for vision tasks.