Skip to main content
llm.info

LLM Models

Explore and compare the latest large language models from leading AI companies.

Promoted

179 LLM Models found

Llama 3.3 70B

Meta

Open Weights

Meta's efficiency-optimized model delivering Llama 3.1 405B-level quality at fraction of the cost. Released December 2024 with breakthrough cost-performance ratio - generates responses nearly 5x more cost-efficiently than 405B while maintaining similar output quality. Features 128K context window and achieves 86% on MMLU Chat, matching 3.1 70B. Excels at math (77% MATH, up from 67.8%), multilingual reasoning (91.1% MGSM vs 86.9%), and instruction following (92.1% IFEval). Pretrained on 15T tokens using 39.3M H100 GPU hours. Runs on accessible hardware (2-4 A100s).

Context: 128,000 tokens
Tools

Llama 3.1 8B

Meta

Open Weights

Meta's smallest and most efficient Llama 3.1 model, optimized for edge deployment and resource-constrained environments. Features 128K context window despite compact size. Can run on consumer GPUs (single RTX 4090 or similar) and even high-end consumer CPUs with quantization. Maintains surprisingly strong performance for its size while enabling local deployment, privacy-focused applications, and low-cost API hosting. Ideal for applications requiring on-device AI or minimal latency.

Context: 128,000 tokens
Tools

Llama 3.1 70B

Meta

Open Weights

Meta's mid-size open-weights model offering excellent performance-to-cost ratio. Balances capability and efficiency, making it practical for self-hosting on more accessible hardware (2-4 A100 GPUs). Features 128K context window and 8-language support. Optimized for multilingual dialogue and outperforms many available open-source models on industry benchmarks. Significantly more accessible than 405B while maintaining strong quality for most tasks.

Context: 128,000 tokens
Tools

Llama 3.1 405B

Meta

Open Weights

Meta's largest and most capable open-weights model, rivaling top proprietary models. First openly available model competitive with GPT-4, Claude 3 Opus, and Gemini 1.5 Pro. Trained on over 16 trillion tokens using 16,000+ H100 GPUs. Features 128K context window (16x larger than Llama 3) and supports 8 languages. 87.3% on general knowledge benchmark exceeds GPT-4 Turbo (86.5%) and Claude 3 Opus (86.8%). Enables commercial use with permissive license, allowing self-hosting and fine-tuning.

Context: 128,000 tokens
Tools

Gemini 2.0 Flash

Google

Google's next-generation model designed for the agentic era with superior speed and enhanced capabilities. Released December 2024, generally available February 2025. Outperforms Gemini 1.5 Pro on key benchmarks at 2x the speed. Features 1M token context window, multimodal inputs (text, image, audio, video), and groundbreaking multimodal output including natively generated images mixed with text and steerable text-to-speech in multiple languages. Built-in tool use, function calling, and advanced reasoning. Beats 1.5 Flash on MMLU Pro (77.6% vs 67.3%) and Bird SQL (57.4% vs 45.6%).

Context: 1,000,000 tokens
Vision
Audio
Tools

Gemini 1.5 Flash

Google

Google's fast, cost-efficient model created through knowledge distillation from Gemini 1.5 Pro. Designed for high-volume, latency-sensitive applications while maintaining strong capabilities. Features 1 million token context window and multimodal support (text, image, audio, video). Excels at summarization, chat, image/video captioning, and data extraction from long documents. Massive price reductions throughout 2024 (78% input, 71% output) make it extremely cost-effective. Improved performance: 7% MMLU-Pro gain, 20% math boost.

Context: 1,000,000 tokens
Vision
Audio
Tools

Gemini 1.5 Pro

Google

Google's flagship model with industry-leading 2 million token context window - the longest available. Sparse mixture-of-experts (MoE) Transformer architecture enables processing massive amounts of data: 2 hours of video, 19 hours of audio, codebases with 60,000 lines, or 2,000 pages of text. Matches or outperforms Gemini 1.0 Ultra on standard benchmarks while using significantly less training compute. Near-perfect retrieval (>99%) up to 10 million tokens in research tests. Major price reductions in October 2024.

Context: 2,000,000 tokens
Vision
Audio
Tools

Gemma 3 27B

Google

Open Weights

Google's flagship open-weights model in the Gemma 3 family (March 2025), with vision support, a 128K-token context window, and multilingual coverage. Runs efficiently on a single GPU/TPU and is free to self-host under the Gemma license.

Context: 128,000 tokens
Vision
Tools

Gemini 2.5 Flash-Lite

Google

Google's fastest and most budget-friendly multimodal model of 2025, optimized for cost-sensitive, high-throughput workloads with full multimodality and a large context window.

Context: 1,000,000 tokens
Vision
Audio
Video
Tools

Gemini 2.5 Flash

Google

Google's best price-performance model of 2025 for low-latency, high-volume tasks, with a 1M-token context window and full multimodality. A workhorse for production applications.

Context: 1,000,000 tokens
Vision
Audio
Video
Tools

Gemini 2.5 Pro

Google

Google's 2025 flagship reasoning model with a 1M-token context window and strong performance on complex reasoning, coding, and STEM tasks. Fully multimodal (text, image, audio, video). Superseded by the Gemini 3.x generation.

Context: 1,000,000 tokens
Vision
Audio
Video
Tools

Gemini 3.1 Flash-Lite

Google

Google's most budget-friendly Gemini 3.1 model (2026), delivering frontier-class capabilities at very low cost for high-volume, latency-sensitive workloads. Fully multimodal with a large context window.

Context: 1,000,000 tokens
Vision
Audio
Video
Tools