Skip to main content
llm.info

LLM Models

Explore and compare the latest large language models from leading AI companies.

Promoted

179 LLM Models found

Command R+

Cohere

Cohere's enterprise-focused flagship model with 104B parameters and 128K context window. Released April 2024, updated August 2024. Optimized for advanced RAG with industry-leading grounded generation capabilities providing in-line citations to mitigate hallucinations. Trained with specific grounding capabilities - generates responses based on supplied document snippets including citation spans indicating information source. Supports 10 key business languages (trained on 23, evaluated on 10). Powerful Tool Use API for automating complex workflows. Beats GPT-4 Turbo on enterprise benchmarks. Available on Azure, OCI, and Cohere API. Pricing: $3 per 1M input, $15 per 1M output.

Context: 128,000 tokens
Tools

CodeLlama 34B

Meta

Open Weights

Meta's specialized code generation model with 34B parameters built on Llama 2 architecture. Released August 2023, trained on 500 billion tokens of code and code-related data from January-July 2023. Features extended 100,000 token context window for handling large codebases. Supports Python, C++, Java, PHP, TypeScript/JavaScript, C#, and Bash. Three variants: Base (code completion), Python (Python-specialized), Instruct (instruction-following). Achieves 53.7% HumanEval and 56.2% MBPP. Trained on Meta's Research Super Cluster requiring 400K GPU hours on A100-80GB hardware across all 9 CodeLlama models. Open weights enable customization and self-hosting.

Context: 100,000 tokens

Llama 2 70B

Meta

Open Weights

Meta's largest open-weights Llama 2 model with 70B parameters. Released July 2023, trained on 2 trillion tokens with 40% more data than Llama 1. Features 4K context length (expanded from 2K in Llama 1) and Grouped-Query Attention for improved inference scalability. Supports both research and commercial use with permissive license. Pretrained on data through September 2022 with tuning data through July 2023. Global batch size of 4M tokens during training. Fine-tuned chat variant optimized for dialogue using supervised fine-tuning and RLHF. Widely adopted as foundation for custom models and applications. Superseded by Llama 3.1 70B (July 2024).

Context: 4,096 tokens

Mistral 7B

Mistral AI

Open Weights

Mistral AI's compact but powerful 7.3B parameter model released September 2023. First model from Mistral AI, setting new standards for efficiency and performance in its size class. Outperforms Llama 2 13B on all benchmarks and Llama 1 34B on many benchmarks despite being significantly smaller. Approaches CodeLlama 7B code performance while maintaining strong English capabilities. Features Grouped-Query Attention for faster inference and Sliding Window Attention (4096 tokens) for efficient long sequences. Apache 2.0 license. Deployable anywhere including locally, cloud (AWS/GCP/Azure), available on HuggingFace.

Context: 8,192 tokens
Tools

Mixtral 8x7B

Mistral AI

Open Weights

Mistral AI's groundbreaking Sparse Mixture of Experts model with 46.7B total parameters. Released December 2023, uses 13B active parameters per token while achieving performance of 6x larger models. Outperforms Llama 2 70B on most benchmarks with 6x faster inference. Matches or exceeds GPT-3.5 across all evaluated benchmarks. Features 32K context window with sliding window attention enabling theoretical 128K token span. Best open-weights chatbot model as of December 2023 per MT-Bench. Supports 5 languages (English, French, German, Spanish, Italian). Apache 2.0 license enables unrestricted commercial use.

Context: 32,000 tokens
Tools

Claude 3 Sonnet

Anthropic

Anthropic's balanced model striking ideal equilibrium between intelligence and speed for enterprise workloads. Released March 2024 alongside Opus and Haiku in Claude 3 family. Delivers 2x faster performance than Claude 2/2.1 with higher intelligence levels. Excels at tasks demanding rapid responses like knowledge retrieval or sales automation. Strong performance at lower cost compared to peers, engineered for high endurance in large-scale AI deployments. Features 200K context window and sophisticated vision capabilities processing photos, charts, graphs, technical diagrams. Positioned between Haiku (speed) and Opus (capability).

Context: 200,000 tokens
Vision
Tools

GPT-3.5 Turbo

OpenAI

OpenAI's widely-adopted cost-effective model released March 2023. Optimized for chat via Chat Completions API with natural language and code generation capabilities. Features 16K context window (20 pages of text). Significantly cheaper than GPT-4 family while maintaining solid performance for most common use cases. Priced at $0.50 per 1M input tokens and $1.50 per 1M output (60% cheaper than previous pricing). Enhanced with better steerability, improved function calling, and extended context. Widely deployed in production applications globally. GPT-4o-mini now recommended upgrade for better performance at similar cost.

Context: 16,385 tokens
Tools

Stable Diffusion XL

Stability AI

Open Weights

Stability AI's flagship open-source text-to-image generation model. Features 3.5B parameter base model with 6.6B parameter refiner in ensemble pipeline. Native 1024x1024 resolution (2x larger than SD 1.5) with improved generation for limbs, text, faces, and overall image quality. Uses dual CLIP networks (CLIP1 + CLIP2) for superior semantic understanding vs single CLIP. Achieves 89% prompt adherence vs SD 1.5's 71%. Supports image-to-image, inpainting, and outpainting workflows. Runs on consumer hardware (RTX 3060+ with 8GB VRAM).

Gemma 2 9B

Google

Open Weights

Google's efficient mid-size open-weights model with 9B parameters. Released June-July 2024, trained on 8 trillion tokens primarily from web documents, code, and mathematics. Delivers class-leading performance, outperforming Llama 3 8B and other open models in its size category. Features Grouped-Query Attention, interleaved sliding window (4096 tokens) and global attention (8192 tokens). Trained with knowledge distillation from larger 27B model for improved quality. 8K context length. Optimized to run at incredible speed on range of hardware - gaming laptops, high-end desktops, cloud setups. Commercially-friendly license. Available via Google AI Studio, Kaggle, Hugging Face.

Context: 8,192 tokens
Tools

Gemma 2 27B

Google

Open Weights

Google's largest open-weights model in Gemma 2 family with 27B parameters. Released June-July 2024, trained on 13 trillion tokens of web documents, code, and mathematics. Features interleaved local/global attention, logit soft-capping for stability, Grouped Query Attention for efficiency. Outperforms Qwen1.5 32B baseline and performs "only a few percent below" Llama 3 70B despite being significantly smaller. Optimized to run full precision on single Google Cloud TPU host, NVIDIA A100 80GB, or H100. 256K vocabulary with 4K context using RoPE. Commercially-friendly license. Available via Google AI Studio, Kaggle, Hugging Face.

Context: 4,096 tokens
Tools

ChatGLM-4

Zhipu AI

Zhipu AI's flagship multilingual model released January 2024 (API) with open-source GLM-4-9B variants in June 2024. Pretrained on 10 trillion tokens primarily Chinese/English plus 24 additional languages. Features 128K context length for both GLM-4 and GLM-4 All Tools. Advanced capabilities: web browsing, code execution, custom tool calls (Function Call), long context reasoning. Approaches state-of-the-art models (GPT-4 Turbo, Gemini 1.5 Pro, Claude 3 Opus) on standard benchmarks. Best-in-class for Chinese language tasks. GLM-4-9B open variant enables research and deployment on consumer hardware.

Context: 128,000 tokens
Tools

Yi-Large

01.AI

01.AI's proprietary flagship model with 70B parameters, released May 2024. Chinese AI powerhouse trained from scratch with significant improvements over Yi-34B open-source model. Performs on par with GPT-4 and Claude 3 across benchmarks. Exceptional multilingual capabilities - excels in Spanish, Chinese, Japanese, German, and French per LMSYS multilingual leaderboard. Positioned first on AlpacaEval 2.0 shortly after launch. Trained on 3.1T tokens of English and Chinese using cascaded deduplication and quality filtering. Uses decoder-only transformer with pre-normalization, SwiGLU, RoPE, and Grouped Query Attention.

Context: 32,000 tokens
Tools