LLM Models
Explore and compare the latest large language models from leading AI companies.
179 LLM Models found
Grok 4.3
xAI
xAI's flagship Grok model (April 2026), recommended for chat and coding as its most intelligent and fastest model. Features a 1M-token context window, native video input (an xAI first), and tight integration with real-time search tools.
DeepSeek-V4-Flash
DeepSeek
The efficient tier of DeepSeek's V4 generation (April 2026): a 284B total / 13B active MoE with a 1M-token context window and the same hybrid attention design as V4-Pro. MIT-licensed for low-cost, high-throughput deployments.
DeepSeek-V4-Pro
DeepSeek
DeepSeek's frontier open-weights MoE model (April 2026): 1.6T total / 49B active parameters with a 1M-token context window and a hybrid attention mechanism (Compressed Sparse Attention + Heavily Compressed Attention) for efficient long context. MIT-licensed and competitive with top closed models on math, STEM, and coding.
Llama 4 Scout
Meta
Meta's long-context Llama 4 model (April 2025): a mixture-of-experts design with 17B active parameters across 16 experts (109B total) and an industry-leading 10M-token context window. Natively multimodal and open-weight under the Llama 4 Community License.
Llama 4 Maverick
Meta
Meta's natively-multimodal Llama 4 mid-tier model (April 2025): a mixture-of-experts design with 17B active parameters across 128 experts (400B total) and a 1M-token context window. Open weights under the Llama 4 Community License with early-fusion multimodality.
Vicuna 33B
LMSYS
LMSYS's open-source chatbot fine-tuned from LLaMA with supervised instruction fine-tuning on ~125K user-shared conversations from ShareGPT.com. Auto-regressive language model based on transformer architecture with 33B parameters. Evaluated with standard benchmarks, human preference studies, and LLM-as-a-judge (MT-Bench). Achieved 1091 ELO rating on Chatbot Arena. Scores: 62.1 Arc, 83 HellaSwag, 59.2 MMLU, 56.2 TruthfulQA, 77 WinoGrande, 13.7 GSM8K. Released as part of FastChat platform for training and evaluating large language models. Represents important early work in open-source chat model fine-tuning. Non-commercial license. Part of research leading to Chatbot Arena benchmark.
MPT-30B
MosaicML
MosaicML's (now Databricks) efficient 30B parameter decoder-style transformer pretrained from scratch on 1T tokens of English text and code. Features 8K token context window with support for extrapolation via ALiBi. Trained initially on 1T tokens using 2K sequences then additional 50B tokens of 8K sequences. First LLM trained on NVIDIA H100 GPUs. Outperforms originally published GPT-3 using ~1/6th the parameters. Competitive with LLaMa-30B and Falcon-40B on open-source benchmarks. Surpasses purpose-built models like StarCoder on HumanEval. Designed for easy deployment on single GPU - 1xA100-80GB (16-bit) or 1xA100-40GB (8-bit). Features FlashAttention for efficient inference. Released under Apache 2.0. Available in base, instruct, and chat variants.
Stable Audio 2.5
Stability AI
Stability AI's enterprise audio model (September 2025) for high-quality sound and music production at scale. A generation breakthrough cuts inference from 50 steps to 8, producing three-minute tracks in seconds, and it supports text-to-audio, audio-to-audio, and audio inpainting on user files.
Cartesia Sonic 3
Cartesia
Cartesia's real-time text-to-speech model (late 2025), built on a state-space architecture for extremely low latency (~90ms model latency) with strong emotional expressiveness, multilingual support, and instant voice cloning from short samples - ideal for live voice agents.
Deepgram Nova-3
Deepgram
Deepgram's speech-to-text model (January 2025), the replacement for Nova-2, delivering large word-error-rate reductions (54.3% streaming, 47.4% batch) versus competitors. Nova-3 is the first voice AI model with real-time multilingual transcription and self-serve keyterm customization without retraining.
ElevenLabs v3
ElevenLabs
ElevenLabs' most expressive text-to-speech model (GA February 2026), producing natural, lifelike speech with high emotional range and contextual understanding across many languages. Eleven v3 adds a Text to Dialogue API for multi-speaker, emotionally rich conversations.
Udio
Udio
Udio's AI music generation model, producing high-fidelity songs with vocals from text prompts. Udio v1.5 brought 48kHz stereo output, key guidance, and global language support (including Mandarin), with later updates improving stem separation and vocal pitch control.