Skip to main content
llm.info

LLM Models

Explore and compare the latest large language models from leading AI companies.

Promoted

179 LLM Models found

GLM-4.6

Zhipu AI

Open Weights

Zhipu AI's open-weight coding and agentic model (September 2025), a 357B-parameter mixture-of-experts with 32B active parameters and a 200K-token context window. GLM-4.6 improves real-world coding, long-context processing, reasoning, search, and agentic use over GLM-4.5.

Context: 200,000 tokens
Tools

Kimi K2

Moonshot AI

Open Weights

Moonshot AI's open-weight flagship (July 2025), a 1-trillion-parameter mixture-of-experts that activates 32B parameters per token, built for agentic coding and tool use. Released under a Modified MIT license, Kimi K2 is one of the strongest open models for autonomous, multi-step workflows.

Context: 131,072 tokens
Tools

NVIDIA Nemotron-3 Super 120B

NVIDIA

Open Weights

NVIDIA's open Nemotron-3 model, a hybrid Mamba-Transformer latent mixture-of-experts with 12B active parameters and a 1M-token context window. Part of NVIDIA's fully open Nemotron family (open weights, training data, and recipes), it targets efficient long-context reasoning.

Context: 1,000,000 tokens
Tools

DBRX

Databricks

Open Weights

Databricks' open MoE model (March 2024), a state-of-the-art open LLM at release with 132B total parameters and 36B active per input via a fine-grained mixture-of-experts. Pretrained on 12T tokens of text and code, DBRX serves a 32K context window under the Databricks Open Model License.

Context: 32,768 tokens
Tools

Amazon Nova Premier

Amazon

Amazon's most capable multimodal model (October 2025) on AWS Bedrock, built for complex reasoning tasks and as a teacher model for distilling custom Nova models. Nova Premier processes text, image, and video with a context window up to 1M tokens.

Context: 1,000,000 tokens
Vision
Tools

Amazon Nova Micro

Amazon

Amazon's text-only, lowest-latency Nova model (November 2024) on AWS Bedrock, delivering the fastest responses at the lowest cost per inference. Nova Micro suits high-throughput text tasks like classification, routing, and simple generation with a 128K context window.

Context: 128,000 tokens
Tools

Amazon Nova Lite

Amazon

Amazon's low-cost multimodal model (November 2024) on AWS Bedrock, built for fast processing of image, video, and text at very low price. Nova Lite suits high-volume multimodal workloads where cost and latency matter more than peak accuracy.

Context: 300,000 tokens
Vision
Tools

Amazon Nova Pro

Amazon

Amazon's highly capable multimodal model (November 2024) on AWS Bedrock, balancing accuracy, speed, and cost. Nova Pro processes text, image, and video inputs with a 300K-token context window, targeting document analysis, agentic workflows, and enterprise applications.

Context: 300,000 tokens
Vision
Tools

Falcon 3 10B

Technology Innovation Institute

Open Weights

TII's open decoder-only model (December 2024) from the Falcon 3 family (1B-10B), trained on 14 trillion tokens of web, code, STEM, and multilingual data. Falcon 3 10B offers a 32K context window and strong efficiency for its size, deployable locally via llama.cpp and MLX.

Context: 32,768 tokens
Tools

Yi-1.5 34B

01.AI

Open Weights

01.AI's open bilingual model (May 2024), continuously pretrained on Yi with 500B additional high-quality tokens and fine-tuned on 3M samples for stronger coding, math, reasoning, and instruction following. Available in 6B/9B/34B sizes with context options up to 32K.

Context: 32,768 tokens
Tools

OLMo 2 32B

Allen Institute for AI

Open Weights

Ai2's fully open 32B model (March 2025) - the first fully open model (all data, code, weights, and training details public) to outperform GPT-3.5 Turbo and GPT-4o mini on a suite of multi-skill academic benchmarks. Trained on up to 6T tokens and post-trained with Tulu 3.1.

Context: 4,096 tokens
Tools

Llama 3.2 3B

Meta

Open Weights

Meta's compact 3B open-weight model (September 2024), part of the Llama 3.2 edge collection, optimized for on-device summarization, rewriting, instruction following, and agentic retrieval. It offers a 128K context window and multilingual support at a size that runs on phones and laptops.

Context: 131,072 tokens
Tools