Stable Diffusion XL
Open Weights
Stability AI
Stability AI's flagship open-source text-to-image generation model. Features 3.5B parameter base model with 6.6B parameter refiner in ensemble pipeline. Native 1024x1024 resolution (2x larger than SD 1.5) with improved generation for limbs, text, faces, and overall image quality. Uses dual CLIP networks (CLIP1 + CLIP2) for superior semantic understanding vs single CLIP. Achieves 89% prompt adherence vs SD 1.5's 71%. Supports image-to-image, inpainting, and outpainting workflows. Runs on consumer hardware (RTX 3060+ with 8GB VRAM).
Strengths
- Open-source with permissive CreativeML license - free to use and modify
- Native 1024x1024 resolution (2x SD 1.5, 1.33x SD 2.0)
- 89% prompt adherence vs SD 1.5's 71% (27% improvement)
- Dual CLIP networks capture semantic meaning better than single CLIP
- Runs on consumer GPUs - RTX 3060 or better with 8GB VRAM
- Fast generation on modern hardware - 3.5 seconds for 1024x1024 on RTX 3060+
- Supports image-to-image, inpainting, outpainting workflows
- Improved rendering of hands, text, faces, colors, contrast, shadows
Caveats
- 8GB VRAM minimum requirement excludes older/budget GPUs
- Slower than proprietary models (DALL-E 3, Midjourney) for equivalent quality
- Requires technical setup compared to hosted services
- Limited to specific resolutions (1024x1024, 1024x1792, 1792x1024)
- No built-in safety filters - user responsibility for content policy
- Ensemble pipeline (base + refiner) increases complexity
- Vendor
- Stability AI
- License
- CreativeML Open RAIL++-M License
- Release Date
- 2023-07-26
- Modalities
- textimage
| Benchmark | Score |
|---|---|
| Prompt Adherence | 89.0 |
Capabilities
Vision
Audio
Video
Tool Use
Pricing
Open source - free to download and self-host. Inference costs: ~$0.0013 per image on RTX 3090/4090