llama.cpp
Open Source
Library
Run LLMs locally in C/C++
About
llama.cpp is a C/C++ implementation for running LLM inference efficiently on CPUs and GPUs, powering much of the local-AI ecosystem. It supports GGUF quantized models and runs on laptops, phones, and edge devices.
Compatibility
Supported Languages
c
c++
Details
- Category
- Library
- License
- MIT