Local mirror for https://github.com/ggml-org/llama.cpp
- C++ 55.7%
- C 16.1%
- Python 7.2%
- Cuda 5.5%
- TypeScript 4.4%
- Other 10.9%
|
Some checks failed
build-cann.yml / mtmd: fix Granite4 Vision image sequence assembly (#26653) (push) Failing after 0s
ui-publish.yml / mtmd: fix Granite4 Vision image sequence assembly (#26653) (push) Failing after 0s
CI (3rd-party) / ubuntu-24-llguidance (push) Has been cancelled
CI (android) / default (push) Has been cancelled
CI (android) / ndk (push) Has been cancelled
CI (android) / arm64 (push) Has been cancelled
CI (apple) / macos-latest-arm64 (push) Has been cancelled
CI (cpu) / windows (x64, x64-cpu-static, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DBUILD_SHARED_LIBS=OFF) (push) Has been cancelled
CI (cpu) / windows (x64, x64-openblas, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_OPENMP=OFF -DGGML_BLAS=ON -DGGML_BLA… (push) Has been cancelled
CI (apple) / macos-latest-x64 (push) Has been cancelled
CI (apple) / macos-latest-ios-xcode (push) Has been cancelled
CI (apple) / macos-latest-tvos (push) Has been cancelled
CI (apple) / macos-latest-visionos (push) Has been cancelled
CI (apple) / macos-latest-swift (generic/platform=iOS) (push) Has been cancelled
CI (apple) / macos-latest-swift (generic/platform=macOS) (push) Has been cancelled
CI (apple) / macos-latest-swift (generic/platform=tvOS) (push) Has been cancelled
CI (cpu) / build-cmake-pkg (push) Has been cancelled
CI (cpu) / ubuntu (arm64, ubuntu-24.04-arm) (push) Has been cancelled
CI (cpu) / ubuntu (x64, ubuntu-22.04) (push) Has been cancelled
CI (cpu) / windows (arm64, arm64, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON) (push) Has been cancelled
CI (cpu) / windows (x64, x64-vulkan, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/x64-windows-llvm.cmake -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=OFF -DLLAMA_BUILD_SERVER=ON -DGGML_RPC=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DGGML_VULKAN=ON) (push) Has been cancelled
CI (CUDA, ubuntu) / cuda (push) Has been cancelled
CI (CUDA, ubuntu) / hip (push) Has been cancelled
CI (CUDA, ubuntu) / musa (push) Has been cancelled
CI (ibm) / ubuntu-24-s390x (push) Has been cancelled
CI (ibm) / ubuntu-24-ppc64le (push) Has been cancelled
CI (opencl) / windows-2025-opencl-adreno (push) Has been cancelled
CI (openvino) / ubuntu-24-openvino (push) Has been cancelled
CI (openvino) / openvino-windows-2022 (push) Has been cancelled
CI (riscv) / ubuntu-cpu-riscv64-native (push) Has been cancelled
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, ADDRESS) (push) Has been cancelled
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, THREAD) (push) Has been cancelled
CI (riscv) / ubuntu-riscv64-native-sanitizer (Debug, UNDEFINED) (push) Has been cancelled
CI (rpc) / ubuntu-24-rpc (push) Has been cancelled
CI (sanitize) / ctest ([self-hosted X64 Linux], UNDEFINED) (push) Has been cancelled
CI (sanitize) / ctest (ubuntu-24.04, ADDRESS) (push) Has been cancelled
CI (sanitize) / ctest (ubuntu-24.04, THREAD) (push) Has been cancelled
CI (self-hosted) / gpu-cuda (push) Has been cancelled
CI (self-hosted) / gpu-rocm (push) Has been cancelled
CI (self-hosted) / gpu-vulkan-nvidia-cm (push) Has been cancelled
CI (self-hosted) / gpu-vulkan-nvidia-cm2 (push) Has been cancelled
CI (self-hosted) / gpu-webgpu-nvidia (push) Has been cancelled
CI (self-hosted) / gpu-metal (push) Has been cancelled
CI (self-hosted) / gpu-webgpu-apple (push) Has been cancelled
CI (self-hosted) / gpu-vulkan-apple (push) Has been cancelled
CI (self-hosted) / gpu-vulkan-intel-linux (push) Has been cancelled
CI (self-hosted) / gpu-vulkan-intel-windows (push) Has been cancelled
CI (self-hosted) / gpu-openvino-low-perf (push) Has been cancelled
CI (self-hosted) / cpu-x64-high-perf (push) Has been cancelled
CI (self-hosted) / cpu-arm64-high-perf-graviton4 (push) Has been cancelled
CI (self-hosted) / cpu-arm64-graviton4-kleidiai (push) Has been cancelled
CI (sycl) / ubuntu-24-sycl (fp16, ON) (push) Has been cancelled
CI (sycl) / ubuntu-24-sycl (fp32, OFF) (push) Has been cancelled
CI (sycl) / windows-latest-sycl (push) Has been cancelled
CI (virtgpu) / ubuntu-24-virtgpu (push) Has been cancelled
CI (webgpu) / macos (push) Has been cancelled
CI (vulkan) / ubuntu-arm64 (push) Has been cancelled
CI (vulkan) / ubuntu-llvmpipe (push) Has been cancelled
CI (wasm) / ubuntu-webgpu (push) Has been cancelled
CI (webgpu) / format (push) Has been cancelled
Release / windows-openvino (push) Has been cancelled
CI (webgpu) / ubuntu (push) Has been cancelled
Code Style Checker / model-naming (push) Has been cancelled
EditorConfig Checker / editorconfig (push) Has been cancelled
Release / check-release (push) Has been cancelled
Release / get-version (push) Has been cancelled
Release / macos-cpu (arm64, arm64, -DGGML_METAL_EMBED_LIBRARY=ON -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-26) (push) Has been cancelled
Release / macos-cpu (x64, x64, -DGGML_METAL=OFF -DCMAKE_OSX_DEPLOYMENT_TARGET=13.3, macos-15-intel) (push) Has been cancelled
Release / ubuntu-cpu (arm64, ubuntu-24.04-arm) (push) Has been cancelled
Release / ubuntu-cpu (s390x, ubuntu-24.04-s390x) (push) Has been cancelled
Release / ubuntu-cpu (x64, ubuntu-22.04) (push) Has been cancelled
Release / ubuntu-vulkan (arm64, ubuntu-24.04-arm) (push) Has been cancelled
Release / ubuntu-vulkan (x64, ubuntu-22.04) (push) Has been cancelled
Release / android-arm64 (push) Has been cancelled
Release / ubuntu-24-openvino (push) Has been cancelled
Release / windows-cpu (arm64) (push) Has been cancelled
Release / windows-cpu (x64) (push) Has been cancelled
Release / windows-rocm (7.14.0, x64, gfx1010;gfx1011;gfx1012;gfx1030;gfx1031;gfx1032;gfx1033;gfx1034;gfx1035;gfx1036;gfx1100;gfx1101;gfx1102;gfx1103;gfx1150;gfx1151;gfx1152;gfx1153;gfx1200;gfx1201) (push) Has been cancelled
Release / windows (arm64, opencl-adreno, -G "Ninja Multi-Config" -D CMAKE_TOOLCHAIN_FILE=cmake/arm64-windows-llvm.cmake -DCMAKE_PREFIX_PATH="$env:RUNNER_TEMP/opencl-arm64-release" -DGGML_OPENCL=ON -DGGML_OPENCL_USE_ADRENO_KERNELS=ON, ggml-opencl) (push) Has been cancelled
Release / windows (x64, vulkan, -DGGML_VULKAN=ON, ggml-vulkan) (push) Has been cancelled
Server (self-hosted) / server-metal (push) Has been cancelled
Release / windows-cuda (13.4, arm64) (push) Has been cancelled
Release / windows-cuda (12.4, x64) (push) Has been cancelled
Release / windows-cuda (13.3, x64) (push) Has been cancelled
Release / windows-sycl (push) Has been cancelled
Release / ubuntu-24-sycl (fp16, ON) (push) Has been cancelled
Release / ubuntu-24-sycl (fp32, OFF) (push) Has been cancelled
Server (self-hosted) / server-cuda (push) Has been cancelled
Release / ios-xcode (push) Has been cancelled
Release / ui-build (push) Has been cancelled
Release / release (push) Has been cancelled
Release / ui-publish (push) Has been cancelled
Server (sanitize) / server (RelWithDebInfo, ADDRESS) (push) Has been cancelled
Server (sanitize) / server (RelWithDebInfo, UNDEFINED) (push) Has been cancelled
Server (self-hosted) / server-kleidiai (push) Has been cancelled
Server / ubuntu (push) Has been cancelled
Server / windows (push) Has been cancelled
* mtmd: fix granite 4v grid assembly (cherry picked from commit 91f82eb1b489bff0dc3f649edf2cb9e0b7f656e9) * mtmd: fix truncation for scaled image height and width before unpad Signed-off-by: Hemanth Battu <hbattu@ibm.com> * mtmd: remove MTMD_DUMP_EMBD debug scaffolding Signed-off-by: Hemanth Battu <hbattu@ibm.com> * clean up comments, clarify about anyres_info excluded from serialization * add_newline is now dead code --------- Signed-off-by: Hemanth Battu <hbattu@ibm.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: Hemanth Battu <hbattu@ibm.com> |
||
|---|---|---|
| .devops | ||
| .gemini | ||
| .github | ||
| .pi/gg | ||
| app | ||
| benches | ||
| ci | ||
| cmake | ||
| common | ||
| conversion | ||
| docs | ||
| examples | ||
| ggml | ||
| gguf-py | ||
| grammars | ||
| include | ||
| licenses | ||
| media | ||
| models | ||
| pocs | ||
| requirements | ||
| scripts | ||
| skills | ||
| src | ||
| tests | ||
| tools | ||
| vendor | ||
| .clang-format | ||
| .clang-tidy | ||
| .dockerignore | ||
| .ecrc | ||
| .editorconfig | ||
| .flake8 | ||
| .gitignore | ||
| .gitmodules | ||
| .pre-commit-config.yaml | ||
| AGENTS.md | ||
| AUTHORS | ||
| build-xcframework.sh | ||
| CLAUDE.md | ||
| CMakeLists.txt | ||
| CMakePresets.json | ||
| CODEOWNERS | ||
| CONTRIBUTING.md | ||
| convert_hf_to_gguf.py | ||
| convert_hf_to_gguf_update.py | ||
| convert_llama_ggml_to_gguf.py | ||
| convert_lora_to_gguf.py | ||
| flake.nix | ||
| LICENSE | ||
| Makefile | ||
| mypy.ini | ||
| pyproject.toml | ||
| pyrightconfig.json | ||
| README.md | ||
| requirements.txt | ||
| SECURITY.md | ||
| ty.toml | ||
llama.cpp
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs / compile times / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine:
- Visit https://llama.app and follow the instructions
- Run with Docker - see our Docker documentation
- Download pre-built binaries from the releases page
- Build from source by cloning this repository - check out our build guide
Once installed:
# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF
# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
|
|
|
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on
a wide range of hardware - locally and in the cloud.
- Plain C/C++ implementation without any dependencies
- Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
- AVX, AVX2, AVX512 and AMX support for x86 architectures
- RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
- 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
- Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
- Vulkan and SYCL backend support
- CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
| Backend | Target devices |
|---|---|
| BLAS | All |
| BLIS | All |
| CANN | Ascend NPU |
| CUDA | Nvidia GPU |
| HIP | AMD GPU |
| Hexagon [In Progress] | Snapdragon |
| IBM zDNN | IBM Z & LinuxONE |
| MUSA | Moore Threads GPU |
| Metal | Apple Silicon |
| OpenCL | Adreno GPU |
| OpenVINO [In Progress] | Intel CPUs, GPUs, and NPUs |
| RPC | All |
| SYCL | Intel GPU |
| VirtGPU | VirtGPU APIR |
| Vulkan | GPU |
| WebGPU | All |
| ZenDNN | AMD CPU |
Documentation
Tools
Development
- How to build
- Running on Docker
- Build on Android
- Multi-GPU usage
- Performance troubleshooting
- GGML tips & tricks
- XCFramework
- Completions
- Models
- Release process
Contributing
- Contributors can open PRs
- Collaborators will be invited based on contributions
- Maintainers can push to branches in the
llama.cpprepo and merge PRs into themasterbranch - Any help with managing issues, PRs and projects is very appreciated!
- Read the CONTRIBUTING.md for more information
Acknowledgements
- yhirose/cpp-httplib - Single-header HTTP server, used by
llama-server- MIT license - stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
- nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
- miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
- subprocess.h - Single-header process launching solution for C and C++ - Public domain