A C++ inference engine for Vision-Language-Action (VLA) models, built on llama.cpp.
It runs the open VLA policies - SmolVLA, π0, BitVLA, Evo-1, GR00T N1.5/1.6/1.7 and more -
under one runtime, each packaged as a single self-contained GGUF that needs no Python or
PyTorch at inference time. The binaries drive robots on CPU, Apple Silicon, CUDA -
from consumer GPUs down to Jetson-class boards - Intel GPUs and NPUs via
SYCL and OpenVINO, or Qualcomm Snapdragon CPUs, Adreno GPUs and Hexagon NPUs
via OpenCL and the Hexagon backend.
Learn vla.cpp walks through the engine design and how each policy is implemented on ggml.
Download a prebuilt release, or build from source for any other platform or backend.
Each release has a
vla.cpp-<tag>-<platform>.tar.gz with libvla, vla.h and the binaries:
| Platform | Needs |
|---|---|
linux-x86_64-cpu |
AVX2 (Haswell or newer) |
linux-x86_64-cuda-12.8 |
AVX2; sm_75/80/86/89/90/120 |
linux-x86_64-cuda-13.4 |
AVX2; sm_75/80/86/89/90/120, driver 580 or newer |
linux-aarch64-cpu |
ARMv8.2-A with dotprod and fp16 (Cortex-A76, Neoverse N1 or newer) |
linux-aarch64-cuda-13.4 |
sm_87 (Orin), sm_110 (Thor), sm_121 (DGX Spark); a CUDA 13 driver |
macos-arm64-metal |
vla-cli and vla-bench only |
TAG=<release tag>
curl -LO https://github.com/VinRobotics/vla.cpp/releases/download/$TAG/vla.cpp-$TAG-linux-x86_64-cpu.tar.gz
tar -xzf vla.cpp-$TAG-linux-x86_64-cpu.tar.gzThe Linux tarballs need Ubuntu 24.04's glibc or newer, libzmq5 for the
servers, and a CUDA runtime for the CUDA ones (shipped as a separate cudart-
tarball). Details, and the macOS and Windows notes, are in
docs/PREBUILT.md.
- CMake ≥ 3.22
- A C++17 compiler (GCC 11+ or Clang 14+)
- CUDA 12.x or 13.x (optional - required only for CUDA GPU builds)
- Intel oneAPI 2025.x + GPU compute runtime (optional - only for Intel GPU builds, see docs/backend/sycl.md)
- OpenVINO 2026.x runtime (optional - only for Intel CPU/GPU/NPU builds via OpenVINO, see docs/backend/ov.md)
libzmq3-dev,cppzmq-dev,libprotobuf-dev,protobuf-compiler
sudo apt-get install -y libzmq3-dev cppzmq-dev libprotobuf-dev protobuf-compilerIdentify your machine CUDA architecture:
| GPU family | Example cards | CUDA_ARCHITECTURE |
|---|---|---|
| Ampere (Jetson) | Orin Nano, Orin NX | 87 |
| Ampere (consumer) | RTX 30-series, A40 | 86 |
| Ada Lovelace | RTX 40-series, L40 | 89 |
| Hopper | H100, H200 | 90 |
| Blackwell (consumer) | RTX 50-series | 120 |
| Blackwell (datacenter) | B100, B200, GB200 | 100 |
| Blackwell (Jetson) | Jetson Thor | 110 |
| Blackwell (DGX Spark) | GB10 | 121 |
Then configure and build. CMake fetches and pins llama.cpp automatically (no patch, no submodule):
# CPU build:
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
# CUDA build (set CMAKE_CUDA_ARCHITECTURES for your GPU):
cmake -B build \
-DGGML_CUDA=ON \
-DCMAKE_BUILD_TYPE=Release \
-DCMAKE_CUDA_ARCHITECTURES=$CUDA_ARCHITECTURE
cmake --build build -j$(nproc)If CMake cannot find CUDA, point the environment at it explicitly:
export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH-DVLA_BUILD_SERVER=OFF -DVLA_SPM=OFF builds vla-cli, vla-bench and
libvla without protobuf, ZeroMQ or SentencePiece, so none of the apt packages
above are needed. Without SentencePiece, pass Octo --tokens instead of --text.
cmake --install build --prefix <dir> copies the binaries, libraries and
share/vla/tokenize_prompt.py into <dir>; the result does not need the build
tree. pip install ./bindings/python builds the Python bindings, see
bindings/python/README.md.
Check docs/backend for compiling vla.cpp on other platforms.
WSL2, Apple Silicon, and Intel GPU are all tested.
To build and run in containers instead, see docs/DOCKER.md.
Once the binaries are built, run one CPU prediction without a server or simulator:
pip install -U "huggingface_hub[cli]" transformers
# -hf fetches and caches the checkpoint (under $VLA_CACHE, default ~/.cache/vla)
./build/vla-cli -hf vrfai/smolvla-libero-gguf \
--image assets/front.jpg --text "pick up the black bowl" --pretty
# or point at a file you already have
./build/vla-cli --ckpt models/smolvla/smolvla-libero.gguf \
--image assets/front.jpg --text "pick up the black bowl" --prettyWith a prebuilt tarball, run vla-cli from the extracted
vla.cpp-<tag>-<platform>/ directory instead of ./build/.
--pretty prints one action row per line and --state sets proprioception.
Prompt tokenization, -hf tags, vla-server, runtime flags and environment
variables are in docs/USAGE.md.
Models (rows) against platforms (columns). Legend: Y =
supported (released and benchmarked), ~ = in progress, - = planned.
| Model | CPU (x86-64 / ARM) | CUDA | SYCL (Intel) | Metal | OpenVINO | Hexagon |
|---|---|---|---|---|---|---|
| SmolVLA | Y | Y | Y | Y | Y | Y |
| π0 | Y | Y | - | Y | Y | ~ |
| π0.5 | Y | Y | - | Y | Y | Y |
| GR00T N1.5 | Y | Y | - | Y | Y | Y |
| GR00T N1.6 | Y | Y | - | Y | Y | Y |
| GR00T N1.7 | Y | Y | - | Y | Y | Y |
| BitVLA | Y | Y | - | ~ | - | - |
| Evo-1 | Y | Y | Y | Y | Y | Y |
| VLA-Adapter | Y | Y | ~ | Y | Y | Y |
| OpenVLA-OFT | Y | Y | - | Y | Y | - |
| VLA-JEPA | Y | Y | - | Y | Y | ~ |
| Octo-Small | Y | Y | Y | Y | - | Y |
| TurboVLA | Y | Y | Y | Y | Y | Y |
khanhnd61-vr/lerobot is a
LeRobot fork whose lerobot-vla-cpp client drives an SO-101 arm against a
vla-server. The client does the per-arch preprocessing (tokenize, resize,
normalize), so it covers SmolVLA, π0, π0.5 and GR00T N1.5/1.6/1.7.
git clone -b vla-simd https://github.com/khanhnd61-vr/lerobot.git
cd lerobot && pip install -e ".[vla-cpp]"Serve the policy, converted to GGUF as in docs/MODELS.md, on a GPU host. SmolVLA takes 38 ms per query on an RTX 5090 and 218 ms on a Jetson AGX Orin, against the 1.67 s of motion a 50-step chunk buys at 30 fps; on a desktop CPU it takes ~1.7 s and the action queue stalls.
./build/vla-server smolvla-so101.gguf # binds tcp://*:5555; --bind moves itCheck the round trip with no arm attached, then run the arm. ROBOT holds the
arm's --robot.* flags; the fork's README sets it up:
# prints round-trip latency next to the recorded actions
lerobot-vla-cpp --server_address=tcp://127.0.0.1:5555 --arch=smolvla \
--replay.repo_id=khanhnd61/so101-multi-task-clean --replay.steps=10
lerobot-vla-cpp --server_address=tcp://127.0.0.1:5555 --arch=smolvla "${ROBOT[@]}" \
--task="Put the tape into the box" --n_action_steps=25 --fps=30 --duration=60--archselects the preprocessing (smolvla,pi0,pi05,gr00t_n1_5/6/7, orpassthroughfor a server that preprocesses itself). The wrong arch loads, runs and returns plausible, wrong actions.--stats_jsonis required forpi05and the GR00T archs; GR00T also takes--embodiment, and--rel_stats_jsonfor an N1.7 checkpoint trained with relative actions.--taskmust match a trained instruction exactly, and the camera keys must stayfrontandwristin that order.vla-serveranswers one request at a time, so the loop is synchronous and--n_action_stepsis the feedback rate: 25 at 30 fps leaves ~0.83 s between observations.
The GR00T paths of the client have not been tested against real checkpoints yet. Wiring, recording, training and queue sizing are in the fork's README.
| Doc | Content |
|---|---|
| docs/PREBUILT.md | Release tarballs: layout, glibc and CUDA runtime requirements, macOS and Windows notes |
| docs/USAGE.md | vla-cli, vla-server, vla-bench: prompt tokenization, -hf tags, runtime flags, environment variables |
| docs/EVAL.md | Installing LIBERO and SimplerEnv, running the eval clients against vla-server |
| docs/MODELS.md | Converting safetensors checkpoints to GGUF, quantizing to Q8_0/Q4_0 |
| docs/QUANTIZATION.md | FoldQuant W8A8 / W4A4: the GGUF contract, the integer arithmetic, converting a FoldQuantVLA quantized model |
| docs/DOCKER.md | Building and running the eval in containers |
| docs/ARCHITECTURE.md | Engine design: layers, the prediction path, backends, adding an architecture |
| docs/backend/ | Per-backend build and run notes: SYCL, OpenVINO, Metal, Hexagon, Hexagon on Windows, WSL2 |
| docs/benchmark/ | Per-device latency and memory for every model, and the fastest flags per device |
| docs/KNOWN_ISSUES.md | Known issues and their resolutions |
| docs/ADOPTION.md | C ABI, Python bindings, release packaging, and what is left |
| docs/UPSTREAMING.md | The ggml-openvino fixes to send upstream to llama.cpp |
| CHANGELOG.md | Per-release changes, with the fastest configuration and success rate per model |
| CONTRIBUTING.md | Proving a change is numerically neutral, adding an architecture |
| bindings/python/ | Python bindings |
| Learn vla.cpp | Walkthrough of the engine design and each policy on ggml |
Licensed under the Apache License, Version 2.0.
Supported VLA models:
- SmolVLA - Hugging Face LeRobot team.
- π0,π0.5 - Physical Intelligence.
- BitVLA - Hongyu Wang et al.
- Evo-1 - Tao Lin et al.
- VLA-Adapter - Yihao Wang et al.
- OpenVLA-OFT - Moo Jin Kim et al.
- GR00T N1.x - NVIDIA Isaac.
- VLA-JEPA - Jingwen Sun et al.
- Octo - Octo Model Team, UC Berkeley RAIL.
- TurboVLA - Hengyi Xie et al.
Built on:
llama.cpp- LLM inference engine in C/C++.- LIBERO - benchmark suite for the success-rate sweeps.
- SimplerEnv - the second simulator in the eval scaffold.
