Skip to content

vla.cpp performance review - #32

Merged
khanhnd61-vr merged 61 commits into
mainfrom
b11223-numerics-perf
Sep 30, 2026
Merged

khanhnd61-vr merged 61 commits into
mainfrom
b11223-numerics-perf

Conversation

@anindex

@anindex anindex commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Bumps llama.cpp to b11223, fixes several numerics bugs against the reference implementations, hardens the servers and loaders, and makes predict faster. It also fixes the release tarballs, which could not start off the CI runner.

Numerics changes

Nine archs now produce different actions on purpose. Each fix was checked against the reference implementation the checkpoint was trained with, then run through a paired LIBERO-Object A/B (same seeds, McNemar test):

Fix Before After p
π0: image tokens were scaled by 1/sqrt(2048) 83/100 90/100 0.17
GR00T N1.7 / VLA-JEPA: Qwen3-VL mergers use erf GELU 97/100 99/100 0.5
SmolVLA: cross-attn k/v kept in F32 90/100 92/100 0.63
TurboVLA: DINOv3 final LayerNorm was skipped 100/100 100/100 1.0
BitVLA: round-to-nearest-even bf16 98/100 99/100 1.0
VLA-Adapter: eval client normalizes proprio 298/300 295/300 0.45
Octo: tanh GELU, JAX eps, stem im2col 89/200 80/200 0.27

None of the differences are significant. The case for each fix is reference parity: TurboVLA now matches PyTorch to 9e-6, Octo matches JAX to 2e-6 on CPU, π0's image rows match lerobot v0.4.4 embed_prefix, and the Qwen3-VL merger matches HF to 6e-6. π0.5 and Evo-1 also move by rounding or small reference-alignment changes.

Speed

Default flags on an RTX 5090. Each number is the minimum over 5 rounds, alternating the 7abe1b4 build and this branch:

Arch ms before ms after
SmolVLA 46.8 38.2 -18%
π0.5 35.0 29.2 -16%
GR00T N1.6 22.8 19.4 -15%
TurboVLA 5.71 4.88 -15%
VLA-Adapter 18.1 16.0 -12%
GR00T N1.7 21.6 19.5 -10%
GR00T N1.5 17.7 16.1 -9%
OpenVLA-OFT 37.9 34.5 -9%
π0 31.6 28.8 -9%
VLA-JEPA 15.6 14.3 -8%
Octo 2.89 2.69 -7%
Evo-1 52.3 49.1 -6%
BitVLA 20.2 20.8 +3%

The gains come from computing step-constant conditioning once at load (π0.5 adaRMS, DiT adaLN) and keeping only the selected GR00T embodiment resident, which saves 1.2 GiB of VRAM. SmolVLA widens its weights at load instead of on every call, and TurboVLA caches the encoded instruction. Vision outputs and constant tables now stay on the device. BitVLA's 3% is the cost of correct bf16 rounding.

About 1300 lines of per-arch copies were moved onto src/layers and src/modules, with byte-identical output for every arch.

Fixes

  • vla-server no longer aborts on a bad request. Previously, each of these killed the process:

    • an out-of-vocab Octo token;
    • 9 to 16 OFT or VLA-Adapter views;
    • a BitVLA prompt past 1024 tokens;
    • Octo stats without a mask.

    Images are now JPEG or PNG only, a request is capped at 64 MPx, and a port already in use gives an error instead of SIGABRT.

  • vlm-server:

    • it read an uninitialized prompt length;
    • a template exception killed it;
    • multi-turn prompts were corrupted;
    • streamed deltas could be invalid UTF-8;
    • network bytes reached the ffmpeg image fallback.
  • Model loading:

    • fixed-size graphs aborted with more views or steps;
    • malformed GGUF metadata divided by zero or read past buffers;
    • WeightLoader::fuse read quantized sources as float;
    • two models in one process shared their precision flags.
  • BitVLA:

    • BF16 and Q8 weights were uploaded as float;
    • VLA_DEVICE was ignored;
    • CUDA errors were dropped;
    • three kernels had shared-memory races.

    The BF16 CUDA hook could also write F32 into a BF16 buffer.

  • Tooling:

    • quantize_gguf.py broke Octo and TurboVLA files and packed the action experts;
    • converters now use torch.load(weights_only=True);
    • the eval client's REQ socket jammed after one timeout.

Packaging and usability

  • Release tarballs ship every shared library with an $ORIGIN rpath, are built with GGML_NATIVE=OFF, and pass a smoke test with the build tree moved away. New linux-x86_64-cuda-13.4 and linux-aarch64-cuda-13.4 builds; the x86 CUDA list gains sm_80 and sm_90. Third-party licenses are included.
  • Docker image is multi-stage and covers sm_75 to sm_121 instead of sm_89 only.
  • Install and bindings: cmake --install works, and pip install ./bindings/python builds a self-contained wheel.
  • vla-cli takes the precision flags and --config. -hf accepts repo:sub/dir/file.gguf and :Q8_0 tags, lists candidates instead of guessing, and downloads only GGUFs.
  • --text builds each arch's real prompt. With scripts/add_tokenizer_to_gguf.py the tokenizer lives in the GGUF and no Python is needed.
  • --num-steps sets the solver step count at load.
  • LIBERO eval can pair episodes with seeded noise, which is what the A/B above used, and now also runs Octo, TurboVLA and VLA-JEPA.

Dependencies

  • llama.cpp b10729 to b11223, with re-anchored CUDA and OpenVINO patches and ggml_prec_set_acc.
  • SentencePiece v0.2.1 (heap-overflow fix).
  • OpenVINO 2026.4 with checksummed drivers.
  • macos-15 runner.
  • transformers 5.x and torch >= 2.6 for the tooling.

Breaking changes

  • TurboVLA: GGUFs without vit.norm fail to load with a message to re-convert, and vrfai/turbovla-libero-gguf needs a re-upload.
  • VLA_OCTO is now VLA_SPM: the old name still works with a warning, and Octo is always built.
  • Release artifacts: linux-x86_64-cuda is now linux-x86_64-cuda-12.8. The macOS tarball ships only vla-cli and vla-bench.
  • Stricter errors: a missing --config file, -hf together with a positional checkpoint, and --act-dtype bf16 on archs other than π0 and Evo-1 are now errors.

Testing

  • Oracle: fixed inputs through vla_predict_check on all 13 archs, CUDA and CPU. Every refactor and perf commit was gated on byte-identical actions, and every numerics fix on a match with its own reviewed build.
  • ctest: passes on CPU and CUDA, plus an ASan+UBSan build.
  • Builds: the minimal build (-DVLA_SPM=OFF -DVLA_BUILD_SERVER=OFF) links no protobuf or ZeroMQ, and the OpenVINO backend compiles in the 2026.4 image.
  • Abuse cases: the server cases above were reproduced with pyzmq before and after.

Follow-ups

  • Re-publish the TurboVLA GGUF.
  • Octo still achieves around 45% on LIBERO-Object, well below published numbers, independent of this change.
  • Giving cached graphs a stable id would save 0.5 to 0.7 ms per predict, but ggml exposes no public API for it.
  • HF library registration and a PyPI publish.

@anindex anindex changed the title llama.cpp b11223, numerics fixes, hardening and faster predict vla.cpp performance review Sep 29, 2026
@khanhnd61-vr
khanhnd61-vr merged commit f7e0f7f into main Sep 30, 2026
5 checks passed
@khanhnd61-vr
khanhnd61-vr deleted the b11223-numerics-perf branch September 30, 2026 11:57
hungho77 added a commit that referenced this pull request Oct 2, 2026
main's performance review (#32) moved pi0.5's Gemma layers into the shared
modules/gemma_expert.h and precomputes the DiT adaLN modulation per
denoising step (FlowTimes, ActionExpert::denoise).

- GR00T: main's per-step adaLN precompute replaces this branch's adaLN
  cache (same idea); the FoldQuant LLM / DiT sites and the GEMM prefetch
  chain are kept on top.
- pi0.5: the FoldQuant paths move into gemma_attn / gemma_mlp /
  gemma_layer. Prefix layers fuse the RMSNorm and folded gamma into the act
  node and the residual into the o / down epilogues; expert layers quantize
  the adaRMS output (gated residuals stay out of the epilogue). pi0 has no
  FoldQuant sites, so its float path is unchanged.
- loader: main's resident-type fuse check, plus the same source type for a
  tensor copied raw (INT8 codes).
- convert_pi05_to_gguf.py: convert() and the OpenPI config merge, plus
  main's QUANTILES-only check.
- Docs: FoldQuant under [Unreleased] in the CHANGELOG, its section in
  docs/MODELS.md, docs/QUANTIZATION.md linked from the README.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants