Repository navigation
vla.cpp performance review - #32
Merged
Merged
Conversation
hungho77
added a commit
that referenced
this pull request
Oct 2, 2026
main's performance review (#32) moved pi0.5's Gemma layers into the shared modules/gemma_expert.h and precomputes the DiT adaLN modulation per denoising step (FlowTimes, ActionExpert::denoise). - GR00T: main's per-step adaLN precompute replaces this branch's adaLN cache (same idea); the FoldQuant LLM / DiT sites and the GEMM prefetch chain are kept on top. - pi0.5: the FoldQuant paths move into gemma_attn / gemma_mlp / gemma_layer. Prefix layers fuse the RMSNorm and folded gamma into the act node and the residual into the o / down epilogues; expert layers quantize the adaRMS output (gated residuals stay out of the epilogue). pi0 has no FoldQuant sites, so its float path is unchanged. - loader: main's resident-type fuse check, plus the same source type for a tensor copied raw (INT8 codes). - convert_pi05_to_gguf.py: convert() and the OpenPI config merge, plus main's QUANTILES-only check. - Docs: FoldQuant under [Unreleased] in the CHANGELOG, its section in docs/MODELS.md, docs/QUANTIZATION.md linked from the README.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps llama.cpp to b11223, fixes several numerics bugs against the reference implementations, hardens the servers and loaders, and makes predict faster. It also fixes the release tarballs, which could not start off the CI runner.
Numerics changes
Nine archs now produce different actions on purpose. Each fix was checked against the reference implementation the checkpoint was trained with, then run through a paired LIBERO-Object A/B (same seeds, McNemar test):
None of the differences are significant. The case for each fix is reference parity: TurboVLA now matches PyTorch to 9e-6, Octo matches JAX to 2e-6 on CPU, π0's image rows match lerobot v0.4.4
embed_prefix, and the Qwen3-VL merger matches HF to 6e-6. π0.5 and Evo-1 also move by rounding or small reference-alignment changes.Speed
Default flags on an RTX 5090. Each number is the minimum over 5 rounds, alternating the 7abe1b4 build and this branch:
The gains come from computing step-constant conditioning once at load (π0.5 adaRMS, DiT adaLN) and keeping only the selected GR00T embodiment resident, which saves 1.2 GiB of VRAM. SmolVLA widens its weights at load instead of on every call, and TurboVLA caches the encoded instruction. Vision outputs and constant tables now stay on the device. BitVLA's 3% is the cost of correct bf16 rounding.
About 1300 lines of per-arch copies were moved onto
src/layersandsrc/modules, with byte-identical output for every arch.Fixes
vla-server no longer aborts on a bad request. Previously, each of these killed the process:
Images are now JPEG or PNG only, a request is capped at 64 MPx, and a port already in use gives an error instead of SIGABRT.
vlm-server:
Model loading:
WeightLoader::fuseread quantized sources as float;BitVLA:
VLA_DEVICEwas ignored;The BF16 CUDA hook could also write F32 into a BF16 buffer.
Tooling:
quantize_gguf.pybroke Octo and TurboVLA files and packed the action experts;torch.load(weights_only=True);Packaging and usability
$ORIGINrpath, are built withGGML_NATIVE=OFF, and pass a smoke test with the build tree moved away. Newlinux-x86_64-cuda-13.4andlinux-aarch64-cuda-13.4builds; the x86 CUDA list gains sm_80 and sm_90. Third-party licenses are included.cmake --installworks, andpip install ./bindings/pythonbuilds a self-contained wheel.--config.-hfacceptsrepo:sub/dir/file.ggufand:Q8_0tags, lists candidates instead of guessing, and downloads only GGUFs.--textbuilds each arch's real prompt. Withscripts/add_tokenizer_to_gguf.pythe tokenizer lives in the GGUF and no Python is needed.--num-stepssets the solver step count at load.Dependencies
ggml_prec_set_acc.macos-15runner.Breaking changes
vit.normfail to load with a message to re-convert, andvrfai/turbovla-libero-ggufneeds a re-upload.VLA_OCTOis nowVLA_SPM: the old name still works with a warning, and Octo is always built.linux-x86_64-cudais nowlinux-x86_64-cuda-12.8. The macOS tarball ships onlyvla-cliandvla-bench.--configfile,-hftogether with a positional checkpoint, and--act-dtype bf16on archs other than π0 and Evo-1 are now errors.Testing
vla_predict_checkon all 13 archs, CUDA and CPU. Every refactor and perf commit was gated on byte-identical actions, and every numerics fix on a match with its own reviewed build.-DVLA_SPM=OFF -DVLA_BUILD_SERVER=OFF) links no protobuf or ZeroMQ, and the OpenVINO backend compiles in the 2026.4 image.Follow-ups