Repository navigation
Add AMD ROCm/HIP backend for SmolVLA and π0.5 on gfx1151 - #35
Open
singlecatlmx wants to merge 5 commits into
Open
singlecatlmx wants to merge 5 commits into
singlecatlmx wants to merge 5 commits into
Conversation
On Linux gfx1151 with ROCm 7.14.60850 and llama.cpp b11223, CPU and HIP Release builds pass 8/8 CTest cases. SmolVLA and pi0.5 CPU BF16 actions remain byte-identical to upstream main@8eec763; HIP actions are finite and byte-identical across five fresh runs each. HIP-vs-CPU BF16 max abs is 1.992e-3 for SmolVLA and 8.792e-4 for pi0.5. An invalid HIP device ordinal falls back to CPU with identical action bytes. CMake rejects HIP with CUDA/Vulkan or dynamic backend loading; HIP builds exclude vla.cpp CUDA objects and NVIDIA runtime links. Synthetic-input HIP P50 is 395.0 ms for SmolVLA and 387.8 ms for pi0.5 under the recorded high clock setting. Other models and task success need validation on this upstream revision.
singlecatlmx
force-pushed
the
pr/rocm-hip-v04
branch
from
October 3, 2026 13:42
3fe9188 to
f332a57
Compare
Limit the ROCm support matrix to SmolVLA and pi0.5 on Linux gfx1151, ROCm 7.14.60850 and llama.cpp b11223. CPU/HIP Release CTest passes 8/8; CPU actions match upstream main@8eec763 byte for byte, and five HIP runs per model are finite and identical. HIP-vs-CPU BF16 max abs is 1.992e-3 and 8.792e-4 respectively. Document GGUF SHA-256 values, build/runtime paths, fallback checks and three-round synthetic latency using stock vla-bench P50, P90 and Vision mean. Other architectures, server/client execution and robot-task success remain unverified on this revision.
ggml-hip also exports GGML_USE_CUDA. Define GGML_USE_HIP for vla_core and use it to select the HIP rung and exclude VLA CUDA-only BF16 code. Verified on gfx1151, ROCm 7.14.60850 and llama.cpp b11223: CPU and HIP Release rebuilds, 8/8 CTest each. SmolVLA and pi0.5 BF16 fixed-input action bytes match their pre-change CPU/HIP baselines; the HIP banner remains ROCm/HIP. Numeric output did not move. The prior CPU/HIP max-abs differences were 1.992e-3 and 8.792e-4. Performance was not remeasured for this macro-only change; the A9 3x20 engine baseline remains the reference.
singlecatlmx
force-pushed
the
pr/rocm-hip-v04
branch
from
October 3, 2026 14:39
f332a57 to
30fb537
Compare
Resolve the Unreleased changelog entries for ROCm/HIP and FoldQuant while retaining the combined CUDA/HIP guard in the registration header. Verified on gfx1151 with ROCm 7.14.60850 and llama.cpp b11223: CPU and HIP Release builds pass 10/10 CTest each. SmolVLA and pi05 BF16 fixed-input CPU/HIP actions remain byte-identical to the pre-merge branch; HIP-vs-CPU max abs errors are 1.992e-3 and 8.792e-4, respectively. Invalid HIP device falls back to CPU with identical actions. Three 20-rep HIP benchmark rounds per model show median P50 395.0/389.1 ms for SmolVLA/pi05. No CUDA-only VLA kernels or NVIDIA libraries are linked in the HIP build.
Document CPU/HIP tests and two-model fixed-input numerics on main 1adf078. Both Release builds pass 10/10 CTest on gfx1151 with ROCm 7.14.60850 and llama.cpp b11223. CPU actions match upstream main byte for byte; HIP actions are finite and repeat across five processes. HIP-vs-CPU max abs errors remain 1.992e-3 for SmolVLA and 8.792e-4 for pi05. Fresh 3x20 synthetic benchmarks report median P50 395.0/389.1 ms and worst P90 396.9/403.1 ms. Mark FoldQuant quantized GGUFs as unmeasured on HIP.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
-DGGML_HIP=ON. Select a device withVLA_DEVICE; log a CPU fallback if initialization fails.GGML_BACKEND_DL=ON. The VLA core usesGGML_USE_HIPto distinguish HIP from ggml-hip's CUDA-named interfaces, excluding vla.cpp's CUDA-only code.docs/backend/rocm.md. The README marks SmolVLA and π0.5 as the two ROCm models validated on this revision.Why
Upstream
GGML_HIP=ONbuilds ggml's HIP backend, but VLA backend selection has no HIP path, so VLA models still use CPU. ggml-hip also exportsGGML_USE_CUDAand CUDA-named entry points; VLA must distinguish those from its NVIDIA-only kernels. This PR adds that path on the current mainline, including the FoldQuant changes inmain@1adf078, and reports the two standard GGUF models validated so far.Verified
-Wall -Wextra.libdirectory onLD_LIBRARY_PATHon this non-system installation.main@1adf078byte for byte for both models. HIP actions are complete, finite and identical across five independent processes per model.Test system: Ryzen AI Max+ 395 / Radeon 8060S (
gfx1151), Linux 6.17.0-40, ROCm 7.14.60850 and pinned llama.cppb11223. The two GGUF SHA-256 digests and exact commands are indocs/backend/rocm.md.Archs and backends tested: SmolVLA and π0.5 on CPU and ROCm/HIP, with BF16 resident weights and F32 activations.
Fixed-input correctness
Both backends used the same GGUF and weight dtype. Each prediction returned a 50 × 32 action array. The following errors compare HIP against CPU BF16:
Synthetic-input engine latency
At GPU performance level
high, three fresh processes per model each ran three warmups and 20 timedpredict()calls. The stockvla-benchreports P50, P90 and mean Vision time per round. The table uses the median round P50, worst round P90 and median round Vision mean. Timings are in milliseconds and exclude the server and simulator.Backend boundaries
An invalid
VLA_DEVICE=999logged a CPU fallback and produced the same action bytes as the CPU build. These combinations each failed at configure time: HIP+CUDA, HIP+Vulkan and HIP+GGML_BACKEND_DL=ON. The HIP build links ggml-hip, hipBLAS, rocBLAS and the HIP runtime, with no vla.cpp CUDA kernel object, CUDA runtime or cuBLAS link.Reproduce
For the exact GGUFs, CPU/HIP
vla_predict_checkcommands, benchmark commands and percentile method, seedocs/backend/rocm.md.Scope and limitations
gfx1151and these two standard GGUFs were revalidated against llama.cppb11223. Other models, FoldQuant quantized GGUFs and AMD targets need their own checks. This PR does not add native HIP FoldQuant integer kernels; the published BitVLA int2 path still requires vla.cpp's CUDA-only kernels.VLA_BUILD_SERVER=OFF: server/client execution, LIBERO task success and a long-duration soak were not retested on this mainline. The synthetic benchmarks and CPU agreement above do not establish robot-task equivalence.