Skip to content

Add AMD ROCm/HIP backend for SmolVLA and π0.5 on gfx1151 - #35

Open
singlecatlmx wants to merge 5 commits into
VinRobotics:mainfrom
singlecatlmx:pr/rocm-hip-v04
Open

singlecatlmx wants to merge 5 commits into
VinRobotics:mainfrom
singlecatlmx:pr/rocm-hip-v04

Conversation

@singlecatlmx

@singlecatlmx singlecatlmx commented Oct 3, 2026 •

Copy link
Copy Markdown

What

  • Route VLA graphs to ggml's ROCm/HIP backend with -DGGML_HIP=ON. Select a device with VLA_DEVICE; log a CPU fallback if initialization fails.
  • Keep backend boundaries explicit. CMake rejects multiple accelerators and GGML_BACKEND_DL=ON. The VLA core uses GGML_USE_HIP to distinguish HIP from ggml-hip's CUDA-named interfaces, excluding vla.cpp's CUDA-only code.
  • Document the build, fixed-input numerics, native-size engine latency and current limits in docs/backend/rocm.md. The README marks SmolVLA and π0.5 as the two ROCm models validated on this revision.

Why

Upstream GGML_HIP=ON builds ggml's HIP backend, but VLA backend selection has no HIP path, so VLA models still use CPU. ggml-hip also exports GGML_USE_CUDA and CUDA-named entry points; VLA must distinguish those from its NVIDIA-only kernels. This PR adds that path on the current mainline, including the FoldQuant changes in main@1adf078, and reports the two standard GGUF models validated so far.

Verified

  • CPU and HIP Release builds compile first-party code under -Wall -Wextra.
  • First-party CTest passes 10/10 on each build. HIP tests need the ROCm lib directory on LD_LIBRARY_PATH on this non-system installation.
  • CPU BF16 actions match upstream main@1adf078 byte for byte for both models. HIP actions are complete, finite and identical across five independent processes per model.

Test system: Ryzen AI Max+ 395 / Radeon 8060S (gfx1151), Linux 6.17.0-40, ROCm 7.14.60850 and pinned llama.cpp b11223. The two GGUF SHA-256 digests and exact commands are in docs/backend/rocm.md.

Archs and backends tested: SmolVLA and π0.5 on CPU and ROCm/HIP, with BF16 resident weights and F32 activations.

Fixed-input correctness

Both backends used the same GGUF and weight dtype. Each prediction returned a 50 × 32 action array. The following errors compare HIP against CPU BF16:

Model / inputs Max abs RMS Cosine
SmolVLA LIBERO, 2 × 512 1.992e-3 2.421e-4 0.9999991
π0.5 LIBERO, 2 × 224 8.792e-4 1.169e-4 0.9999998

Synthetic-input engine latency

At GPU performance level high, three fresh processes per model each ran three warmups and 20 timed predict() calls. The stock vla-bench reports P50, P90 and mean Vision time per round. The table uses the median round P50, worst round P90 and median round Vision mean. Timings are in milliseconds and exclude the server and simulator.

Model / inputs P50 Worst P90 Vision mean
SmolVLA, 2 × 512, 48 tokens 395.0 396.9 210.3
π0.5, 2 × 224, 128 tokens 389.1 403.1 47.1

Backend boundaries

An invalid VLA_DEVICE=999 logged a CPU fallback and produced the same action bytes as the CPU build. These combinations each failed at configure time: HIP+CUDA, HIP+Vulkan and HIP+GGML_BACKEND_DL=ON. The HIP build links ggml-hip, hipBLAS, rocBLAS and the HIP runtime, with no vla.cpp CUDA kernel object, CUDA runtime or cuBLAS link.

Reproduce

HIPCXX="$(hipconfig -l)/clang" HIP_PATH="$(hipconfig -R)" \
cmake -B build-rocm -G Ninja -DCMAKE_BUILD_TYPE=Release \
    -DGGML_HIP=ON -DGPU_TARGETS=gfx1151 -DVLA_BUILD_TESTS=ON \
    -DVLA_BUILD_SERVER=OFF -DVLA_SPM=OFF
cmake --build build-rocm -j"$(nproc)"
export VLA_ROCM_ROOT="$(hipconfig -R)"
export LD_LIBRARY_PATH="$VLA_ROCM_ROOT/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
ctest --test-dir build-rocm --output-on-failure

For the exact GGUFs, CPU/HIP vla_predict_check commands, benchmark commands and percentile method, see docs/backend/rocm.md.

Scope and limitations

  • Only Linux gfx1151 and these two standard GGUFs were revalidated against llama.cpp b11223. Other models, FoldQuant quantized GGUFs and AMD targets need their own checks. This PR does not add native HIP FoldQuant integer kernels; the published BitVLA int2 path still requires vla.cpp's CUDA-only kernels.
  • This run used VLA_BUILD_SERVER=OFF: server/client execution, LIBERO task success and a long-duration soak were not retested on this mainline. The synthetic benchmarks and CPU agreement above do not establish robot-task equivalence.
  • HIP has no per-op CPU scheduler fallback. A graph containing an unsupported HIP op fails prediction.

On Linux gfx1151 with ROCm 7.14.60850 and llama.cpp b11223, CPU and HIP Release builds pass 8/8 CTest cases. SmolVLA and pi0.5 CPU BF16 actions remain byte-identical to upstream main@8eec763; HIP actions are finite and byte-identical across five fresh runs each. HIP-vs-CPU BF16 max abs is 1.992e-3 for SmolVLA and 8.792e-4 for pi0.5.

An invalid HIP device ordinal falls back to CPU with identical action bytes. CMake rejects HIP with CUDA/Vulkan or dynamic backend loading; HIP builds exclude vla.cpp CUDA objects and NVIDIA runtime links. Synthetic-input HIP P50 is 395.0 ms for SmolVLA and 387.8 ms for pi0.5 under the recorded high clock setting. Other models and task success need validation on this upstream revision.
Limit the ROCm support matrix to SmolVLA and pi0.5 on Linux gfx1151, ROCm 7.14.60850 and llama.cpp b11223. CPU/HIP Release CTest passes 8/8; CPU actions match upstream main@8eec763 byte for byte, and five HIP runs per model are finite and identical. HIP-vs-CPU BF16 max abs is 1.992e-3 and 8.792e-4 respectively.

Document GGUF SHA-256 values, build/runtime paths, fallback checks and three-round synthetic latency using stock vla-bench P50, P90 and Vision mean. Other architectures, server/client execution and robot-task success remain unverified on this revision.
ggml-hip also exports GGML_USE_CUDA. Define GGML_USE_HIP for vla_core and use it to select the HIP rung and exclude VLA CUDA-only BF16 code.

Verified on gfx1151, ROCm 7.14.60850 and llama.cpp b11223: CPU and HIP Release rebuilds, 8/8 CTest each. SmolVLA and pi0.5 BF16 fixed-input action bytes match their pre-change CPU/HIP baselines; the HIP banner remains ROCm/HIP. Numeric output did not move. The prior CPU/HIP max-abs differences were 1.992e-3 and 8.792e-4. Performance was not remeasured for this macro-only change; the A9 3x20 engine baseline remains the reference.
Resolve the Unreleased changelog entries for ROCm/HIP and FoldQuant while retaining the combined CUDA/HIP guard in the registration header.

Verified on gfx1151 with ROCm 7.14.60850 and llama.cpp b11223: CPU and HIP Release builds pass 10/10 CTest each. SmolVLA and pi05 BF16 fixed-input CPU/HIP actions remain byte-identical to the pre-merge branch; HIP-vs-CPU max abs errors are 1.992e-3 and 8.792e-4, respectively. Invalid HIP device falls back to CPU with identical actions. Three 20-rep HIP benchmark rounds per model show median P50 395.0/389.1 ms for SmolVLA/pi05. No CUDA-only VLA kernels or NVIDIA libraries are linked in the HIP build.
Document CPU/HIP tests and two-model fixed-input numerics on main 1adf078. Both Release builds pass 10/10 CTest on gfx1151 with ROCm 7.14.60850 and llama.cpp b11223. CPU actions match upstream main byte for byte; HIP actions are finite and repeat across five processes. HIP-vs-CPU max abs errors remain 1.992e-3 for SmolVLA and 8.792e-4 for pi05. Fresh 3x20 synthetic benchmarks report median P50 395.0/389.1 ms and worst P90 396.9/403.1 ms. Mark FoldQuant quantized GGUFs as unmeasured on HIP.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant