Skip to content

feat: support Wan2.2 VACE - #2062

Merged
leejet merged 2 commits into
leejet:masterfrom
losewayy:wan22-vace
Sep 27, 2026
Merged

leejet merged 2 commits into
leejet:masterfrom
losewayy:wan22-vace

Conversation

@losewayy

Copy link
Copy Markdown
Contributor

Summary

Issue #1953 asked about Wan2.2 VACE. Testing showed the existing VACE path already handles the new checkpoints correctly once they are loaded — the real obstacles were two problems in the GGUF reading path:

1. Misleading errors when loading >4-dim tensors. Wan2.2 GGUFs (QuantStack) store conv3d weights as 5-dim tensors (patch_embedding.weight, vace_patch_embedding.weight). gguf_init_from_file rejects them with invalid number of dimensions: 5 > 4 and logs an ERROR, even though the extended reader folds them correctly — which made these models look unsupported. Wan2.1_14B_VACE and Wan2.2-T2V GGUFs carry the same 5-dim tensors, so this affects the whole Wan 14B family.

2. A latent bug in the extended reader's ARRAY metadata parsing. It read each array element as a full key-value pair, but GGUF array elements are raw values (no key/type headers). Any file with array-typed metadata (e.g. the ~32k tokenizer vocabulary in umt5-xxl-encoder) fails in that path — which went unnoticed because the extended reader only runs as a fallback.

Changes:

  • gguf_reader_ext.h: parse ARRAY metadata by element type (fixed-size seeks / per-string skips); add has_tensors_beyond_ggml_limits().
  • gguf_io.cpp: do a cheap header scan first; files declaring >4-dim tensors go straight to the extended reader instead of calling gguf_init_from_file, which is guaranteed to fail. All other files keep the existing behavior (gguf first, reader as fallback).

Also added a docs/wan.md section for Wan2.2 VACE-Fun A14B (MoE expert pair usage) and a generated sample clip.

Related Issue / Discussion

Fix #1953.

Additional Information

Verified locally on an RTX 5070 Ti (CUDA):

  • Wan2.2-VACE-Fun-A14B Q3_K_S high/low-noise experts: R2V produces coherent reference-conditioned output (sample in assets/wan/), T2V runs cleanly
  • Regular GGUFs (umt5-xxl-encoder Q8_0) load unchanged
  • A truncated file still fails cleanly, no crash

Not covered: Vulkan backend (the code path is backend-agnostic; the issue reporter offered to verify), V2V control video (shares the Wan2.1 VACE path).

Checklist

losewayy and others added 2 commits September 26, 2026 02:54
- route GGUFs declaring >4-dim tensors to the extended reader instead of
  gguf_init_from_file, which rejects them with a misleading ERROR
- fix the extended reader's ARRAY metadata parsing: elements are raw
  values, not key-value entries
- document Wan2.2 VACE-Fun A14B (MoE expert pair) and add a sample clip
@leejet
leejet merged commit 168f7b8 into leejet:master Sep 27, 2026
9 checks passed
@leejet

leejet commented Sep 27, 2026

Copy link
Copy Markdown
Owner

I’ve updated the PR title to use fix:. Wan2.2 VACE is already supported by the existing implementation; this PR fixes GGUF metadata parsing, avoids misleading errors when loading five-dimensional tensors, and adds usage documentation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Wan 2.2 VACE support in vid_gen (reference-to-video / video editing)

2 participants