Official code for "Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding" (CVPR 2026).
📄 Paper
conda create -n attpack python=3.10
conda activate attpack
cd src
bash setup.shFor Video-LLaVA, follow the installation instructions in the Video-LLaVA repository.
All examples below use LLaVA-1.5-7B. --rank_k and --rank_v set the compression ranks for the key and value caches (R_k and R_v).
python inference/inference_ocrvqa.py \
--model-path liuhaotian/llava-v1.5-7b \
--use-attpack \
--rank_k 64 \
--rank_v (64, 16) \
--output-path ocrvqa_attpack.jsonpython inference/inference_aokvqa.py \
--model-path liuhaotian/llava-v1.5-7b \
--use-attpack \
--rank_k 64 \
--rank_v 64 \
--output-path aokvqa_attpack.jsonpython inference/mmmu/inference_mmmu.py \
--model-path liuhaotian/llava-v1.5-7b \
--use-attpack \
--rank_k 64 \
--output-path mmmu_attpack.jsonIf you find this work useful, please cite:
@inproceedings{ilhan2026attentionpack,
title = {Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding},
author = {Ilhan, Fatih and Liu, Gaowen and Kompella, Ramana Rao and Tekin, Selim Furkan and Huang, Tiansheng and Yahn, Zachary and Xu, Yichang and Liu, Ling},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}