Prefix-Aware Attention for LLM Decoding
Python 48 3
An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
Python 40 4
C++ 49 64
MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming