Skip to content

Pinned Loading

  1. PAT PAT Public

    Prefix-Aware Attention for LLM Decoding

    Python 48 3

  2. RAGPulse RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    Python 40 4

  3. flash-linear-attention-npu flash-linear-attention-npu Public

    C++ 49 64

Repositories

Showing 5 of 5 repositories
  • flashserve/flash-linear-attention-npu's past year of commit activity
    C++ 49 64 95 112 Updated Sep 29, 2026
  • OLED-MoE Public
    flashserve/OLED-MoE's past year of commit activity
    Python 1 0 0 0 Updated Sep 25, 2026
  • RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    flashserve/RAGPulse's past year of commit activity
    Python 40 MIT 4 0 0 Updated Sep 1, 2026
  • PAT Public

    Prefix-Aware Attention for LLM Decoding

    flashserve/PAT's past year of commit activity
    Python 48 MIT 3 0 0 Updated May 26, 2026
  • Mosaic Public

    MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

    flashserve/Mosaic's past year of commit activity
    Python 6 Apache-2.0 0 0 0 Updated May 23, 2026