Skip to content

Pinned Loading

  1. vllm vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 92.7k 22.7k

  2. vllm-omni vllm-omni Public

    A framework for efficient model inference with omni-modality models

    Python 7.1k 1.8k

  3. recipes recipes Public

    Common recipes to run vLLM

    JavaScript 1k 431

  4. llm-compressor llm-compressor Public

    State-of-the-art LLM compression, built for production inference with vLLM

    Python 3.8k 672

  5. speculators speculators Public

    A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

    Python 853 234

  6. semantic-router semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    Go 5.9k 977

Repositories

Showing 10 of 49 repositories
  • semantic-router Public

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    vllm-project/semantic-router's past year of commit activity
    Go 5,918 Apache-2.0 977 352 (2 issues need help) 166 Updated Sep 26, 2026
  • tpu-inference Public

    TPU inference for vLLM, with unified JAX and PyTorch support.

    vllm-project/tpu-inference's past year of commit activity
    Python 440 Apache-2.0 326 100 (2 issues need help) 412 Updated Sep 26, 2026
  • vllm-metal Public

    Community maintained hardware plugin for vLLM on Apple Silicon

    vllm-project/vllm-metal's past year of commit activity
    Python 1,779 Apache-2.0 268 12 11 Updated Sep 26, 2026
  • vllm-gaudi Public

    Community maintained hardware plugin for vLLM on Intel Gaudi

    vllm-project/vllm-gaudi's past year of commit activity
    Python 58 Apache-2.0 152 8 49 Updated Sep 26, 2026
  • vllm-ascend Public

    Community maintained hardware plugin for vLLM on Huawei Ascend

    vllm-project/vllm-ascend's past year of commit activity
    Python 2,892 Apache-2.0 2,367 1,529 (77 issues need help) 2,100 Updated Sep 26, 2026
  • vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    vllm-project/vllm's past year of commit activity
    Python 92,701 Apache-2.0 22,679 2,524 (32 issues need help) 5,000+ Updated Sep 26, 2026
  • vllm-omni Public

    A framework for efficient model inference with omni-modality models

    vllm-project/vllm-omni's past year of commit activity
    Python 7,067 Apache-2.0 1,792 837 (152 issues need help) 1,214 Updated Sep 26, 2026
  • recipes Public

    Common recipes to run vLLM

    vllm-project/recipes's past year of commit activity
    JavaScript 1,029 Apache-2.0 431 57 131 Updated Sep 26, 2026
  • guidellm Public

    Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs

    vllm-project/guidellm's past year of commit activity
    Python 1,645 Apache-2.0 239 48 24 Updated Sep 26, 2026
  • humming Public

    Humming is a high-performance, lightweight, and highly flexible JIT (Just-In-Time) compiled GEMM kernel library specifically designed for quantized inference.

    vllm-project/humming's past year of commit activity
    Python 238 Apache-2.0 43 6 13 Updated Sep 26, 2026