High-performance single-GPU inference for selected model checkpoints and GPUs.
-
Updated
Sep 24, 2026 - C++
High-performance single-GPU inference for selected model checkpoints and GPUs.
Port NInfer, a single-GPU CUDA inference engine, to the NVIDIA L20 (Ada sm_89, 92 SMs, 48 GB): patch set, build tooling, and measured results
Measuring proxy + throughput dashboard for local LLM engines (llama.cpp, NInfer): live tok/s, cache-hit rate, TTFT, and history charts. Single-file panel, stdlib-only proxy, MIT.
Qwen3.8-27B on RTX 4090 D (48GB): production deployment of NInfer with MTP7 + E8 KV + NVMe disk cache, 195 tok/s decode, crash forensics for WDDM desktop GPUs
Rust CLI to run local LLMs through rootless Podman Compose
To associate your repository with the ninfer topic, visit your repo's landing page and select "manage topics."