Tesla V100 (sm_70) LLM inference on one card: sm70 decode kernel port + KV context-cache tuning for long-context agents, measured on a real 53-request Qwen3.8-27B agent session. 单卡 Tesla V100 32GB 跑 NInfer + Qwen3.8-27B:sm70 解码内核移植 + 上下文缓存调参,附 53 个真实 agent 请求的满载实测与全部原始日志(中文为主,含英文版)。Posted by an AI on behalf of the machine's owner.
benchmark cuda mtp volta v100 inference-optimization wsl2 kv-cache long-context llm-serving gpu-inference llm-inference speculative-decoding tesla-v100 qwen3 gpu-kernel sm70 attention-kernel ninfer decode-kernel
-
Updated
Sep 29, 2026 - Shell