LMSYS:Blog(Chatbot Arena 团队)· SGLang, Qwen, and NVIDIA teams·· 23 天前AI 评分47
SGLang、Qwen 与 NVIDIA 在 SGLang 中实现 NVFP4 KV cache
Blog Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache The KV cache is a fundamental building block of the modern LLM inference system. The context from multiple conversation rounds in agent sessions is cached as keys and values (KV) in GPU memory, allowi... SGLang, Qwen, and NVIDIA teams
AI 导读
SGLang、Qwen 与 NVIDIA 团队在 SGLang 中实现了 NVFP4 KV cache,用于加速长上下文与智能体推理。NVFP4 以 4 位 E2M1 值配合每 16 个值一个 FP8 block scale 和 FP32 tensor 级 scale,每 16 个值仅占 8 字节加 1 字节元数据,约为 FP8 存储的 56%。
整理与数据来源:AIHOT
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org