推理与部署 KV Cache 自回归生成过程中保存先前词元注意力键和值的缓存。它避免重复计算完整前缀,但往往成为主要内存消耗项。 时间节点 2017–present 最近审查 2026-07-30 关键术语 keysvaluesprefilldecode 项目内关联 基于 Transformer 被用于 Continuous Batching / PagedAttentionSpeculative DecodingKimi K3 相关技术 QuantizationTokenizationReasoning and test-time compute 历史背景 大模型时代 (2020年代—现在) 一手来源 Transformers cache explanation 这是经过选编的技术图谱,并非对 AI 技术的完整覆盖声明。