推論與部署 KV Cache 自回歸生成過程中保存先前詞元注意力鍵和值的快取。它避免重複計算完整前綴,但往往成為主要記憶體消耗項。 時間節點 2017–present 最近審查 2026-07-30 關鍵術語 keysvaluesprefilldecode 專案內關聯 基於 Transformer 被用於 Continuous Batching / PagedAttentionSpeculative DecodingKimi K3 相關技術 QuantizationTokenizationReasoning and test-time compute 歷史背景 大模型時代 (2020年代—現在) 第一手來源 Transformers cache explanation 這是經過選編的技術圖譜,並非對 AI 技術的完整涵蓋聲明。