Tags: kvtc* + compression*

0 bookmark(s) - Sort by: Date โ†“ / Title /

  1. This paper introduces KVTC, a lightweight transform coder designed to compress key-value (KV) caches, which are crucial for efficient large language model (LLM) serving. KV caches enable reuse across conversation turns, but can consume significant GPU memory. KVTC addresses this by applying techniques from classical media compression โ€“ PCA-based decorrelation, adaptive quantization, and entropy coding โ€“ to reduce cache size without requiring changes to the underlying model. The authors demonstrate that KVTC achieves up to 20x compression while maintaining reasoning accuracy and long-context performance, and even higher compression for specific applications.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: tagged with "kvtc+compression"

About - Propulsed by SemanticScuttle