KV cache: from tensor shapes to capacity budgets
Autoregressive decoding keeps historical keys and values instead of recomputing that state at each step. The payload depends on layers, KV heads, head dimension, storage precision and retained tokens. This note distinguishes a logical tensor budget from GPU allocation and provides a reproducible capacity calculator.
INTERACTIVE / CAPACITY
KV capacity workbench
Illustrative parameters, not a named model. Calculates total logical KV payload, excluding weights, sharding, sharing and engine overhead.
- Per token · all layers
- 128.00 KiB
- Per sequence
- 1.00 GiB
2 × 32 × 8 × 128 × 2 × 8192 × 4 = 4294967296 bytes
Integer multiplication gives bytes; KiB = 2¹⁰ bytes and GiB = 2³⁰ bytes. Displays round to two decimal places.
Inputs: the editable illustrative fields above. Method: the integer formula on this page.