Glossary › Architecture

KV Cache

Stores computed Key and Value tensors from previous tokens during autoregressive generation. Eliminates recomputation of past attention — O(n) per new token instead of O(n²). Memory grows linearly with sequence length.

Where this is taught

Related terms

Attention Mechanism · Inference

More in Architecture

Attention Mechanism · Transformer · RAG (Retrieval-Augmented Generation) · Encoder-Decoder · Mixture of Experts (MoE)

Learn this by building

DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary