Glossary › Infrastructure

vLLM

High-throughput inference engine using PagedAttention for efficient KV cache management. 2-4× faster than HuggingFace transformers. Supports continuous batching, tensor parallelism, and quantization.

Where this is taught

Related terms

Inference · KV Cache

More in Infrastructure

Vector Store · ANN (Approximate Nearest Neighbor) · Inference · Quantization · Speculative Decoding

Learn this by building

DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary