Glossary › Deployment

Latency

Time from request to first token (TTFT) or full response. TTFT depends on prompt processing. Token generation rate (tokens/s) determines total latency. Optimized by batching, quantization, speculative decoding.

Where this is taught

Related terms

Inference · Throughput

More in Deployment

Throughput · Streaming

Learn this by building

DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary