Figures › Deployment

Inference Latency: Prefill vs. Decode

Token-by-token generation timeline showing the first token paying the prefill cost while subsequent tokens decode in a fraction of the time.

FIG_026 · Animated diagram · DeVenture Academy

Terms in this diagram

Token · Inference · Latency

Learn the concept behind the diagram

This figure comes from the DeVenture Academy curriculum — project-based AI engineering courses where every lesson is paired with a hands-on lab you run on your own machine.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary