Glossary › Deployment
Number of tokens or requests processed per second. Increased by batching, continuous batching (vLLM), and tensor parallelism. Trade-off with latency: higher throughput often means higher per-request latency.
DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary