Glossary › Infrastructure
Reducing model precision from FP16 to INT8, INT4, or lower. Reduces memory and speeds up inference with minimal quality loss. Methods: GPTQ, AWQ, GGUF, bitsandbytes. Essential for local deployment.
Vector Store · ANN (Approximate Nearest Neighbor) · Inference · vLLM · Speculative Decoding
DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary