Glossary › Architecture
Architecture with multiple specialized sub-networks (experts) and a gating network that routes tokens to the best expert(s). Increases parameters without proportional compute. Used in Mixtral, GPT-4.
Attention Mechanism · Transformer · RAG (Retrieval-Augmented Generation) · Encoder-Decoder · KV Cache
DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary