Glossary › Architecture
Neural network architecture based on self-attention (Vaswani et al., 2017). Processes all tokens in parallel (not sequential like RNNs). Foundation of GPT, BERT, and all modern LLMs. Key innovation: attention is all you need.
Attention Mechanism · Encoder-Decoder
Attention Mechanism · RAG (Retrieval-Augmented Generation) · Encoder-Decoder · Mixture of Experts (MoE) · KV Cache
DeVenture Academy teaches AI engineering through projects — every lesson is paired with a hands-on lab you run on your own machine with real tools.
Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.
All courses · Pricing · About · FAQ · Glossary