Figures › Transformers

Self-Attention Matrix

N×N attention grid lighting up token-to-token relationships with softmax weighting.

FIG_003 · Animated diagram · DeVenture Academy

Terms in this diagram

Token · Softmax

More Transformers diagrams

BPE Tokenizer: Merge Rules in Action · Transformer Block Data Flow · KV Cache Growth During Generation

Learn the concept behind the diagram

This figure comes from the DeVenture Academy curriculum — project-based AI engineering courses where every lesson is paired with a hands-on lab you run on your own machine.

Create a free account — the opening phases of 24 of 30 courses are free, no credit card. Or see Pro pricing.

All courses · Pricing · About · FAQ · Glossary