A Transformer uses **self-attention** to process all tokens in a sequence simultaneously (unlike RNNs which process sequentially).
Key components:
1. **Embedding layer** — converts tokens to vectors
2. **Self-attention** — each token attends to all other tokens
3. **Multi-head attention** — multiple attention patterns in parallel
4. **Feed-forward layers** — process attended features
5. **Layer normalisation** — stabilises training
GPT, BERT, LLaMA, and Gemini are all based on transformers.