- MantraMindAI
- Courses
- Machine Learning & Data Science
- Attention Mechanisms and Transformers
Attention Mechanisms and Transformers
For engineers and data scientists who know basic neural networks and want to understand what actually happens inside a transformer. You will work through attention by hand, build and debug a small encoder-decoder in PyTorch, apply transformers to text and images, and fine-tune BERT and GPT models into a working classifier.
Course Content
Why recurrent models hit a wall, and how attention (worked through by hand, number by number) removes it. 3 lessons, about 40 minutes.
The full transformer piece by piece — encoder and decoder, positional encoding, residuals and normalisation, masking and optimisation — ending with a small encoder-decoder you build and debug yourself. 4 lessons, about 60 minutes.
How the same mechanism handles images (ViT) and language (BERT, GPT, T5), and how to fine-tune pre-trained models without destroying what they already know. 3 lessons, about 45 minutes.
Build a complete, honest sentiment classifier on IMDB — from data checks through fine-tuning, error analysis and a serving endpoint. 1 lesson, about 15 minutes.