- MantraMindAI
- Courses
- AI Career Readiness
- Transformer Architecture Q&A
Transformer Architecture Q&A
For engineers preparing for ML, LLM and AI-engineering interviews who need to explain how Transformers work, not just use them. You will be able to answer questions on attention, masking, positional encodings, normalisation, feed-forward blocks, the KV cache and GPT-style models, working the key calculations by hand.
Course Content
Prepares you to explain a Transformer forward pass end to end and to work scaled dot-product attention by hand, from Q, K and V to the final weighted sum.
Prepares you to explain how masks control what each token can see and how sinusoidal, learned, RoPE and ALiBi positions tell a Transformer about word order.
Prepares you to compare encoder, decoder and encoder-decoder models and to explain what residuals, normalisation and the feed-forward block (GELU, SwiGLU) each contribute inside a Transformer layer.
Prepares you to explain why LLM inference is slow and memory-hungry — the KV cache, prefill versus decode, context limits — and how GQA, FlashAttention, sparse and ring attention, and mixture-of-experts reduce that cost.
Prepares you to explain what happens inside a GPT attention layer — heads and their sizes, Q/K/V and the output projection, masking, softmax and dropout — well enough to spot a bug in someone else's implementation.
Prepares you to trace a GPT model from token ids through embeddings, blocks and the LM head to generated text, and to do the parameter, shape and decoding reasoning interviewers ask for.