Subscribe
Applied Deep LearningTransformers & Self-Attention Architecture
Knowledge Checkpoint Quiz15 minutes

Transformers & Self-Attention Architecture

Lesson Summary: Test your mastery of Query-Key-Value projections and multi-head attention mechanics.
Checkpoint Score
0 / 2 Correct
Novice(0 XP)

Q1.In Scaled Dot-Product Attention, why are dot products scaled by 1/sqrt(d_k)?

mcq
To counteract large values pushing the softmax function into regions with tiny gradients
To ensure that the output tensor is strictly orthogonal
To reduce memory usage on CUDA device allocations
To convert logits into boolean masks

Q2.What enables Transformers to process all tokens in a prompt in parallel during training, unlike RNNs?

mcq
Self-attention operates across all token pairs simultaneously as matrix multiplications
Transformers do not use matrix multiplication
Hidden state is passed sequentially from left to right
Transformers discard word order entirely