Course Overview
Applied Deep LearningScaled Dot-Product & Multi-Head Self-Attention
0% Done
Video Lesson34 minutes
Scaled Dot-Product & Multi-Head Self-Attention
Lesson Summary: The math behind Query, Key, and Value projections, softmax temperature scaling, and attention masks.
Scaled Dot-Product Attention Implementation Lab
Deep LearningMatched to lessonCompute Q, K, V dot-products, scale attention scores by sqrt(d_k), apply softmax normalization, and extract context vectors.
Labs:
Scaled Dot-Product Attention Implementation
Pyodide Wasm
Interactive Challenge: Modify code inputs, click Run Code to execute live in WebAssembly.
Wasm Terminal Output
Click Run Code to execute this algorithm in the browser sandbox.
The Transformer architecture replaced recurrence with attention, allowing complete parallelization across entire context windows.
Mathematical Formulation
Given query matrix , key matrix , and value matrix :
The factor counteracts the tendency of dot products to grow large in high-dimensional spaces, which would otherwise push softmax into regions with vanishingly small gradients.
import torch
import torch.nn.functional as F
def scaled_dot_product_attention(Q, K, V, mask=None):
d_k = Q.size(-1)
# Compute attention scores
scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)
if mask is not None:
scores = scores.masked_fill(mask == 0, -1e9)
weights = F.softmax(scores, dim=-1)
output = torch.matmul(weights, V)
return output, weights 