Subscribe
Applied Deep LearningScaled Dot-Product & Multi-Head Self-Attention
Video Lesson34 minutes

Scaled Dot-Product & Multi-Head Self-Attention

Lesson Summary: The math behind Query, Key, and Value projections, softmax temperature scaling, and attention masks.

Scaled Dot-Product Attention Implementation Lab

Deep LearningMatched to lesson

Compute Q, K, V dot-products, scale attention scores by sqrt(d_k), apply softmax normalization, and extract context vectors.

Labs:
Scaled Dot-Product Attention Implementation
Pyodide Wasm
Interactive Challenge: Modify code inputs, click Run Code to execute live in WebAssembly.
+25 XP Reward
Wasm Terminal Output

Click Run Code to execute this algorithm in the browser sandbox.

The Transformer architecture replaced recurrence with attention, allowing complete parallelization across entire context windows.

Mathematical Formulation

Given query matrix QQ, key matrix KK, and value matrix VV:

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V

The factor 1dk\frac{1}{\sqrt{d_k}} counteracts the tendency of dot products to grow large in high-dimensional spaces, which would otherwise push softmax into regions with vanishingly small gradients.

import torch
import torch.nn.functional as F

def scaled_dot_product_attention(Q, K, V, mask=None):
    d_k = Q.size(-1)
    # Compute attention scores
    scores = torch.matmul(Q, K.transpose(-2, -1)) / (d_k ** 0.5)
    
    if mask is not None:
        scores = scores.masked_fill(mask == 0, -1e9)
        
    weights = F.softmax(scores, dim=-1)
    output = torch.matmul(weights, V)
    return output, weights