EncodeEdge LabsInteractive Engineering & Research Hub
Interactive AI Simulators & Notation Decoder
Explore deep learning mechanics, spatial convolutions, transformer attention heatmaps, and mathematical notation with real-time visual controls.
Select Simulation Lab:Available in browser & within course lessons
Deep LearningLab ID: attention-visualizer
Transformer Self-Attention & Heatmap Visualizer
Transformer Self-Attention & Heatmap Visualizer
Transformers & AttentionExplore Query-Key dot products, multi-head attention patterns, and the impact of softmax temperature scaling.
Click any token below to view its incoming/outgoing attention weights
Active Query Token:"The"(index #0)Weights sum to 1.00 (Softmax normalized)
Softmax Temperature ($\tau$):1.00Balanced Sampling
0.2 (Sharp)1.0 (Standard)2.5 (Diffuse)
Full Attention Matrix ($N \times N$)
Rows: Query Tokens ($Q$) | Columns: Key Tokens ($K$)Q \ K
The
animal
did
not
cross
the
street
because
it
was
15%
11%
11%
9%
8%
7%
8%
10%
11%
10%
11%
15%
9%
8%
8%
10%
11%
11%
9%
8%
10%
8%
16%
9%
11%
12%
10%
9%
8%
9%
8%
8%
10%
16%
11%
9%
8%
8%
10%
11%
8%
10%
11%
10%
15%
7%
8%
10%
11%
10%
11%
11%
9%
8%
8%
15%
11%
11%
9%
8%
11%
9%
8%
9%
11%
12%
16%
9%
8%
8%
8%
8%
9%
11%
11%
10%
8%
15%
9%
11%
5%
39%
7%
7%
6%
5%
7%
6%
10%
7%
11%
11%
9%
8%
7%
8%
10%
11%
10%
15%
Mathematical Attention Formulation:
Each Query token qᵢ projects a dot product onto all Key tokens kⱼ, divided by scaling factor √dₖ = 8.0. Softmax normalizes the raw logits into positive probabilities that sum strictly to 1.0 across every row, weighting which semantic context vectors are extracted into the output representation.
