Subscribe
EncodeEdge LabsInteractive Engineering & Research Hub

Interactive AI Simulators & Notation Decoder

Explore deep learning mechanics, spatial convolutions, transformer attention heatmaps, and mathematical notation with real-time visual controls.

Select Simulation Lab:Available in browser & within course lessons
Deep LearningLab ID: attention-visualizer

Transformer Self-Attention & Heatmap Visualizer

View inside lesson: Scaled Dot-Product Attention

Transformer Self-Attention & Heatmap Visualizer

Transformers & Attention

Explore Query-Key dot products, multi-head attention patterns, and the impact of softmax temperature scaling.

Click any token below to view its incoming/outgoing attention weights
Active Query Token:"The"(index #0)Weights sum to 1.00 (Softmax normalized)
Softmax Temperature ($\tau$):1.00Balanced Sampling
0.2 (Sharp)1.0 (Standard)2.5 (Diffuse)

Full Attention Matrix ($N \times N$)

Rows: Query Tokens ($Q$) | Columns: Key Tokens ($K$)
Q \ K
The
animal
did
not
cross
the
street
because
it
was
15%
11%
11%
9%
8%
7%
8%
10%
11%
10%
11%
15%
9%
8%
8%
10%
11%
11%
9%
8%
10%
8%
16%
9%
11%
12%
10%
9%
8%
9%
8%
8%
10%
16%
11%
9%
8%
8%
10%
11%
8%
10%
11%
10%
15%
7%
8%
10%
11%
10%
11%
11%
9%
8%
8%
15%
11%
11%
9%
8%
11%
9%
8%
9%
11%
12%
16%
9%
8%
8%
8%
8%
9%
11%
11%
10%
8%
15%
9%
11%
5%
39%
7%
7%
6%
5%
7%
6%
10%
7%
11%
11%
9%
8%
7%
8%
10%
11%
10%
15%
Mathematical Attention Formulation:

Each Query token qᵢ projects a dot product onto all Key tokens kⱼ, divided by scaling factor √dₖ = 8.0. Softmax normalizes the raw logits into positive probabilities that sum strictly to 1.0 across every row, weighting which semantic context vectors are extracted into the output representation.