freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #43 of 120

attention · the seq2seq fix

NeuraVSThe Overfit Ogre
Neura saysAttention lets the decoder look back at all input positions, weighted by relevance.

Instead of one fixed context vector, attention lets the model, at each output step, compute a weighted blend of all input positions — focusing on the relevant words. It dissolved the seq2seq bottleneck and produced interpretable alignment (which input word maps to which output). Attention was the key idea that, scaled up and used alone, became the transformer.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleExplain how attention fixes seq2seq’s bottleneck.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>each output word → weighted look at ALL input words
focus where relevant → no single bottleneck vector</pre></body></html>
▶ Open the interactive comic issue
‹ Seq2seq · Encoder-DecoderTransformers · Attention Is All You Need ›