freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #42 of 120

seq2seq · encoder-decoder

NeuraVSThe Overfit Ogre
Neura saysSeq2seq uses an encoder to read input and a decoder to generate output — for translation.

Sequence-to-sequence models pair an encoder (compresses the input sequence into a context vector) with a decoder (generates the output sequence from it) — the original neural translation architecture. The bottleneck: cramming a whole sentence into one fixed vector lost information for long inputs. That bottleneck is exactly what attention was invented to fix.

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleExplain the information bottleneck in vanilla seq2seq.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>encoder: sentence → one context vector (bottleneck)
decoder: context vector → output words
long input → info lost → attention fixes it</pre></body></html>
▶ Open the interactive comic issue
‹ Grus · Simpler Gated RecurrenceAttention · The Seq2seq Fix ›