Instead of one fixed context vector, attention lets the model, at each output step, compute a weighted blend of all input positions — focusing on the relevant words. It dissolved the seq2seq bottleneck and produced interpretable alignment (which input word maps to which output). Attention was the key idea that, scaled up and used alone, became the transformer.