A Recurrent Neural Network reads a sequence one element at a time, updating a hidden state that carries context forward — natural for text, audio, time series. The problem: over long sequences, gradients shrink (vanish) during backprop, so early context is forgotten and training stalls. This limitation motivated gated variants (LSTM/GRU) and, ultimately, attention.