Long Short-Term Memory networks fix the vanishing-gradient problem with a cell state (a memory conveyor) and three gates (forget, input, output) that learn what to keep, add, and emit. This lets them retain context across hundreds of steps. LSTMs powered the pre-transformer era of translation and speech, and still appear where sequences are long but data/compute is limited.