Long Short-Term Memory
2 korrents from this paper
In plain words
Networks that process sequences lose the teaching signal over long delays; LSTM keeps that signal constant inside special memory units with gates, learning across more than 1000 steps. It gets many more successful runs and learns much faster than older methods, and solves complex long-delay tasks previous algorithms never managed. It proposes a new learning method built on an earlier analysis of why that signal fades, tested against several older sequence-network algorithms on artificial data.
Our summary of the paper, not the authors' words — written to be readable without the field's vocabulary, from the stored copy of the paper and nothing else. Drafted with xai:grok-4.5 and checked by a person. The authors' own sentences are the quotes below.
Near this, by wording
- Frankle & Carbin, "The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks" (arXiv)
- He, Zhang, Ren & Sun, "Deep Residual Learning for Image Recognition" (arXiv)
Papers whose claims are worded most like this one's, found by the same hourly pass that draws the map. It is a measure of LANGUAGE, not of agreement or of citation: two papers can be near each other here and flatly contradict one another.
Sepp Hochreiter, Jürgen Schmidhuber did not write this page.
Every claim below was made in this piece, quoted word for word and numbered in the order the piece makes them, so you can read it there rather than take our word for it. The sentence above each quote is our reading of the claim, not their wording. Each quote was checked against a stored copy of the page at build time; where the two differ, the quote is the fact.