Understanding LSTM Networks (2015)
colah.github.io
colah.github.io
Additionally, GRU may be simpler to understand memory unit.
From this article I can correctly and easily read off the model architecture it describes, as a composition of smooth maps on modules over the real numbers, which is more than what one can say about a lot of papers.
As you note, the main contribution of this article is the diagrams -- but I think it's a bit deeper than being pretty. Previously, diagrams of LSTMs had lots of loops in them. This made them hard to understand. (In fact, they were ambiguous because of unclear execution order.) The diagrams in this article unrolled the loops, which seems to be a much easier way to reason about LSTMs. They also work at the abstraction of layers instead of weights and matrix multiplies, which make the diagram less noisy and focuses the reader on the important ideas.
Since diagrams were previously challenging to consume and understand, explanations prior to my article tended to focus on the six equations than define an LSTM. While that's a relatively clear way to think about LSTMs once you've really internalized them, it can be very hard when you are trying to understand them for the first time. In particular, the equations introduce ~15 variables. If you aren't already familiar with those variables, it's extremely challenging to keep them all in your working memory and understand what's going on.
In contrast, the diagrams don't require a new reader to hold lots of things in working memory, and visually hint a great deal of the narrative (like the horizontal line at the top).