It's not used a lot in ambiguous tasks like image recognition and audio because you would need to have a state for every single possible wave form or image variation. You would have to have a similarity preserving hash or some other related method to reduce related states to a single common state. That's why it's better to use them near the ends of DL networks. If you didn't do that, every image of the same apple with just one pixel different would be a new state.