That's right - most literature does show that encoder-decoder architectures outperform CTC. I think one of the main reasons for this is that CTC assumes the label outputs are conditionally independent of each other, which is a pretty big flaw in that loss function.
The blog does mention Listen-Attend-Spell (which is an encoder-decoder architecture) as an alternative to the CTC model.