Exactly. The translation back to words is the final step, so in a way very similar to what the post describes.
Improvements in model performance have been made exactly by having intermediate steps stay in the form of internal representations rather than words.