Exactly. The premise was true years ago, definitely. But today it's not hard to get correct spelling out of the top models. Ask DALL-E 3 to make a picture with some text, and it will spit out 4 image. Usually 2 or 3 are perfectly spelled. Lesser or older diffusion models (whatever OP is using) sometimes mix cyrillic and latin letters, or invent plausible-looking letters that don't exist in any language. But think about how they work - they are trained to turn pure pixel noise into a more plausible array of pixels for the prompt. It's pretty close to plausible text - misspelling is a nitpick for what it's getting right. Technology progresses over time.