It's reasonable to call minimising the training loss function the "objective" of the LLM, I think (said loss functions are often called objective functions, after all), though we must careful that the overloading such term in an anthropomorphic fashion. Whether "producing clear and coherent text" is a good characterisation of a "predict the next token" loss function is another question though. I would say it probably isn't, but it's a resonable hypotheses to say that the nature of the training is why the output tends to bias towards plausible but incorrect text as opposed to some variation of "I don't know".