I see this argument all the time. Why are you assuming that this technology just "stops" at the LLM level?
If I'm openAI or Google or whatever, I'm definitely going to run extra classifiers on top of the output of the LLM to determine & improve accuracy of results.
You can layer on all kinds of interesting models to make a thing that's generally useful & also truthful.