But sometimes it makes up answers that are wrong but sound plausible to you. You have no way to tell when it does this, and neither does it.
It's just a matter of priorities for the company designing the models.
> Maybe have the network which generates the answer be moderated by another network that assesses the truthiness of it.
Like a GAN? Sometimes you can do that, but it seems not always.
If this was simple and obvious, they'd have done it as soon as the first one was interesting-but-wrong.