ChatGPT interfaces with semantics, and not logic.
--
That means that any emergent behavior that appears logically sound is only an artifact of the logical soundness of its training data. It can only echo reason.
The trouble is, it can't choose which reason to echo! The entire purpose of ChatGPT is to disambiguate, but it will always do so by choosing the most semantically popular result.
It just so happens that the overwhelming majority of semantic relationships also happen to be logical relationships. That's an emergent effect of the fact that we are usually using words to express logic. So if you mimic human speech well enough to look semantically interesting, you are guaranteed to also appear logically sound.
--
I don't see any way to seed such a system to always produce logically correct results. You could feed it every correct statement about every subject, but as soon as you merge two subjects, you are right back to gambling semantics as logic.
I also don't see a scalable way to filter the output to be logically sound every time, because that would be like brute-forcing a hash table.
OP considers something in the middle, but that's still pretty messy. They essentially want a dialogue between ChatGPT and WolphramAlpha, but that depends entirely on how logically sound the questions generated by ChatGPT are, before they are sent to WolphramAlpha. It also depends on how capable WolphramAlpha was at parsing them.
But we already know that ChatGPT is prone to semantic off-by-one errors, so we already know that ChatGPT is incapable of generating logically sound questions.
--
As I see it, there is clearly no way to advance ChatGPT into anything more than it is today. Impressive as it is, the curtain is wide open for all to see, and the art can be viewed plainly as what it truly is: magic, and nothing more.