To be fair, the claim wasn't that it always produced the wrong answer, just that there exists circumstances where it does. A pair of examples where it was correct hardly justifies a "demonstrably false" response.
If you want a more scientific answer there is this recent paper: https://machinelearning.apple.com/research/gsm-symbolic
Maybe some HN commenters will finally learn the value of uncertainty then.