There are two conclusions I took from scanning through this and trying to reproduce a few of the reported failures.
1. The author is bad at prompting. There are many ways to reduce hallucinations and provoke better thinking paths for the model.
2. The author is using ChatGPT's GPT-4, leading him to conflate "GPT-4" with "ChatGPT". While you can consider this a shared failure with OpenAI, due to OpenAI's poor communication, anybody doing serious work evaluating these models would know that the first thing you need to do is use the API and pin the model version. In the author's case, he should have used gpt-4-0314 or gpt-4-0613. What I suspect he did is that he just used ChatGPT's GPT-4, and likely the default model at that. (Nobody should ever use the Default model. It's their most heavily performance optimized model and performs worse on reasoning tasks than the Plugins model, even on within-context-size tasks.)
There are huge problems with that, because OpenAI has done both a ton of fine tuning and performance optimization continuously on the default ChatGPT model over time that its performance has ranged anywhere from "I'm pretty sure this is gpt-3.5" to "whoa, this is damn good" (the latter being mostly the model at launch, which was probably the same as gpt-4-0314).
If the author has been working seriously at evaluating models, specifying the model is the first thing he'd do. Perhaps he should explain his reasoning.