models are still having trouble reasoning – current tricks still depend of feeding back an intermediate answer with a different prompt, but it is still an LLM, and most times the same LLM
I think it's always going to be a bit shit, no matter how large they make them it's always going to be an obvious faker, always going to confabulate and lie etc.
As long as we are using a technology that merely fakes intelligence (using statistics to generate the most likely next token based on training data is not intelligence), it is going to be obvious that it's fake. This whole bubble is going to burst when people start realizing how overrated it is.