ChatGPT was released in 2022! It doesn't feel like that, but it's been out for a long time and we've only seen marginal improvements since, and the wider public has simply not seen ANY improvement.
It's obvious that the technology has hit a brick wall and the farce which is to spend double the tokens to first come up with a plan and call that "reasoning" has not moved the needle either.
I build systems with GenAI at work daily in a FAANG, I use LLMs in the real world, not in benchmarks. There hasn't been any improvement since ChatGPT first release and equivalent models. We haven't even bothered upgrading to newer models because our evals show they don't perform better at all.
If nothing else, that technique has cut down drastically on hallucinations.
to a skilled user of a model, the model won't just make shit up.
Chatbots will of course answer unanswerable questions because they're still software. But why are you paying attention to software when you have the whole internet available to you? Are you dumb? You must be if you aren't on wikipedia right now. It's empowering to admit this. Say it with me: "i am so dumb wikipedia has no draw to me". If you can say this with a straight face, you're now equipped with everything you need to be a venture capitalist. You are now an employee of Y Combinator. Congratulations.
Sometimes you have to admit the questions you're asking are unlikely to be answered by the core training documents and you'll get garbled responses. confabulations. Adjust your queries accordingly. This is the answer to 99% of issues product engineers have with llms.
If you're regularly hitting random bullshit you're prompting it wrong. Models will only yield results if they get prompts they're already familiar with. Find a better model or ask better questions.
Of course, none of this is news to people who actually, regularly talk to other humans. This is just normal behavior. Hey maybe if you hit the software more it'll respond kindly! Too bad you can't abuse a model.
But that doesn't mean the warped slide rule and a super computer capable of finite element analysis are equally useful or powerful.
(Alas, Dall-E also lacks, so I couldn't generate a picture of deadly slide rule kung fu. At least none that wasn't unintentionally hilarious.)