And this is inherent to how LLMs work.
And this is inherent to how LLMs work.
Having a natural language interface where you need to go out of your way to specify that you want an accurate answer rather than just a plausible one defeats the entire purpose of it being a natural language interface for normal people. In certain professional contexts, it can be useful, but I don't buy it at all that it makes sense to ask everyone in their everyday lives to go out of their way to specify that they actually want correct answers to their questions.
You don't. It goes in the system prompt.
Take that away and you'll barely be able to make an app that display a pigeon riding a bicycle (or whatever you ppl are doing these days).
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.