I suppose this is similar to humans and probably why my school teachers always told me to show my work, but I'm curious if this has been documented and if there are any explanations for why it works this way with LLMs.
I suppose this is similar to humans and probably why my school teachers always told me to show my work, but I'm curious if this has been documented and if there are any explanations for why it works this way with LLMs.
By their very nature they only "know" what they have written down and must infer the final answer from that token by token.
They fundamentally can't do certain things such as complex iteration or track back.
When you ask for chain of thought thinking, you allow the LLM to create a "buffer space" and break down the task into more manageable substeps thereby improving the quality of the results.
We know this because it happily told us, including the json format it uses internally.
When using gpt-4 directly through the API we can emulate this behavior
https://www.make-safe-ai.com/is-bing-chat-safe/Prompts_Conve...
Really, just asking again is a fine way to expose all sorts of "hallucinations" in a LM.
First you wrap the user query with "the user asked you: ... . What are the reasoning steps you need?" and then you prompt with "considering `<previous answer>` now answer <user prompt>"
Obviously this is clearly hackable so it would need improvements.
Start at 7:30 to see example of backtracking.
If the model makes some mistake in the beginning, it now needs to explain / make sense of that mistake.
Kind of like a split-brain patient whom you ask why they got up, and they then say, to get a Coke. [1] In psychology, that is called confabulation. In machine learning, they use “hallucination“, probably so they can use the term across several disciplines, like language, audio, vision, etc.
[1] https://www.brainscape.com/flashcards/chapter-4-hemispheric-...
Just because a theory sounds nice and seems to make sense doesn’t mean it’s scientific
Video: https://youtu.be/qbIk7-JPB2c