When do you need chain-of-thought prompting for ChatGPT?
arxiv.org
arxiv.org
I imagine it like handing the LLM a piece of scratch paper.
...Hence, it is plausible that ChatGPT has already been trained on these tasks with CoT and thus memorized the instruction so it implicitly follows such an instruction when applied to the same queries, even without CoT. Our analysis reflects a potential risk of overfitting/bias toward instructions introduced in IFT, which becomes more common in training LLMs..."
Also I'm thinking would be interesting to train a LoRa, using synthetic prompts.
What may be here to stay might be more task-specific prompting to break a problem down. For example, a variation might be "Let's think step by step, making sure we do or consider X before we do Y".
By way of analogy, when my toddler was really young, if he couldn't figure something out, I'd ask him to break it down into steps (very similar to general Cot prompting). Now, he rarely needs that, but still needs more occasional, task-specific "prompting", eg "Let's think about the weather, and what we're going to be doing, before deciding what to wear".
Maybe the LLM gets so smart that it doesn't need to do this, but that I would definitely want to call a super-intelligence, because I sure can't solve big problems without breaking them down into smaller parts.
But maybe you are on to something when you say there are some things that we've just done so much that we don't really have to think them through to accomplish them anymore.
I don't know what that would say about an AI that was smart enough it could just spit out the right answer to literally any question we could devise. That would be quite a feat of engineering.
It’s like multiplication tables for thought. Using the “what to wear” example, you’ve essentially already just memorized what to wear for every occasion so you rarely need to think about it. Eg you know what is work appropriate, and if it’s raining or snowing or hot you’re still fine. But a wedding for a foreign culture you’re not familiar with in a city with weather you’re not familiar with, and you might not immediately know what to wear, so you have to break it down.
Even if they are "stochastic parrots" that are no more effective at elucidating truth than Ouija boards, they are still better than the autocorrect on my phone right now.
I’ve been explaining it to people as “if ChatGPT can think, it can only think out loud”. Because of the way LLMs were trained, nudging it to do this will continue to be valuable.
Most writing about anything difficult is product, not process. Articles get drafts before being published. People think about answers before writing them down. How to Solve It does a great job explaining this about math problems. The steps to the proof are not the steps to creating the proof.
So when you go to solve a problem by mimicking the solutions to problems, something is missing. People will either need to train LLMs to do it, or continue to ask for it in prompts.
Bing's free one is better in that sense in that you can use all 200 over a shorter timespan.