Simplest thing GPT4 can’t do?
lesswrong.com
lesswrong.com
Perhaps any tasks that require a great deal of dynamic memory/comparison would be expected to struggle in this architecture.
Since "write down all possible permutations of these letters" is like n! "cross reference the permutations with a dictionary/definition if you feel like it" must be performed m*n times, maybe other such n! tasks are doomed until more parameters are feasible.
It almost certainly could produce a bash script to do that based on what I've seen.
That should be prompted, but it's lazy about sorting through it's own outputs like this.
Say you explain the rules and it guesses, "pecks",
> P E _ _ _
> The word contains a C
> The word doesn't contain K, S
Re-promoting and accumulating those details like this, ChatGPT 3.5 managed to solve yesterday's wordle successfully, but it took 10 tries to land on this prompting method.
Maybe more parallel GPTs instructed like, "find 5-10 key facets of reasoning about the problem presented", and then "extract those facets from the conversation, organize them into a table, and use them to supplement decision-making" might help prevent the needed parameters from blowing up too much.
If you think of it as a statistically-derived processor/algorithm, unless there's some great task of short-term memory/memorization it's forced to undergo it's going to be biased towards more direct responses.
Another thing to consider, regarding the fine-tuning, is that "instruction" may be biased towards sentence ordering, so that the actual order with which you present your prompt can be important.
For example Prompt: <detailed situational specification>. <task to be performed> might perform considerably worse than <task to be performed>. <detailed situational specification>.
I suspect it's just a short-mid-term resource constraint though since:
- demonstrating extremely sharp memory performance conversationally isn't that common
- total bout / token width isn't particularly wide in training
- presumably fewer "transistors" than humans -- basically even 200B parameters isn't all that many
Another concern that I'm not sure how to consider is that 4-letter tokenization could be an obvious reason ChatGPT is such a poor speller.