With the simpler prompt, all the answers were wrong, most of them ridiculously wrong.
With the simpler prompt, all the answers were wrong, most of them ridiculously wrong.
Ultimately I feel it is fairer to benchmark llm’s by what they can be prompted into. After all, we let people carefully work through a problem during exams so it seems fair to hold llm’s to the same standard.
Oh wait, forgot something:
Think it through step by step.
Phew, close one.
I keep seeing comments and posts on HN that significantly downplay GPT-4's capabilities. Are people actually using GPT-4 or are they using a 3rd party service that claims to be GPT-4?
I got:
>Sally has 3 brothers, and each of those brothers has 2 sisters. One of those sisters is Sally herself, and the other one is Sally's sister. So, Sally has 1 sister.
> Sally has 2 sisters. Each of her 3 brothers has 2 sisters, and those sisters would be Sally and her 2 sisters.