At least, for people who need large context windows, they would not be the first choice anymore.
At least, for people who need large context windows, they would not be the first choice anymore.
The ChatGPT equivalent is 3x speed and was somewhere between ChatGPT and GPT4 on my TriviaQA benchmark replication I did
Couple tweets with data and examples. Note they’re from 8 weeks ago, I know Claude got a version bump, GPT3.5/4 accessible via API seem the same.
[1] brief and graphical summary of speed and TriviaQA https://twitter.com/jpohhhh/status/1638362982131351552?s=46&...
[2] ad hoc side by sides https://twitter.com/jpohhhh/status/1637316127314305024?s=46&...
GPT3.5 just got an update a few days ago that resulted in a pretty good improvement on its creativity. I saved some sample outputs from the previous March model, and for the same prompt the difference is quite dramatic. Prose is much less formulaic overall.
Random Q, I don’t use the ChatGPT front end much past month or two, used it a week back and it seemed blazingly faster than my integration: do you have a sense of if it got faster too?
It can't even solve simplified versions of problems it had zero issues with just a week ago.
It’s worse at everything.
Impressions:
Bad enough compared to GPT-4 that I default to GPT-4. I think if I had api access I’d use it instead, right now it requires more coaxing, and using Poe.
I did find “long-term” chats went better, was really impressed with how it held up when I was asking it a nasty problem that was hard to even communicate verbally. Wrong at first, but as I conversed it was a real conversation.
GPT4 seems to circle a lower optima. My academic guess it’s what Anthropic calls it “sycophancy” in its papers, tldr GPT really really wants to do more like what’s in the context, so the longer the conversation with initial errors goes, it’s actually harder to talk it out of the errors.