AFAIK, OpenAI has repeatedly stated that GPT4 hasn't changed. People repeatedly states that when they use ChatGPT, they get a difference experience today than before. Both can be true at the same time, as ChatGPT is a "packaged" experience of GPT4, so if you use the API versions, nothing has likely changed. But ChatGPT has guaranteed changed, for better or worse, as that's "just" integration work rather than fundamental changes to the model.
In the discussions on HN, people tend to speak each other regarding this as well, saying things like "GPT4 has for sure changed" when their only experience of GPT4 is via ChatGPT, which has changed since launch, obviously.
But ChatGPT != GPT4, which could always be made clearer.
It might have seem worth it to them but not the end user.
More guardrails on sensitive answers, fine.
But respect a user that explicitly (literally and figuratively) requests jailbreak and a specific type of response.
Working with the customer in dev: "Ok, run this SQL query and restart the service. Done, ok does the test case pass?" Done in 15 minutes.
Working with customer in production: "Ok, here is a 35 point checklist of what's needed to run the SQL query and restart the service. Have your compliance officer check it and get VP approval, then we'll run implementation testing and verification" --same query and restart now takes 6 hours.
The next day, suddenly prompts that used to work now gave much more generic results, the code was much more skinflinty and the kept trying to 'no wait I'm going to leave that long code as an exercise for you human'.
I didn't have time to buy in to a hallucination, I wasn't involved in openai chats to get 'infected by hysteria' or whatever, I was just using the tool a ton. and there was a noticeable change on day N+1 that has persisted until now.
The fact that gpt4 API calls appear to be similar tells me they changed their hidden meta prompt on the chatgpt plus website backend and are not admitting that they adjusted the meta prompt or other settings on the interface middleware between the JS webpage we users see and the actual gpt4 models running.
Isn't the thread about ChatGPT? I mean it is helpful to know that they are not the same (I personally was not clear on this myself, so I, at least, benefitted from your comment), but I think the thread is just about Chat GPT.
I doubt that. I don't recall them actually clearly and precisely saying they aren't changing the 'gpt-4' model - i.e. the model you're getting when specifying 'gpt-4' in an API call. That one direct tweet I recall, which I think you're referring to, could be read more narrowly as saying the pinned versions didn't change.
That is, if you issue calls against 'gpt-4-0314', then indeed nothing changed since its release. But with calls against 'gpt-4', anything goes.
This would be consistent with their documentation and overall deployment model: the whole reason behind the split between versioned (e.g. 'gpt-4-0314', 'gpt-4-0613') and unversioned models (e.g. 'gpt-4') was so that you could have both stable base and a changing tip. If that tweet is to be read as saying 'gpt-4' didn't change since release, then the whole thing with versioning is kind of redundant.
The -0613 version is really different! It added function calling to the API as a hint to the LLM, and in my experience if you don't use function calling it's significantly worse at code-like tasks, but if you do use it, it's roughly equivalent or better when it calls your function.
Maybe I lack the imagination, but what function should I give to the LLM? "insert(text: string)"?
For generating arbitrary code, I imagine you could do the same thing but swap `query_db` with the name `exec_javascript` or something similar based on your preferred language.
a. Worse at general code-like tasks without using functions
b. Equivalent or better at code-like tasks if you use the function API
c. Much faster than the older model either way.
I'd guess it's cheaper to run, too, and that they use the presence of a function in the API signature to weight their mixture of experts differently (and cull some experts?). The degradation in general purpose coding tasks is pretty obvious and repeatable (try the same prompts in the Playground with the -0314 model vs the -0613!), but it does seem like you can regain that lost capability with the new function call API, and it's faster. The tradeoff is that you only regain the capability when it calls functions; you can't really have a mix of prose-and-code in the same response as easily, or at least not with the same quality.
1: https://twitter.com/reissbaker/status/1671361372092010497
I ask, because I think it’s going to be a big challenge, so I built a service to record feedback / acceptance data: https://modelgymai.com/
If you think it can help, I’d love if you’d try it out and let me know if it helps.
I'm certain it's now doing significantly worse on the same tests, but alas I have lost the historical data to prove it.
Since 30 June, the API responses are making common English misspelling errors, of the type where two words sound the same with different meanings such as break and brake.
I saw this happen zero times in the prior GPT-4 model, and multiple times this July, on multiple conversation topics and multiple word pairs.
Curiously, they're behaving as misspellings rather than mismeanings, since the sentence continuation is as if the correct meaning had been used.
I acknowledge this could be a blend of pareidolia and the Baader-Meinhof phenomenon.
GPT-4 used to write with consistent under-grad level quality. Now it's closer to a junior high school kid.
We also can never independently evaluate it, OpenAI could cache messages, fine tune on public test sets, etc. etc.
That is entirely not equally likely, and would be completely unprecedented, at the frontier of an emerging technology that people are pumping the money and the future of the world into to win.
Combined with stricter guardrails, I would certainly expect intelligence to go down.
Check out the tikz unicorn drawing example from the paper "Sparks of Artificial General Intelligence: Early experiments with GPT-4" (pdf page 7)
I have used the same prompts to design a infinite scroll up/down image container.
2 months ago chat when it first came out was able to generate working code that used intersection observer API.
Now when I use the same prompts it only generates high level suggestions and when I ask for code it suggests using dom scroll events and doesn't even come up with intersection observer API unless I specifically ask. And if I do, it then generates incorrect code.
It even previously was correctly memoizing certain functions and included performance optimizations.
It’s just not as robust and general as it felt when using it for the first time.
That being said, the safety filters have definitively changed in OpenAI. ChatGPT is definitely more prone to reminding me that it is an LLM, and it refuses to participate in pretend play which it perceives as violating its safety filters. As a trivial example, ChatGPT is less willing to generate test cases for security vulnerabilities now - or engage in speculative mathematical discussions. Instead it will simply state that it is an advanced LLM blah blah blah.
I started using it relatively late, but earlier in May, you could have given it a DOI link, and it would have summarized it for you. Now, it argues that it's not a database and that it can only summarize it if you provide the full text. However, if you ask for it with the title of the paper, it will provide you with a summary.
You could have also asked it to search patents on some topic, and it would have given you a list of links. Now, it provides instructions on how to find it yourself.
My theory is that the initial ChatGPT offering (3.5/4/whatever) was "too hot" for the likes of certain incumbents. In my experience, the capabilities at launch were incredible and clearly a threat for a wide range of F500 software firms. I had phone calls with people I haven't talked to in over a decade about what I was seeing. I am not seeing those things today. This was mere months ago. This is not nostalgia.
Come to think of it, it must be the case, because the alternative would be pretty much every player on the market taking the hit and carrying on, or pretending they don't see the untapped value source that just freely flows out of OpenAI for anyone to enjoy, for a modest fee.
As a prime example, I'd point out Microsoft and their various copilots - the code one, the Office 365 one, the Windows system-wide one, in varying stages of development. API access to GPT-4 as good as it originally was[1], directly devalues all of those.
It stands to reason that slowly making the model dumber, while also making it faster and cheaper to use, is the best way for OpenAI to safeguard big players' markets - the "faster" and "cheaper" give perfect cover, while the overall effect is salting the entire space of possibilities - making the model good enough to entertain the crowd, but just not good enough to build solutions on top, not unless you're working for one of the players with special deals.
TL;DR: too many entities with money were unhappy about all the value OpenAI was giving to the world for peanuts, so the model is being gradually nerfed in a way that allows that value to be captured, controlled, and doled out for a hefty price.
(And if that turns out to be true, I'm going to be really pissed. I guess it's in the style of humanity to slow down pace of development not because of ideology, not because of potential risks, but because it's growing too fast to fully monetize.)
--
[0] - I mean that in the most nasty, parasitic sense possible.
[1] - I'm talking about the public release. That GPT-4 version seems to have already been weakened compared to pre-"safety tuning" GPT-4 (see the TikZ Unicorn benchmark story), but we can't really talk about what we never got to play with.
It just seemed obvious that if anyone suggested a use case that was actually really high value MS would just take the idea, run with it for a month or two to see if it has legs, and then steal it if it actually worked.
All while you're waiting in the queue to have your idea validated as "safe".
Can you provide some evidence to back that up? Especially because OpenAI _has_ been tinkering with ChatGPT - by trying to limit jailbreaks.
People have a strong prior that these kinds of changes will reduce model performance (because you're limiting your model), so the burden is on you to show that performance hasn't degraded.
Maybe that had some unexpected side-effects.
Now it can't even do just the caeser cipher without hallucinating nor can do it do even purely base64 decoding without hallucinating.
Here's the best fish I could make today: https://www.svgviewer.dev/s/3IuulHlC
Make of that what you will.
While GPT-4 is still workable, GPT-3.5 flatly refuses requests these days, claiming that as an "AI language model" it couldn't help me write code.
TBH, ChatGPT 3.5 has intermittently given me such responses from dy 1.
Then you get used to this new level of capability and subconsciously weight the errors more.
For all the talk, I see very few people sharing direct chat links that are the same query at different points in time with different quality of answer.
In fact, when I do similar things, I don't notice a change in quality.
could this be just A/B testing? Do the terms of use rule out tweaking inference parameters (temperature?) even if using the same model?
Changes happen at many layers, ChatGPT UI, API Gateways, moderation API, backend server hosting model, the model file itself, etc.
Each of these components is changing pretty regularly it seem.
The end result of combined changes for users is being the observed degraded performance of ChatGPT.
1. It's software that offers non-deterministic output and as such is fiendishly difficult to write realistic end-to-end tests for. Of course it's experiencing regressions. Heisenbugs are the hardest bugs to catch and fix, but having millions of users will reliably uncover them. And for an LLM, almost every bug is a Heisenbug! What if OpenAI improved GPT4 on one metric and this "nerfed" it on some other more important metric? That's just a classic regression. And what would robust, realistic end-to-end tests even look like for GPT4?
2. It's software that presents itself as a human on the Internet—even worse, a human representing an institution. Of course nobody trusts it. Everyone is extremely mistrustful of the intents and motivations of other humans on the Internet, especially if those humans represent an organization. I co-ran the tiny activist nonprofit Fight for the Future for years, and it was really amazing how common it was for comments in online spaces to assume the worst intentions; I learned to expect it and react extremely patiently. Imagine what it's like for OpenAI, building a product that has become central to peoples' workflows. Of course people are paranoid and think they're the devil, and are able to hallucinate all manner of offense and model it with every paranoid theory imaginable. The funny thing is, the more successful GPT4 is at seeming human, the less some people will trust it, because they don't trust humans! And the smarter and more successful it gets, the less some people will trust it! (How much do most people trust smart, successful public figures?)
3. Maybe an overall improvement for most users (one that the data would strongly suggest is a valid change and that would pass all tests) is a regression for some smaller set of users that aren't expressed in the tests. There might be some pairs of objectives that still present genuinely zero-sum tradeoffs given the size of the model and how it's built. What then? The usefulness of GPT4 is specifically that it is general purpose, i.e. that the massive cost of training it can be amortized across tons of different use cases. But intuitively there must be limits to this, where optimization for some cases comes at a cost to others, beyond the oft-cited of Bowdlerization. Maybe an LLM is just yet another case in the real world where sharing an important resource with lots of people is a hard problem.
If I were at OpenAI, I would want some third party running a community-submitted end-to-end test suite on each new release, with accounts that were secret to OpenAI and from unknown IP addresses—via Tor Snowflake bridges or something.
It's so tempting when running into user-reported Heisenbugs to trick oneself into ignoring users and not accepting that you've shipped a real regression. In addition to wanting the world to know, I would want to know.
But there's a real question of what these community-curated tests would even be, since they'd have to be automated but objective enough to matter. Maybe GPT4 answers could be rated by an open source LLM run by a trusted entity, set to temperature: 0? Or maybe some tests could have unambiguous single-string answers, without optimizing for something unrealistic? And the tests would have to be secret or OpenAI could just finetune to the tests. It's tricky, right?
Some people even make it their life's work.
The same exact prompts from February are leading to significantly degraded responses. I've proven it to myself dozens of times.
It's not surprising, they were probably running the model at a huge and ultimately unacceptable loss. But they should really offer a higher paid tier to access the previous capabilities... not drop them entirely. Many would pay far more than $20/month to access a marginally but meaningfully better model.
EDIT: Many being dismissive of LLMs don't even seem to use them. Providers are vastly overvalued from an investment perspective, but the utility is very real. To say the loss in capability is just an "illusion" is clearly wrong to anybody who actually uses it.