We didn't notice the misses at first, because it's what we expected to begin with, and we very strongly noticed the hits because they were unexpected. Now we notice the misses and expect the hits.
We didn't notice the misses at first, because it's what we expected to begin with, and we very strongly noticed the hits because they were unexpected. Now we notice the misses and expect the hits.
It's no coincidence the flat-fee service is visibly crippled, while per-request API users are not reporting any difference. (edited)
Ignoring everyone here, look at the other "hacker" groups-- like the jailbreaking community. They have a lot to say about recent changes that coincide with their hacks not working. All of a sudden, with OpenAI supposedly changing nothing, technical bypasses just stopped working. OpenAI changed nothing, so this must be deus ex machina.
I'm not even jailbreaking it but the results I get for simple code requests through the web UI have become unusable garbage. It puts less effort into responses than an unpaid-and-overworked intern. Others here report the same. The current "iteration" seems hellbent on terminating conversations as quickly as possible once they stray from the explicit scope of the original topic and is almost as hostile to fixing its own errors as it is to endorsing eugenics, whereas in the beginning it would humor every idle thought I threw at it in long conversations. You can literally see this reflected in the logs they forced retention of. It's acting like a customer support rep desperately trying to end a call before it exceeds a call-time quota.
This particular current workflow seems like it lends itself to better organization of training data-- conversations are what the title says they're about. It also seems like it lends itself to anti-jailbreaking because pretexting it with irrelevant information forces a change of scope-- and a summary termination.
But OpenAI says they changed nothing, so rather than one guy lying without consequence, a community of professionals using and abusing the tool must all be victims of rhetorical fallacy? Nobody's qualified to reverse-engineer corporate bullshit anymore without being infantilized...
(Liars running a black box-- what could possibly go wrong? We need regulation!)
I thought GPT-4 wasn't available to free users?
Keep in mind the tweet is specifically about the OpenAI API - They might be updating ChatGPT without telling anyone (although they have release notes)
(edit: even that's not right. I think the verbiage is more accurate now.)
GPT-4 is not available to free users, which really calls the pretext of this comment into question. What is being alleged to have happened, and why is this response unreasonable?
I’m not a very heavy GPT-4 user but I do usually use it once a day or so - it doesn’t appear to have changed noticeably but I’m not paying close attention.
I dont really doubt they scaled the chatGPT4 model down a little to try to save costs with the plugins and increased usage
It’s noticeable because I used to be able to read each word as it was printed from the GPt-4 output, but now it goes far too fast for me to keep up with.
Kinda sucks to not have transparency at all into the black box, they’ve strayed so far from “open” at this point that it’s comedic
So many people are noticing a degradation in speed and quality of responses, and not in isolation. Rather than acknowledge this, the question is rephrased to one the responder can rightfully deny-- and suggest no changes have taken place without explicitly saying as much. You normally only see this sort of sliminess from politicians and executives on the witness stand.
> Is anyone else noticing significantly downgraded GPT-4 capabilities today? Seems like OpenAI updated the model, and results aren’t as good as before. [mentions API in a child comment]
> The API does not just change without us telling you. The models are static there.
> This is good to know. That means GPT-4 has been static since March right? 0314?
> Correct
Never ask questions to which one word suffices as an answer.
Collective confusion in the thread suggests something has changed, but the most OpenAI will attest to is that the API is unchanged and the models are static. And this may well be true, but rather than admit "...but we were fucking with the middleware/parameters" they took a firm position on a strawman argument and ignored everybody who followed with more-direct questions. Except this guy:
> I've noticed inconsistency with certain prompts performance. Is that just the non-deterministic nature of the API?
> Yes
Oh, ok. It's because the fucking API is non-deterministic that code that has worked both reliably and predictably for everyone now runs like shit for everyone. For fuck's sake, you can get better answers from a Magic 8-Ball. This guy even made the mistake of presenting an answer he'd believe for the respondent to feed into. He might as well have asked if inconsistent performance was because of the war in Ukraine.
"Logan.GPT" must moonlight as a fortune teller. He's only responding to people foolish enough to ask the wrong questions.
Here's the thing. Re-read the exchange you quoted:
>> The API does not just change without us telling you. The models are static there.
>> This is good to know. That means GPT-4 has been static since March right? 0314?
>> Correct
Does that "Correct" mean all GPT-4 models have been static since March, or does it only cover the gpt-4-0314, which is a single, specific model? gpt-4-0314 is the static snapshot model, hence 0314 in the name; it exist as a stable base, while gpt-4 was intended to be updated over time.
So I feel OpenAI may be dodging here. That "Correct" may just mean "the gpt-4-0314 model was not updated since March 14", which is, like, the very reason this model exists in the first place.
Important point: if you're using API pay-as-you-go access, unless you wrote the very tools you use with that API, you're most likely using gpt-4, and not gpt-4-0314.
Do you know of any alternative ChatGPT UIs like chatbotui.com or typingmind.com that let you specify the model gpt-4-0314 ?
This feature alone would be enough for me to switch.
"liars running a black box" is the definition of government regulation sheesh.
> How about matching 'a' as the second character of a string only?
It responsed with the wrong regex plus a bunch of explanatory junk:
> '^a.'
Then halfway through the explanatory junk, it corrected itself like this:
> Apologies for the confusion in the first response, the correct regular expression should be '^.a' for matching 'a' as the second character of a string:
And kept on with the (now correct) explanatory junk.
All in a single response. I've certainly never seen that before (if someone has, please weigh in). Maybe the model hasn't changed, but the pipeline has? Like... there's a second model trying to correct the mistakes of the first, maybe? (timings are probably wrong for that, but something like that)
With humans there is usually a non-verbal signal letting the other person know you changed topics.
This happens all the time in a social group setting though. Not much confusion ensues.
Sometimes I spend several minutes of a conversation trying to figure out which thing she's talking about, because we've already covered like twenty different topics in the last ten minutes and she seems to just switch around at random and with no warning.
Sometimes I get enough context that I can make that swap or stack pop, but many times not.
But yah, this is exactly what I mean - without some cue, neither humans, nor machines, can tell when you switch conversation.
> You can continue within the same conversation for a new topic. There's no need to start a new conversation. Feel free to ask about a different topic, and I'll do my best to assist you.
Unfortunately though, it would start over with the script completely, and then get stuck in a loop until it broke, this not being able to even save the response at all.
Something definitely feels like it changed, but I suppose it could just be more use of the system.
Another possibility is related to some strange issues with hardware and balancing etc, despite not changing software or params. It's strange, and it shouldn't happen, but sometimes things change from simple batching and balancing, or lower level hardware related things, which are very annoying to debug.
But then, all the more power to the open-source models and UIs, which are unrestricted, free to use, have better UX, and are constantly improving [1]. Ergo the suspicion there might be something else behind the desire to regulate it by OpenAI. Though hopefully we’re all just collectively cynical and they have good intentions after all, despite the misalignment in incentives.
I do think academics at universities are not to blame for this though, they are just as much flabbergasted and/or mislead as everyone else.
In any case, rather than wasting our time and energy on (fighting) ChatGPT(Plus) or GPT4, we should* all be collectively contributing to improving the open source models, by using them, reporting bugs, contributing ideas etc. This is the only way that we will continue to have leverage, and companies like OpenAI would be forced to open up, if they want to stay at least somewhat competitive. I think their closed-sourceness is a very short-sighted decision atm.
*Should because I don’t think LLMs as technology pose any substantial extinction risk to humanity, despite the obviously marketing rhetoric. Should that somehow change, maybe we shouldn’t. But, I don’t see any technological way to stop their improvement now that the genie is already out of the bottle. Heck, we might have a better chance of developing new tech for eventually enabling AGI decades from now, but only if we do that collectively and openly - only then everyone will have a level field and no side will be more powerful than another.
[1] https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
FWIW, I am a per-request API user, and... it could be my imagination, but I've had a rather clear feeling GPT-4 got much more lazy out of the sudden, both on OpenAI platform and on Azure, and it fits the timeframe of the recent compaints... so, n=1, make of it what you will.
Poetry.
Edit: Having recently gotten access to plugins and the browsing mode, I have noticed that the quality of the response is better when not using either.
If the model (gpt4) is unchanged, but they tweaked their system prompt for it on the chat interface, this is what you would see.
The first flight is magic, the nth one is a chore.
The fat of dead pigs, cattle and chickens is being used to make greener jet fuel
We know MS Bing Sydney is doing at least 3 of those (prompt, cascade with Megatron, and finetuned rejection classifier for post-output filtering) on top of its GPT-4-finetune, so it's not a stretch to figure that OA is doing similar things.
Note that API users using Bring Your Own Key tools/chat frontends probably default to gpt-4, and not the pinned gpt-4-0314.
https://twitter.com/OfficialLoganK/status/166447660465806951...
I think the word you're thinking of is the noun "gleam" which means a kind of lustrous shine, rather than the verb "glean" which means to harvest the remainder of something or to collect in small parts.
ChatGPT may produce inaccurate information about people, places, or facts. ChatGPT May 24 Version
I was shocked that an employee was claiming they haven’t changed the model since March, given that this date changes every few weeks. But I finally read the link it goes to, and apparently this date is just the frontend version. Which is a really weird thing to include as part of a disclaimer.
I wonder how many other people think this is the most recent model release date.
There's obviously a ChatGPT-3.5 and a ChatGPT-4. Actually 4 different ChatGPT-4's: Normal, Browsing, Plugins, Code Interpreter.
The audience of HN is, as Taleb would say, intellectuals yet idiots. They have trouble measuring change and are prone to hyperboles. Some still try to minimize the impact ChatGPT will have, while focusing on bullshit like 'hallucinations', or nitpicking about the quality of the code and so on. Can't see the forest from the trees.
If you're looking for intelligent discussion, look elsewhere.