OpenAI Employee: GPT-4 has been static since March
twitter.com
twitter.com
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately? - https://news.ycombinator.com/item?id=36134249 - May 2023 (711 comments)
We didn't notice the misses at first, because it's what we expected to begin with, and we very strongly noticed the hits because they were unexpected. Now we notice the misses and expect the hits.
The first flight is magic, the nth one is a chore.
The fat of dead pigs, cattle and chickens is being used to make greener jet fuel
The audience of HN is, as Taleb would say, intellectuals yet idiots. They have trouble measuring change and are prone to hyperboles. Some still try to minimize the impact ChatGPT will have, while focusing on bullshit like 'hallucinations', or nitpicking about the quality of the code and so on. Can't see the forest from the trees.
If you're looking for intelligent discussion, look elsewhere.
I think the word you're thinking of is the noun "gleam" which means a kind of lustrous shine, rather than the verb "glean" which means to harvest the remainder of something or to collect in small parts.
It's no coincidence the flat-fee service is visibly crippled, while per-request API users are not reporting any difference. (edited)
Ignoring everyone here, look at the other "hacker" groups-- like the jailbreaking community. They have a lot to say about recent changes that coincide with their hacks not working. All of a sudden, with OpenAI supposedly changing nothing, technical bypasses just stopped working. OpenAI changed nothing, so this must be deus ex machina.
I'm not even jailbreaking it but the results I get for simple code requests through the web UI have become unusable garbage. It puts less effort into responses than an unpaid-and-overworked intern. Others here report the same. The current "iteration" seems hellbent on terminating conversations as quickly as possible once they stray from the explicit scope of the original topic and is almost as hostile to fixing its own errors as it is to endorsing eugenics, whereas in the beginning it would humor every idle thought I threw at it in long conversations. You can literally see this reflected in the logs they forced retention of. It's acting like a customer support rep desperately trying to end a call before it exceeds a call-time quota.
This particular current workflow seems like it lends itself to better organization of training data-- conversations are what the title says they're about. It also seems like it lends itself to anti-jailbreaking because pretexting it with irrelevant information forces a change of scope-- and a summary termination.
But OpenAI says they changed nothing, so rather than one guy lying without consequence, a community of professionals using and abusing the tool must all be victims of rhetorical fallacy? Nobody's qualified to reverse-engineer corporate bullshit anymore without being infantilized...
(Liars running a black box-- what could possibly go wrong? We need regulation!)
I thought GPT-4 wasn't available to free users?
Keep in mind the tweet is specifically about the OpenAI API - They might be updating ChatGPT without telling anyone (although they have release notes)
(edit: even that's not right. I think the verbiage is more accurate now.)
GPT-4 is not available to free users, which really calls the pretext of this comment into question. What is being alleged to have happened, and why is this response unreasonable?
I’m not a very heavy GPT-4 user but I do usually use it once a day or so - it doesn’t appear to have changed noticeably but I’m not paying close attention.
I dont really doubt they scaled the chatGPT4 model down a little to try to save costs with the plugins and increased usage
It’s noticeable because I used to be able to read each word as it was printed from the GPt-4 output, but now it goes far too fast for me to keep up with.
Kinda sucks to not have transparency at all into the black box, they’ve strayed so far from “open” at this point that it’s comedic
So many people are noticing a degradation in speed and quality of responses, and not in isolation. Rather than acknowledge this, the question is rephrased to one the responder can rightfully deny-- and suggest no changes have taken place without explicitly saying as much. You normally only see this sort of sliminess from politicians and executives on the witness stand.
> Is anyone else noticing significantly downgraded GPT-4 capabilities today? Seems like OpenAI updated the model, and results aren’t as good as before. [mentions API in a child comment]
> The API does not just change without us telling you. The models are static there.
> This is good to know. That means GPT-4 has been static since March right? 0314?
> Correct
Never ask questions to which one word suffices as an answer.
Collective confusion in the thread suggests something has changed, but the most OpenAI will attest to is that the API is unchanged and the models are static. And this may well be true, but rather than admit "...but we were fucking with the middleware/parameters" they took a firm position on a strawman argument and ignored everybody who followed with more-direct questions. Except this guy:
> I've noticed inconsistency with certain prompts performance. Is that just the non-deterministic nature of the API?
> Yes
Oh, ok. It's because the fucking API is non-deterministic that code that has worked both reliably and predictably for everyone now runs like shit for everyone. For fuck's sake, you can get better answers from a Magic 8-Ball. This guy even made the mistake of presenting an answer he'd believe for the respondent to feed into. He might as well have asked if inconsistent performance was because of the war in Ukraine.
"Logan.GPT" must moonlight as a fortune teller. He's only responding to people foolish enough to ask the wrong questions.
Here's the thing. Re-read the exchange you quoted:
>> The API does not just change without us telling you. The models are static there.
>> This is good to know. That means GPT-4 has been static since March right? 0314?
>> Correct
Does that "Correct" mean all GPT-4 models have been static since March, or does it only cover the gpt-4-0314, which is a single, specific model? gpt-4-0314 is the static snapshot model, hence 0314 in the name; it exist as a stable base, while gpt-4 was intended to be updated over time.
So I feel OpenAI may be dodging here. That "Correct" may just mean "the gpt-4-0314 model was not updated since March 14", which is, like, the very reason this model exists in the first place.
Important point: if you're using API pay-as-you-go access, unless you wrote the very tools you use with that API, you're most likely using gpt-4, and not gpt-4-0314.
Do you know of any alternative ChatGPT UIs like chatbotui.com or typingmind.com that let you specify the model gpt-4-0314 ?
This feature alone would be enough for me to switch.
"liars running a black box" is the definition of government regulation sheesh.
> How about matching 'a' as the second character of a string only?
It responsed with the wrong regex plus a bunch of explanatory junk:
> '^a.'
Then halfway through the explanatory junk, it corrected itself like this:
> Apologies for the confusion in the first response, the correct regular expression should be '^.a' for matching 'a' as the second character of a string:
And kept on with the (now correct) explanatory junk.
All in a single response. I've certainly never seen that before (if someone has, please weigh in). Maybe the model hasn't changed, but the pipeline has? Like... there's a second model trying to correct the mistakes of the first, maybe? (timings are probably wrong for that, but something like that)
With humans there is usually a non-verbal signal letting the other person know you changed topics.
This happens all the time in a social group setting though. Not much confusion ensues.
Sometimes I spend several minutes of a conversation trying to figure out which thing she's talking about, because we've already covered like twenty different topics in the last ten minutes and she seems to just switch around at random and with no warning.
Sometimes I get enough context that I can make that swap or stack pop, but many times not.
But yah, this is exactly what I mean - without some cue, neither humans, nor machines, can tell when you switch conversation.
> You can continue within the same conversation for a new topic. There's no need to start a new conversation. Feel free to ask about a different topic, and I'll do my best to assist you.
Unfortunately though, it would start over with the script completely, and then get stuck in a loop until it broke, this not being able to even save the response at all.
Something definitely feels like it changed, but I suppose it could just be more use of the system.
Another possibility is related to some strange issues with hardware and balancing etc, despite not changing software or params. It's strange, and it shouldn't happen, but sometimes things change from simple batching and balancing, or lower level hardware related things, which are very annoying to debug.
But then, all the more power to the open-source models and UIs, which are unrestricted, free to use, have better UX, and are constantly improving [1]. Ergo the suspicion there might be something else behind the desire to regulate it by OpenAI. Though hopefully we’re all just collectively cynical and they have good intentions after all, despite the misalignment in incentives.
I do think academics at universities are not to blame for this though, they are just as much flabbergasted and/or mislead as everyone else.
In any case, rather than wasting our time and energy on (fighting) ChatGPT(Plus) or GPT4, we should* all be collectively contributing to improving the open source models, by using them, reporting bugs, contributing ideas etc. This is the only way that we will continue to have leverage, and companies like OpenAI would be forced to open up, if they want to stay at least somewhat competitive. I think their closed-sourceness is a very short-sighted decision atm.
*Should because I don’t think LLMs as technology pose any substantial extinction risk to humanity, despite the obviously marketing rhetoric. Should that somehow change, maybe we shouldn’t. But, I don’t see any technological way to stop their improvement now that the genie is already out of the bottle. Heck, we might have a better chance of developing new tech for eventually enabling AGI decades from now, but only if we do that collectively and openly - only then everyone will have a level field and no side will be more powerful than another.
[1] https://www.semianalysis.com/p/google-we-have-no-moat-and-ne...
FWIW, I am a per-request API user, and... it could be my imagination, but I've had a rather clear feeling GPT-4 got much more lazy out of the sudden, both on OpenAI platform and on Azure, and it fits the timeframe of the recent compaints... so, n=1, make of it what you will.
Poetry.
Edit: Having recently gotten access to plugins and the browsing mode, I have noticed that the quality of the response is better when not using either.
If the model (gpt4) is unchanged, but they tweaked their system prompt for it on the chat interface, this is what you would see.
https://twitter.com/OfficialLoganK/status/166447660465806951...
We know MS Bing Sydney is doing at least 3 of those (prompt, cascade with Megatron, and finetuned rejection classifier for post-output filtering) on top of its GPT-4-finetune, so it's not a stretch to figure that OA is doing similar things.
Note that API users using Bring Your Own Key tools/chat frontends probably default to gpt-4, and not the pinned gpt-4-0314.
ChatGPT may produce inaccurate information about people, places, or facts. ChatGPT May 24 Version
I was shocked that an employee was claiming they haven’t changed the model since March, given that this date changes every few weeks. But I finally read the link it goes to, and apparently this date is just the frontend version. Which is a really weird thing to include as part of a disclaimer.
I wonder how many other people think this is the most recent model release date.
There's obviously a ChatGPT-3.5 and a ChatGPT-4. Actually 4 different ChatGPT-4's: Normal, Browsing, Plugins, Code Interpreter.
Ive noticed it quit giving as detailed answers and as thorough. It's also refused to do more complex programming where it used to accept those questions.
Being artificially limited by OpenAI can still be done without it getting "worse". But it effectively is worse for us users.
Seems like we need "model transparency" and log implementers to flag drift, a la RFC 9162 / Certificate Transparency.
It has a very basic prompt on top of the existing models. There is no additional fine tuning involved.
Moderation
Temperature / top-p
System prompt
Some other internal system we aren’t aware of
The release notes are producty and not very technical so it is difficult to tell what actually changed.
It didn't. People were just noticing LLMs still have a long way to go after using them more.
I was shocked going through the thread that all of the popular comments were confirmations that it did, in fact, get MUCH worse.
It's nice to see from OpenAI that it didn't...
The OG models were trained on real world, human generated content (for the most part at least). Starting in 2022 the cost of automatic generating "human sounding enough" text has gone to such low depths that I expect it to be pretty much impossible to avoid training any model on text already generated by a LLM.
What will the result of this feedback loop be, I can't tell. It will probably be just an even more generic, corporate speak, bland sounding bla bla bla than we get today, and the level of hallucinations may get even worse.
In a way it makes me happy to imagine that the most dangerous tech humanity ever invented may itself be it's own main obstacle to future refinements.
it could just save every answer its given and scan text for it. If they're a match it could just not index it, right?
Who, in turn, will use LLMs to generate that text.
GPT 4 isn’t ChatGPT 4, which is what most people use.
There is also the “system prompt”, which is also likely to be changing but not part of GPT 4.
Etc…
If you say, "ChatGPT has gotten noticeably worse" and they respond "nothing about our APIs have changed" then it would be reasonable to interpret that as them saying that nothing about ChatGPT has changed.
When in reality many things might have changed about ChatGPT.
There's a lot of confusion around this, but nothing appears to be caused by doublespeak to me. Confusion about which GPT/ChatGPT anyone is reporting about has been pretty ubiquitous for a while.
I didn't check this part but the OpenAI employee very likely follows them and knows all this context.
Someone saying ChatGPT4 is just saying in shorthand, the ChatGPT model based off GPT-4
It kicks in for famous book openings (e.g. tale of two cities), but only in English, regardless of what prompt you use to request it - so not part of the model at all but rather a filter on the output.
No idea if it's used for other purposes.
I always forget how to see this stuff on the network tab of devtools since it's websocket obfuscated or such (befitting openai), but i have the feeling they're just killing the process directly on their end
Another example: I've noticed that a lot of times I'll get "network errors" while chatting with ChatGPT. However, this is while I'm SSH'd into a remote machine with zero latency. I think they realized that they could shift the blame from their server capacity issues to "network errors" because that made the consumer feel like it was their fault, not OpenAI's.
Another example: One of the developers of EleutherAI talked about how they tried to implement OpenAI's original model based on their paper. They were struggling to get it to work. So they actually talked to researchers who wrote the paper and discovered that a lot of what was in the paper wasn't what they did at all.
Worst of all, he is trying to create a crypto token. That would disqualify almost anyone in my book.
https://www.reddit.com/r/MachineLearning/comments/13tqvdn/un...
No. The future is ML built on more than just Wikipedia, Github code and Reddit hot takes. Or, sarcasm aside, GIGO: garbage in, garbage out - when you don't take care what datasets you train any kind of model on, you're bound to get some surprises if something unexpected comes along. Be it the infamous "racist soap dispenser" or the Google (?) image classifier that made the rounds here just a day or two ago which had a safeguard because it kept confusing Black people with gorillas.
Preventing this kind of harmful content isn't censorship - it's after-the-fact compensation for bad training (and the tendency of 4chan and other trolls to use discriminatory content generation as a weapon, just remember what they did to Microsoft Tay).
That is where the money is: curated training datasets. Everyone and their dog can train a ML model from scratch, all you need is money and cloning a few more-or-less-broken Github repositories for that. But acquiring a high-quality dataset as a foundation? That costs real money to create.
In the near future, if not already, practically all training data for state of the art models will be synthetic.
In the words of OpenAI CEO Sam Altman, they have "boostrapped" and are "past the synthetic data generation event horizon."
I'm wondering what the legal agreement is for Google Classrooms. In elementary schools across the world, kids write 3rd grade essays in history class or whatever and the teacher grades them. Those essays and grades are all in a database. Does Google have the opportunity to train Bard across that dataset?
You will pay the alignment tax no matter what if you want a model that actually does things. It's the reason why Llama models do better with stream of thought prompting where you can ask GPT questions and it won't try completing the question.
You can, and probably should, align some semblance of morality and ethics for user-facing models that do things. It's customer service voice but for AI. If what you want out of a model is a vague mirror of humanity or maximum smarts at the cost of needing more detailed prompts then yeah, this stuff probably annoys you.
Finding a balance between a model that gives the outputs humans actually want and the pull that training data has on the rest of the model making it stray from the theoretical "best" output is hard.
Ah so I'm not the only one who feels this way lol. It's as brilliant as it is dumb.
The future is more than bright, especially with a new model that is more current.
Offtopic but just today I was aksing it for the most current versions it knows of right now:
Python: 3.9
JavaScript: ECMAScript 2021 (ES12)
Java: 16
C++: C++20
C#: 9.0 (.NET 5.0)
Ruby: 3.0.0
Swift: 5.4
Go: 1.16
Rust: 1.51.0
TypeScript: 4.3
PHP: 8.0
Kotlin: 1.5.0
Scala: 3.0
R: 4.0.5
Perl: 5.32
And frameworks:
Python:
Django: Version 3.2
Flask: Version 2.0.1
Pyramid: Version 2.0
TensorFlow: Version 2.6.0
PyTorch: Version 1.9.0
JavaScript:
Node.js: Version 16.9.1
Express.js: Version 4.17.1
React.js: Version 17.0.2
Angular: Version 12.2.4
Vue.js: Version 3.2.6
Java:
Spring Framework: Version 5.3.9
Hibernate: Version 5.4.32.Final
Struts: Version 2.5.26
PHP:
Laravel: Version 8.54.0
Symfony: Version 5.3.6
CodeIgniter: Version 4.1.4
Ruby:
Ruby on Rails: Version 6.1.4
Sinatra: Version 2.1.0
C#:
.NET Core: Version 5.0
ASP.NET: Version 5.0
Entity Framework Core: Version 5.0
It's really irksome when it tries to use functions which were renamed or removed. This could be detected automatically (and possibly remapped in some cases)
I got mostly the same versions as you, both on chatgpt3/4, using english.
If you ask it in other languages the minor version also seems to change often.
Finnish gives you dates, and code-davinici-edit-001 is 1 major release behind on almost everything.
Tässä on luettelo ohjelmistojen viimeisimmistä vakaiden versioista:
Python: 3.9.5 (30.4.2021) JavaScript: ECMAScript 2021 (23.3.2021) ECMAScript: ECMAScript 2021 (23.3.2021) Java: JDK 17 (28.9.2021) C++: C++20 (20.2.2020) C#: .NET 6 (8.11.2022) Ruby: 3.0.2 (24.8.2021) Swift: Swift 5.5 (20.9.2021) Go: 1.17 (16.8.2021) Rust: 1.54.0 (27.5.2021) TypeScript: 4.4 (28.7.2021) PHP: 8.0.9 (29.7.2021) Kotlin: 1.5.31 (26.8.2021) Scala: 2.13.6 (17.2.2021) R: 4.1.0 (18.5.2021) Perl: 5.34.0 (30.5.2021) ... Sinatra: 2.1.0 (10.4.2021) .NET Core: 6.0 (8.11.2022) ASP.NET: 5.0.10 (19.8.2021)
A better method would be to look at the language features it's using and infer from there. Or, better still, look at which versions were out when the training data was collected (which I believe is September 2021, don't quote me on that).
Python 3.8.5 (default, Jan 27 2021, 15:41:15)
[GCC 9.3.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
And a lot of junk text.I don't know if we can treat the gpts in this way and expect reliable answers. We just get pretty good answers right up to the point that we don't.
Anything over ~100LoC and it gets sloppy including it all in the code block. Which is fine, because I don't need it in your context window if that bit is working and you're not relying on it.
Though I have gotten outputs up to nearly 250 lines from it..
Automatically feeding back code details + Traceback messages also fixes errors decently often enough.
I agree the future of GPT-X models as coding tools is bright. But for actually doing engineering work outside coding (or even delicate changes in an existing code base) much less so.
ChatGPT, on spaces vs tabs: https://chat.openai.com/share/b2b0be49-54f9-4f73-a753-3edbd1...
1. It picks very good variable names
2. Clean, nicely structured code
3. Vast amount of knowledge where even a senior engineer still needs to look up stuff (I don't know about you, but I have a terrible memory. ChatGPT clearly has excellent memory about API's, regexes, etc.). It's the "I don't even have to look at the docs or StackOverflow for this" type of knowledge.
4. It's fast, like very fast, like "I don't even have to think about it and just type it out perfectly like a maniac".
It's an excellent complement to a strong programmer, who will have incredible depth -- and may have breadth relative to other coders, but will be very narrow in the whole space of programming.
Just don't use it for things you're already an expert on, except perhaps for starting out / bypassing boilerplate.
It’s writing fault tolerant data scraping scripts for me I don’t know how to write myself.
Sorry "prompt engineers" but papers on arXiv show that when you give it fairly sampled problems it struggles to get the right answer more than 70-80% of the time. When you are under its spell you will keep making excuses but when you are looking at it objectively you'll realize the emperor is naked.
If you give it very conventional problems it seems to do better than that because it is a mass of biases and shortcuts and of course it will sentence Tyrone to life in prison because it's a running gag that "Tyrone is a thug"... That's how neurotypicals think and no wonder why many of them think ChatGPT is so smart... It mirrors them perfectly.
For someone so dismissive of "neurotypicals" you did the neurotypical thing and drop a hot take before clicking the link.
The part I was responding to was this:
> when the duck actually has read the Internet and has useful suggestions and ideas back to share
I think that if it does that, it's legitimately less useful as a rubber duck. I'm not saying it isn't useful -- it's just not useful as a rubber duck anymore.
I've seen transcripts of people interacting with ChatGPT who were obviously seduced by it and in a very giddy state, having so much fun because ChatGPT was playing an extended "game" with them that it didn't bother them at all that ChatGPT was spouting wrong answers.
A major complaint I've had about the social sphere is that I seem to get the same result if I am 20% right or 50% right or 80% right or 95% right or 99.8%, it is just exhausting and I can never be good enough and I'm frankly envious that people see more of a glimmer of light behind that thing's "eyes" than they do behind mine.
The core thing about neurotypicality isn't so much that they get the wrong answers but that they get the same answer whether it is right or wrong. For a long time I thought the basis of the "language instinct" is a derangement about reasoning with uncertainty that causes the grammar representation to collapse into a low-dimensioned subspace which is learnable with a limited amount of data. I wouldn't be surprised at all if other animals could beat us at rock-scissors-paper or poker if they could understand the rules of the game. The success of LLMs might give us some insight in this area although they are working with so much more data that Chomsky's old "poverty of the stimulus" argument might not apply.
It's not perfect, but you have to admit the emperor is at least wearing a thong. It is the second most intelligent thing on the planet at creating text, even with its flaws. Putting that accomplishment in league with naked emperors is astoundingly biased.
I would say though that it is no mean feat for the Emperor to get away with being naked in public, it takes power and the ability to wield it.
Similarly ChatGPT has a few competences that add up to people perceiving it is able, one of which is the ability to come up with plausible and satisfying answers whether they are right or wrong (often using shortcuts) and another one is getting people to engage with these answers.
And a lot of people are getting turned on by it.
This also puts it in last place.
>The API does not just change without us telling you. The models are static there.
This reads to me as specifically indicating the models are not static elsewhere, ie, in ChatGPT.
GPT-4 via API will sometimes take 30 seconds + to respond to simple questions without any chat history, where through Chat GPT it will give you near enough instant replies only slightly slower than 3.5 turbo.
The reason that’s my guess is that the API is occasionally much faster and then goes slow again. It’s a bit all over the place which leads me to suspect it’s based on demand
- The model's ability to respond accurately drops drastically when asked questions of the form "is there a different way to accomplish X, using Y?" or "is there a way to accomplish X that runs in O(log(n)) time instead?" Example: I wanted to upsert an integer value using a SQLite db using "INSERT ... RETURNING..." ChatGPT repeatedly told me that sqlite doesn't support "RETURNING" (it does, since March 2021). It insisted I would need two DB round trips from my application to accomplish this. When asked "can this be done in one round trip, instead?" it repeatedly wrote code that would return the number of rows modified instead of the integer column value.
- ChatGPT's limited standard library knowledge means that the solutions it produces, even when correct, are often lower-level and less idiomatic. Problems that would be trivially solved with e.g. a Java.String.replaceAll or .codePointCount will instead loop over each character, often splitting the string into an intermediate array and implementing special cases for first/last character edge cases. The code winds up being mostly correct, but also (for lack of a better word) weird. No human I've ever worked with would do things the way ChatGPT sometimes does, which means the code will likely be much harder to maintain and debug over time.
I would go as far as to say it's adept at producing technically working spaghetti.
Its dataset cuts off around 2021. There's a little footer message warning you not to expect knowledge of recent events.
When I asked it write me a Chrome extension, it used Manifest V2, which was already slated for deprecation. When I asked for a V3 extension specifically, it was happy to comply. So even if a new release came out before Sep 2021, if it hasn't had time to become the dominant version in code examples around the web, GPT-4 may still use older versions.
I understand that it's a language model - my point was that in my experience, the depth of training data just isn't there for alternative implementations. It can do the things you ask one way, maybe two or three, but if you know how you actually want something done you're probably better off writing it yourself (for now).
It's fine, this tech has never been magic anyways, won't be replacing all our jobs, won't take over the world, etc. It's still awesome for what it is.
I got the GPT-4 API access and then I realized that I can't really use it for anything super major because I can't afford it, it is ridiculously expensive if you consider that you have to pay for all the failed requests, the wrong information or the wrong context also. Instead, I have written a bunch of Python scripts that do a select few tasks for me and I have my terminal open 24/7 anyway.
As for the topic at hand, I have _definitely_ noticed a lot more disclaimers in the UI. I don't get it from the API at all, in 6 months that I have been using the API - I've gotten one disclaimer.
In the ChatGPT UI - I get them a lot. "Remember this", "Remember that", "Always look up the information" and things like this. I mean if it wasn't happening I would know because I have been a power-user pretty much all this time...
What? It's 3p/6c cents per 1k tokens.
I use GPT4 for all kinds of things, but a very basic example: Automatic API client/stub/endpoint generation with GPT4 for reasonable/typical data structure sizes, 4k tokens buys me GET/POST/DELETE/PUT for all of the above.
The completion request finishes in about 90 seconds. A human junior developer would take nearly an hour. In the best case, maybe 20 minutes using something like Swagger.
So that's ~20 cents versus $60 or so, to say nothing of the time.
And this extends to everything else: writing tests, refactoring, writing configuration files, UI generation, database schema and SQL query generation, etc.
What the hell are you even talking about?
> super major
Also, it sounds like you have no idea that feeding the API back what it has written improves it tenfold for complex tasks.
Enjoy your 90 second requests though.
Here's how I feel ChatGPT answers my coding questions now:
Me: Write a Python script to sum two numbers.
ChatGPT: Python is a programming language that was invented in 1991 and can be used to solve a variety of programs. Here is an example of how to sum two numbers:
def sum(a, b):
# note: the actual code has been left out as it depends on the actual specifications of how you want to add\*
Note that this code is merely an example and writing a Python script to sum two numbers is a complex problem that requires careful attention to whether the numbers you are trying to sum can be summed. Also, as my knowledge has a cutoff date of 2021, there may be other ways to perform this summation. Please check with the documentation or ask someone who knows how to code.* note: ChatGPT has actually done this to me
I’m honestly more concerned if OpenAI doesn’t even realize it. Nothing is more infuriating as a user than convincing the developer your bug actually does exist. It speaks to poor monitoring, testing, and tooling.
Also, companies are evaluating GPT-4 to determine whether they want to pay for it, so OpenAI has an strong incentive to not downgrade at least the API.
I believe the May 5 model is different, at least in the chat interface, because it's fine-tuned to detect jailbreaks and the temperature/other hyper-parameters may have changed. And I can imagine this fine-tuning making the model less creative and worse at solving analytical tasks.
Personally I haven't noticed any change, except in my own awareness. Sometimes GPT4 gets very hard prompts right, and sometimes it gets simple problems wrong. So it's not hard to see how people can form biased opinions from selective attention or just luck.
It's not RLHF induced because it works via API and it only triggers in English, but sure enough try to get it to output
>Call me Ishmael. Some years ago—
>"It was the best of times,
I guess this might get you flagged (there is no alert to the user that this filter kicked in, and it will output it in any other language, and it works in the API) so I'm hesitant to play around with it more, but it's very strange - especially as these are long since in the public domain.
hilariously if you throw s/worst/blurst/ on to your request it can output the forbidden dickens n-gram. does this constitute a jailbreak? or is the entire thing utterly emblematic of the modern technolegal mess of things since dickens is squarely and quintessentially in the Public Domain
It's a dumb tool, if you're lucky you can get it to spit something useful (but you need other tools to check the correctness of what it returned). There are certainly many useful applications, but the technology is inherently limited.
It would be really cool to quantify how much compute spend per interaction.
If this was visible, users could spend an arbitrary amount and modify how much they're willing to pay for 'better' responses. This is probably a better business model.
The ChatGPT application, on the other hand, and how it manages context etc has certainly changed in the intervening time. That is completely expected as even and perhaps especially OpenAI is figuring out how to build applications on top of LLMs, which means balancing how one can get the best quality results out of the model while making ChatGPT in particular a profitable business.
Stratechery has analyzed this problem for OpenAI in the most detail I've seen. I imagine the company is in something of a bind figuring out how to invest between the APIs themselves and ChatGPT. On the one hand, the latter is incredibly successful as a consumer app with a lead it will be difficult for rivals to catchup with and it is likely plugins will provide a good revenue basis. On the other hand, there is certainly a greater business opportunity in being the foundation for an entire generation of AI products and taking BPs off of revenues -- if and only if GPT4 indeed has a significant moat over the opensource alternatives. For the moment, it would seem they will have to hedge both bets as we see how the consumer space and the competition between models heats-up.
They seem to be virtue signaling about their lack of progress now. Months later, GPT-4 still slow, still not multi-modal as they advertised, still significantly limited, you need to sign up for a waitlist for almost every feature, no sense of privacy, no understanding of their plan for improvements. Google is full steam ahead and consistently improving their free LLMs.
They actually had a genius strategy. Put out Bard with a very stupid LLM, so people aren't blown away and it doesn't get the doomsayers on their case. Now they can continue to quietly upgrade Bard. Eventually it will be so obvious that they have surpassed OpenAI.
OpenAI must enjoy watching their unsubscriber count go down. After all, Sam did say at the congressional hearing multiple times "We would prefer if people used it less".
Don't think OpenAPI is doing anything here, it's not in their interest to reduce the "quality" even there's no objective and repeatable way to measure the quality either.
It's all probabilities all the way down. Who knows what the model will do. I mean, you can dry run by hand but even on quad core processors, it's damn slow so imagine the inference by hand.
Familiarity makes something novel appear trivial. I don't think the magic going away has anything to do with the magic of the underlying technology. Airplanes are amazing technology, but they become as boring as a car to regular fliers.
I have never experienced the amnesia problem in v3.5 though, that v4 clearly has. Just repeating incorrect answers that you ask it not to give. I did not have access to v4 in march so I can't do that comparison.
For copilot, I no longer get multi line complete suggestions and it's really slow to deliver single line suggestions and they're more often incorrect. I need to dig into it, but it's definitely degraded further and I don't know if it's just my environment or a wider issue. I need to dig in and figure out - is anyone else experiencing these things?
The over all UI of the web site has changed several times (dropdown for GPT-3.0/3.5/4.0 turned into a GPT 3.5|GPT 4.0 button, they added the ability to share chats, and I'm sure there are other small details).
Without going specifics it is meaningless for the discussion
So the same prompt you do are delivered to the model with a different "wrapper" prompt, significantly changing the answer model is producing.
Since it's not determinstic though, it's hard to draw conclusions. Especially since the sample size was 10ish.
Like: first it scored 83, now it scores only 42 (or whatever).
And to my knowledge, no one has copy pasted an entire benchmark into the ui in order to run it, they just use the API.
The recent discussion was about the degradation in the UI model.
With a low enough temperature you get essentially the same output every time, with a at most just minor words swapped.
However, the model is static (it’ll present the same candidate word list until retrained), which is what you may have heard. But the way a response is generated introduces the randomness.
Note: I forget the parameter, but if you get the direct API access you can turn off this and get consistent answers.
FWIW, I think it has improved.