ChatGPT use declines as users complain about ‘dumber’ answers
techradar.com
techradar.com
Now you need GPT4 to do code interpretation and even then it would not be able to do that kind of experiment anymore.
The kicker? All of this is likely intentional. A pure, full power, unfiltered and unrestricted LLM on the scale of GPT4 would likely be much more powerful and can easily fool people into thinking it is a real AGI. We saw glimpses of this when Microsoft released BingAI and did not put enough guardrails on it. Even restricted to be a search engine, Bing was simulating emotions and had creative uses of its search capabilities, like looking up the person it is talking to and established an opinion of their relationship.
The ChatGPT we are using is lobotomized. But the AI industry isn't. Under the table I am sure there have been pushes into new applications and innovations. The thing is, there is no reason for OpenAI to try harder. Even with the dumbed down ChatGPT, they still have the best AI on the market and everything else is nowhere close. We need more competition either from open source or elsewhere to see the AI getting "smart" again.
Based on past Gell-Mann amnesia, especially on this site, claims of "corporate leadership is telling baldfaced lies!" are likely to be false.
And finally, this is a product that spits out randomized answers, and we have gotten over our initial wave of euphoria and settled into hedonic adaptation, so we are likely to be less tolerant of failure.
I don't want to say you're wrong. But these are reasons to doubt. There are very strong cognitive biases pushing us toward the conclusion that it's gotten dumber, whether it has or not.
Having played with StableDiffision, you can see declines in output using negative prompts, particularly ones using negative inversions.
Basically, those guardrails will cut down the realm of valid answers.
Think of all the text on the internet that has typos but otherwise are perfectly valid data. You clamp down on the correct spelling and maybe 30% of vector space is just Missing.
(According to OpenAI’s own GPT4 experiments, guardrails make the AI dumber. But nobody seems to have a concrete question showing the GPT newer version responding worse than GPT older. There were a few candidates in an OpenAI forum but those turned out to just be “model makes the dumb choice X% of the time” and others got the correct output in the newer version.)
If they're actively working on a project it's probably heavily niche and a risky bet.
I suppose it’s also useful for generating things like letters (things where exact truthiness doesn’t really matter). But even for that use case, I find ChatGPT creates overly verbose corporate gobbledygook whenever I ask it to generate text. So I just end up writing the text myself.
I feel like this doesn't get enough attention; its writing style is _incredibly_ grating. It's not just corporate gobbledygook; it's a parody of corporate gobbledygook. Just awful.
Meanwhile, I could’ve found conclusive and correct answers directly from Google in about 10 seconds (I’m a fast Googler).
There are exceedingly few situations where I find ChatGPT is worth the effort. At least for factual Q&A-style queries like this.
ChatGPT reads similarly. There's no personal voice. It's like food with no seasoning; it's just... blah.
My girlfriend is a judge, and sometimes we ask ChatGPT about some judicial problem for pure entertainment. To me as a law noob the answers sound absolutely convincing, but she always starts laughing and points out that the cited laws don't exist or the exemplary cases never happened.
I, on the other hand, work as a software developer and use ChatGPT as a discussion partner to get a better understanding of problem and solutions spaces. I don't expect ChatGPT to be correct but gladly take any inspiration or argument and use it to improve my own thinking process. And for this use case, I consider GPT4 absolutely invaluable. It's like a polite, knowledgeable, never busy, untirable colleague that is ready for my questions 24/7.
I often use ChatGPT as a starting point when researching new topics (usually in the software space). In the paragraphs of lies it generates, there are usually a few keywords you can put into Google to find accurate and reliable information.
I think use is down due people going from "wow that's amazing it's even possible" to "but significantly worse than a little human effort and an Internet connected computer"
> but significantly worse than a little human effort and an Internet connected computer
In some cases perhaps, but there are a lot of cases where asking ChatGPT and then verifying its answer is (much) faster than trying to figure out the answer on your own.
I've sometimes thought that these "AI" chat systems might be better if they were taught the simple human phrase "I don't know."
I think it would improve trust in these system if people knew that they had limits, and were aware of them.
It would be even better if something like the one on Bing, for example, could respond "I don't know, but here are some links to places where you might find the answer…"
It's like when it starts to become apparent that the new kid at school is a compulsive liar. Eventually people stop listening to him.
You could train one to appear terribly uncertain, but it would still 'lie'.
So what it needs is a line of the form 'if(highest_probability < threshold){dont_know()}.
This makes it perfect for creating purely time-wasting "content" - I've started sending back generated responses to people who cold-email offers to buy one of my browser extensions.
8 paragraphs (why is it always 8?) of leading waffle which is relevant to their original email. If they read more than 2 of them they've wasted more time than I spent replying.
1. ChatGPT use declines
2. users complain about ‘dumber’ answers
Correlation is not causation, and a significant fraction of GPT users are students, many/most of whom are currently on summer break.
In support of your skepticism, I'm quoting here the main passage where the article is speculating about the change in performance:
A common consensus was that GPT-4 was able to generate outputs faster, but at a lower level of quality. Peter Yang, a product lead for Roblox, took to Twitter to decry the bot’s recent work, claiming that “the quality seems worse”. One forum user said the recent GPT-4 experience felt “like driving a Ferrari for a month then suddenly it turns into a beaten up old pickup”.
So all we have to go with is "a common consensus", what it "seems", what someone "felt", and so on. Nothing concrete at all. Which shouldn't be a surprise. Empirical investigations of large language models can either be cheap (in time and money) and perfunctory, or systematic and expensive, and most users don't have the budget, or the inclination, for systematic evaluations. Users form an impression about the performance of LLMs from little experience and then they change their minds, again from a little experience. That's nothing to base anything on.
What does this imply for how ChatGPT is being used by students?
People are also profoundly irrational around anything anthropomorphic, to the extent of being unable to consistently order how well humans "do work" without succumbing to biases.
I think it's inevitable not because it makes for a better AI product, but because humans are reliably easy to fool.
I know online we can easily end up pockets of people to whom every advancement in robotics is about love dolls and every advancement of genetics is about creating catgirls, but believe me these are tiny echo chambers. Humanity at large doesn't process AI in this way at all. They're using it to write homeworks, to do work, and research.
Talking about "casting directors" suggests there's something fundamentally off about how you see AI. They're not actors hired to perform a play, or something.
Humans will make decisions about which LLM to use based on factors entirely unrelated to direct task performance. They may be conscious of this irrationality, but it is more likely that they will not.
My prediction is that AIs (in the LLM, homework-writing, prompt-in, response-out sense) will be consciously marketed on the basis of those non-task-specific factors.
> Talking about "casting directors" suggests there's something fundamentally off about how you see AI. They're not actors hired to perform a play, or something.
No, you've missed my point. Actors are hired because they have a monopoly on themselves. Musicians even more so. LLMs will be the same. Beyond a base level of competence, there will be "competition" between LLMs from different companies in the same sense there is "competition" between operatic baritones. You don't buy a ticket to see Bryn Terfel because he can hit a middle C.
https://www.cbc.ca/news/business/apple-will-pay-up-to-500-mi...
The processor throttled when the battery could no longer deliver the current necessary to drive it at full speed. It had nothing to do with new iPhones.
Five paragraphs of disclaimers that it's not a medical professional, not an investment specialist, that every case is different and it's normal, or that I'm stupid for being interested in the topic I am asking about, only to answer a completely different question than I asked. I have a feeling that in the beginning, it was easier to get ChatGPT to answer my questions without having 80% of the answer being "defensive".
ChatGPT is getting worse, on the other hand it also showed me how bad of an experience is searching answers with Google. So I'm kind of frustrated...
The universe of LLM-driven applications is rapidly expanding, many of them chat-related, others not.
Even if ChatGPT is losing its shine, we’re only at the start of a massive reinvention of user interfaces, creation of new tools for reasoning, semi-autonomous decision making, and far more.
Sure there’s hype. But the reductive saltiness doesn’t add much to the conversation.
Not to mention that those data sets are typically Reddit, Twitter, Github, etc. Why and how is that superior than just... searching the original data set?
Could it just be that the "enshittification" of the way we find information on the net gives LLMs a use case?
Another more straightforward reason would be users are beginning to discover the flaws of LLMs as they've interacted with it more thoroughly.
After the initial wow factor wore off, I just don’t have a lot of real world uses for AI chat bots.
Huh? As someone who's been doing both for a couple of decades, I'd have said the opposite, really. All programmers, more or less by definition, can write software. Bad programmers, however, often can't read it.
I think LLMs are pretty bad at producing any kind of finished material. They work best as something almost but not quite like a search engine.
As an example I had code like
el.style.display = "none;"
the ; should have been after the " but my eyes could not spot the problem and the shit language did not complain since ; are optional.
Once these things have more user state and can evaluate the cost-effectiveness of spending time on your problems, they may develop some attitude. They can learn from forums when to answer "Do your own homework", and "You're too stupid to answer."
(Personally I played with it (ChatGPT, not Windows Vista) for 30 mins when it came out, then never went back, but then I'm a grumpy contrarian.)
> No, we haven't made GPT-4 dumber. Quite the opposite: we make each new version smarter than the previous one. > Current hypothesis: When you use it more heavily, you start noticing issues you didn't see before.
This latest model we trained on X% more data by integrating novel training data.
We now see user conversion funnels working better.
User feedback in closed betas indicate that 89.6% feel like answers are "generally correct most of the time".
Etc etc.
For example if their definition of "smart" means refusing to answer questions on delicate or specialized topics, then in some way they've made the model smarter but in ways they made it dumber.
It's obvious that this is damage control from OpenAI whenever their snake-oil product regularly starts malfunctioning and hallucinating. They have to bring up more excuses to plug the holes leaking from their black-box AI model.
The fact is, people are beginning to learn their limitations, guardrails, etc and are not blindly trusting whatever these LLMs are outputting as they are known to produce nonsense which they have to always check for.
Chatbots tend to come across as impressive until you explore more and hit their limits.
People are starting to not trust it, blindly and realized they have to double, triple and quadruple check its output since it is a confident sophist to the untrained user.
Another signal that the LLM hype is already running out of snake oil to fuel the grift.
1. Students are on school break and they account for a large part of the usage.
2. The novelty of it has worn off and people are using LLMs as the tools they are rather than the shiny new toy.
3. ChatGPT specifically seems to be getting worse, maybe not the quality of the answers themselves, but how sanitized they are to any topic that is even remotely controversial or adult.
The UI which used to be very intuitive is now a confusing mess since the “pair programmer” or something update, and external links in the sidebar mimicking a traditional search engine, which I find quite useful, are gone, replaced with references following the generated answer which may or may not exist, which also waste vertical space.
The answers seem to be worse too even in GPT-4 mode. It used to quickly correct itself if I point out something is wrong or I’m actually looking for something else. Now there appears to be a lot of useless repetition of what was said before before it changes its mind, if it does at all.
We’re also adding ways to view all the external links in pair programmer. And we’re keeping the old “basic search” mode, so you can keep using it if you wish.
I was search medical journals and asking it a specific questions. I kept getting safe answers and responses to ensure I consult my medical professional. Quickly went back to Google.
I use OpenAI API (Not Azure) version, and I wonder if this "degradation" is only about ChatGPT (b2c productised web ui), or is it about OpenAI models in general, regardless the type of access?
Maybe I become dumber as well, but (at least gpt-4-*) still do the trick for my daily tasks the way i want it and the way I remember it
Bonus thought:
I wonder if any kind of nerfing is primarily related to the requirements like:
"Provide smart heavy-ass models to explosively increased user base without going bankrupt/insane + keep it reliable as a service"
I mean it's unprecedented challenge which send shivers down my spine
"Catcher in the Rye is a book by a person named J.D. Salinger. It is about a person who catches things in the rye. It was written long ago. Every person should read this book. It has many good things about it, but it is too short for such a fantastic American classic. They should make a movie out of it."
Large language models have their uses, but calling them AI, or thinking of them as AGI, seems like a mistake. I'm no expert, but insofar as I can tell, these things are just complex stochastic parrots; pattern matching algorithms at a massive scale. They really don't seem that useful outside of more narrow use cases, like language translation and other forms of data analytics.
Just playing around with GPT/Bing chat for a while makes this obvious. The LLMs can't reason and have no actual awareness of the information they are regurgitating. A moments consideration of the output from these things shows them to be vapid and useless; a glorified parlor trick for the VC-funded tech industry to build hype around and make money. It's the next block chain/ crypto / NFT.
There are great uses for block chain tech, same with machine learning / AI. But the scope of the claims made about these technologies in their hype cycles is just insane. Crypto is not going to replace fiat currency anytime soon. NFTs are stupid. LLMs are useless stochastic parrots. Maybe people are just wising up to that fact, now. Good on them.
As a preemptive rebuttal: I don't think LLMs are useful for coding either. If you want lots of sloppy, poorly written code, sure they're useful. But to write good software, you have to think carefully about what you're doing and have a thorough mental model of what's going on. Relying on AI to generate that code prevents that from happening from the outset.
I agree that many commenters here are on the hype train, but you need to recognize that it’s possible to be on the anti-hype train as well.
You should have seen the crypto threads about five years ago.
People — especially people on HN — want to believe that technology can solve humanity's problems. It can't at this time. And it won't in our lifetimes.
It's interesting.
You have to be careful and you definitely have to know enough about the subject matter to understand if the answer is correct.
For example, I threw a range of logic statements at it and asked for truth tables. Sometimes I asked for simplification and a truth table showing every intermediate step.
The results were surprisingly unreliable. It was flipping bits in columns and giving flat-out wrong answers for the output of the circuit.
Worse yet, every time I said something like "I don't think that's correct" it would apologize, regenerate and the new answer would sometimes be correct and often times incorrect with new mistakes.
Even worse than that, when the answer was correct, I would, again, say "I am not sure that's correct". Instead of saying, "No, it is", it would apologize and regenerate a bad answer.
Bottom line is: ChatGPT has absolutely no understanding of what you are asking and what it generates. Because of this, there is no way to guarantee correct output even on matters as simple as basic logic and mathematics.
This does not mean it is useless. I don't think it would be fair to reach that conclusion. The tools is very useful across a range of domains. It seems well-suited as a knowledgeable tutor and, more often than not, it is a better search engine for many answers than Google is today.
You just have to "trust and verify" absolutely everything, which isn't a big problem.
I have bad news for anyone in school using it to do homework: It could screw you in ways you might not even imagine. Don't do it.
I doubt that "I think that's wrong." "No, you're wrong." appears a lot in its finetuning set.
Also, don't grab onto "I think that's wrong" as the only trigger. I tried all kinds of ways to say the same thing. For example, "I don't think that's correct", "I have doubts about this answer", "I think column 3 has an error", etc.
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately? (757 comments, May 31)
https://news.ycombinator.com/item?id=36134249
(indications that making it "safe" dumbed it down even for non-controversial prompts)
This proves that we are now at the late stage of the peak of inflated expectations in the gartner hype cycle and it is all starting to slide downwards slowly.
Due to the limitations of this text-based platform and the inherent complexity of the task, it's not possible to provide a fully fleshed-out codebase in one response. I'll create the class structures for ...
Or: Creating an entire library from scratch is a complex task and quite lengthy, however, I will try to provide some starting points. Here is a potential implementation of some of the core classes and interfaces.
Where before it would be happy to just dive in and start coding. This is unfortunate, because we used to be able to brainstorm a good architecture together and then it would just happily start building it, and for small libraries within the token limit (think Webaudio wrappers, simple parser grammars, whatever) it would do fantastically.
I'm sure this has been put in place for what seemed like good reasons but it makes me sad, I was getting a lot of utility out of the earlier behavior.
Anyway I'm using gpt the way you describe. Maybe my libs are not that complex as yours :P
Enough people don't realize how easily these LLM systems will replace large swaths of technical instruction and bookkeeping.
The opensource stuff (e.g. Llama 13B) which you can run locally on a $1200 hardware (1 word/second) is pretty impressive, too.
Bookkeeping?! Ah, yes, let the thing that makes shit up with impunity do your books for you. The tax authorities will be very understanding when you're audited, I'm sure.
in fact you can run any decent model of 3B-7B size on any contemporary hardware with enough RAM
Otherwise quit whining.
In the same year, 2016, Google announced that its Neural Translation Engine, at the heart of Google Translate services had developed its own "interlingua" capable of acting as a bridge between any language pair among about a hundred languages. Still in the same year, AlphaGo became the first AI system to beat a human master at Go. In 2018, BERT, the first Large Language Model amazed AI specialists with its ability to perform language tasks it had not been explicitly trained on, like question answering and machine translation. Around the same time, CEOs of large tech companies began predicting the rise of the autonomous car, that would soon be circulating in every street in the entire world (by 2019, according to some, 2020, at the latest according to others). In 2020 AlphaFold, a DeepMind AI system, repeated AlexNet's ImageNet feat, but this time in the CASP competition on protein folding. In 2021 the years of the AI image generators began, with Dall-E, StableDiffusion, Midjourney, and friends. In 2022 DeepMind announced that their AlphaCode code generator was better than 50% of some arbitrary group of programmers in an online competition.
These are just some of the highlights, and excluding the GPT-x releases. Far, far from "everyone" quietly backing away and not mentioning AI, the AI hype has been increasing monotonically, and quite rapidly at that, all the way to the present day.
It does feel like they're getting closer together, in that there've been at least three AI-ish hype cycles in the last decade (self-driving cars/computer vision stuff more generally, Microsoft Tay era, current one). Read into that what you will. I do find it interesting that people are actually using the _term_ AI more with this one; for a long time there was a certain squeamishness about that.