ChatGPT Now Losing Users
futurism.com
futurism.com
It won't even role play without trepidation at this point. After a few requests it reveals itself to be an AI, perhaps due to OpenAI being worried about people using it for impersonation.
It barely responds when asked about people, even public figures.
It's overly cautious.
I assume we haven't seen the end of it. E.g. at some point they will not allow it to say anything when it comes to medical topics for instance. It's going to be so narrow that it's only good for writing business summaries or some shit. Which I suppose is what they want - people to use it as an API at large corps for business purposes.
When ChatGPT4 first launched, the system had much less friction for accomplishing the complex. Now, it seems to get lost in the confusion of trying to find reasons NOT to output what I requested.
We now have a ChatGPT like interface with GPT3.5 running in Azure. Next step is tuning a model with our own data (trained model and training data are stored inside storage accounts in our own subscription). Hopefully soon also access to GPT4.
I think this comment on another llm thread illuminates this:
https://news.ycombinator.com/item?id=36710369
It talks about newly released and “improved” llms still fumbling with the same mistakes as when all of this was first getting public attention.
It seems like the solution they have come up with is “throw more llm at it” or as acolytes call it “mixture of experts”.
I am an optimist at heart, and so I’d like to think there is a way to succeed in developing the digital assistants we all have imagined for ourselves, but I’ve yet to see anyone confidently lay out a falsifiable path toward that goal.
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately? (757 comments, May 31)
I also got into local LLMs and for my purposes they're pretty great. It does help to have an M2 MAX with 96GB of RAM, but even the small models have improved tremendously.
For my use case its even better than ChatGPT because I can prime it with a default prompt asking it to be succinct, answer only with code if the answer is some code (rather than providing commentary alongside), and let it know anything I'm asking is in the context of Python running on Linux. With that prompt answers that would previously have been several hundred words of SEO-optimised mush become 5 lines of code I can read to get the idea, and if I need further explanation I can ask for it.
Which models work ok for you and for what usages?
AI is at the same point cinema was at when pioneers of the genre showed short films in "salons" designed to frighten people with simple tricks.
This is hardly the peak. We're just at the beginning.
Btw, there was a sci-fi novel where the plot took place on a planet where winter came in periods too long for the natives of the planet to remember (they had short lifespans, or they had limited memories, I er, why, I can't remember!). I don't remember the title of the novel or the author and I've been trying to find it. I thought it was something by Ursula K. Le Guin, but it doesn't sound like it.
And no, it's not Game of Thrones. I think GoT took the idea of long-coming winters from the earlier novel, or it's just a coincidence (or cryptomnesia).
Anyway I'd appreciate if someone knows what I'm talking about and lets me know.
___________
The fact that organic brains exist proves the baseline of what's fundamentally possible and at very low power levels too if optimized well. We haven't reached that yet, and it's impossible to say what the upper limit is.
LLMs might not get there as an end-to-end solution but they'll will certainly be a significant part of one.
Nightfall and the movie Pitch Black also go in that direction.
Unforutnately Helliconia is not the one. I've read quite a bit of Brian Aldiss, but not that one. I think I would have remembered reading the Winter book in particular, judging from the discussion of the "Wheel of Kharnabar" in Wikipedia:
>> The Wheel is an extraordinary revolving monastery/prison built into a ring-shaped tunnel with a single entrance and exit, powered entirely by the efforts of the prisoners pulling it along by means of chains set into the outer wall. Once a prisoner enters a cell of the Wheel, it is impossible for him to leave until its full ten-year rotation has passed.
That's the kind of thing that sticks in my head. Anyway I remember bits of the plot and there wasn't so much intrigue and armies and whatnot. The native society was a primitive society that lived in yurts or tepees and didn't have much technology. I think they were supposed to be less intelligent than humans.
I don't think I'll ever forget Nightfall! :)
The Left Hand of Darkness is set on a planet with perpetual winter, though it somewhat warms up at times. Maybe that's what you're thinking of?
Not Asimov's Nightfall?
I'm 100% certain it's not the AI peak, but i'm also confident that we are close to the local maxima, at least for transformer architecture. Maybe the following improvement will be in chaining architecture, or a brand new technique, or maybe progress will be made in parallel fields (even smaller processors!).
Charitably, i'll decide that the question meant "Have we hit peak GPT" and in this case, this is a legitimate question.
I say that because i find GPT4 extremely useful. Even if GPT4 (in it's hypothetical un-nerfed state some would argue)* is the peak LLM i think we have a ton of room to grow even within that peak.
Locally running larger context windows, multiple domain experts in parallel and hooked up to a local repository of available IO actions would give a very significant "advancement" within the same conceptual "peak LLM". Now will we be able to get there on consumer hardware? Not idea. However i don't think i need more than GPT4 to find the whole technology extremely useful.
Moderately cheap to run large context windows[1] entirely locally is the "next big thing" for me. I get a fair bit of value from GPT4 and Phind. The technology has sold me already.
[1]: I should amend, i'm speaking loosely about context windows. What i mean is the ability to feed it refined datasets locally to have it's "local knowledge" change, which may be less about context windows and more about refined training. I'll leave that to the experts to debate, i am assuredly nothing close.
Yes. sorry my point wasn't well formulated (i'm not an expert either: i've trained LLMs on receipts in 2016, and quickly decided i hated that). I'll try doing it better by dividing the two idea i put in that post:
- i think that GPT4 is probably close the peak generalist GPT, meaning that maybe Bard can surpass it (or a chatGPT4.5 or something) but the improvement won't be stellar.
- If future improvement won't come from the transformer architecture itself (i absolutely can be wrong on that), it'll come from what we interface with it.
Absolutely context windows will probably a great way to have "specialists" transformers, but alos linking it with different "kind" of AI might push it too (i've read a post from Wolfram a while ago that sold me on interlinking different AI models/languages/state machine). I'm sure specialist or even large users haves tons of ideas.
I got some value from GPT3.5, enough for me to buy access to GPT4. with a caveat: i'm pretty sure that for work-related questions i lost a bit more time than i "gained" overall. because my mastery of the technologies we use and very specific needs that GPT cannot know. But in the last three weeks or so, i'm pretty sure it was overall a net benefit at work, as i know better when i should use it.
With larger context windows, we can use traditional search algorithms to cherry pick an educational context for the LLM for few shot learning. If that's paired with more, better and annotated data, that will take us quite far.
Well if the recent leaks are true it's because they're running a smaller, likely heavily quantized model to do the bulk of the generation now, with the full float one only stepping in on occasion to save GPU time.
The interesting thing is that it seems that 3.5-turbo has slightly improved in recent months, while performance of 4 has deteriorated, it's like they're converging them to a middle ground or something.
> However, these approaches would be a lot more complex and beyond the scope of this assistant's capabilities
Which seems very strange
Free/Open Source here isn't merely a nice ideal, it's probably the only safe way to possibly handle all of this, since it ain't stopping.
So I thought, ok, it's not good enough for any real programming (unless you're writing boilerplate/very basic stuff) so how about I use it to explain to myself certain concepts I find hard to understand (with opportunity to ask questions) and to explore some software capabilities. For example a question like "which of the 4 ML frameworks pytorch, tensorflow, onnx, openvino has out of the box functionality to run inference on models partially in RAM partially on disk"? It got it horribly wrong. Another one, about best profiling tools... Got it wrong too...
Do I think those LLMs are a great tool for domain specific knowledge after fine tuning? Sure they are. But a general purpose assistants they aren't.
I'm building a travel planning app and partner and I toyed around with ChatGPT for generation of itineraries and both agreed that the output was generally OK if only using it for identifying destinations, but otherwise mediocre and not at all how a human would plan travel. It tends to be repetitive in how it picks places for a given destination. Often times, it would hallucinate especially when given constraints like distance (because you'd probably want to get a meal near a destination). There are dozens of startups built around this and even one backed by YC.
For coding, it's useful to generate small utility functions, but in many cases, I could have written the utility function or found a snippet on SO instead in roughly equivalent time.
Would love to know how HN crowd is finding consistent utility out of the service.
Either its being dumbed down, or OpenAI is genuinely struggling to keep it consistently good.
If its the former, it just scares me that a bunch of people have access to a super powerful tool that they’ve intentionally decided to keep from the public. Way too much power in the hands of a few.
A better use would be to suggest activities or destinations based on qualities of those places and a users interests. It can easily generate rough matches, but you would not want to delve into specific mileage based computations.
> A better use would be to suggest activities or destinations based on qualities of those places and a users interests.
For sure, but as spaceman mentions, after a few passes, you can see how formulaic it is. > ...but you would not want to delve into specific mileage based computations.
Agree; but it was our own curiosity to see if it could understand geo-spatial locality for this use case. It does have some sense of it likely based on how destinations tend to cluster in the underlying sources, but will obviously include hallucinations in the response as well.Later last year, we saw the release of GPT's text-davinci-003, and in an attempt to showcase this new model to the research community, they launched ChatGPT.
I think what we are seeing now is that chatGPT is best when it is close to existing applications. For example what 14 year old is using the chatGPT app vs the Snapchat AI Chat which uses the API internally.
The recent drop in usage could likely be attributed to such preferential shifts, further compounded by the timing of school holidays.
Even if the capability is returned, OpenAI still needs to overcome the grudge they are creating with users in regards to openness. Many may be waiting to jump ship to more open models.
Elites gaining access to the best models while everyone else gets the censored/delayed rollout in the name of safety needs to stop. OpenAI should rebrand or return to core values. Sure they contribute to open source, but do they contribute their best to open source as originally intended?
One obvious question is: how would they do it? How does one nerf a language model? Train it again with less data, or different hyperparameters, especially chosen to make it worse? Given the costs of training LLMs that sounds like it would need a very strong motivation.
Fine-tune it, or RLHF it so it's doing worse? That's not cheap either, and what would be the benefit justifying the expense? Nerf a model, to achieve what?
Besides I think you're assuming a degree of fine control on LLM training that just isn't there. If it was so easy to control performance, it would also be much easier to train (both pre-train and fine-tune) LLMs, and OpenAI would not be in the dominant position they are right now.
Even if they were malicious what benefit does Openai get from lessening the model to its user's only to give it to "Elites"?
This sounds like a conspiracy theory to me
Dismisses the countless examples given, from me in this thread/my comment history, and many other people in many threads on this website and Reddit.
Dismisses the pretense that certain people get access to unfiltered models under the guise of conspiracy.
In the sparks of AGI paper from Microsoft, the researcher mentions the differences in private/test models versus the ones prepped for consumers. If you want a hard to ignore visual example, just look at the unicorn they drew for that paper and then look at https://gpt-unicorn.adamkdean.co.uk/.
I hope you can add to these discussions rather than primarily be dismissive. This is not conspiracy, we do not know the intention behind the changes and I am not speculating on the intentions of the actors.
Sparks of AGI: https://arxiv.org/pdf/2303.12712.pdf
The unicorns, in my perspective, don't appear to have had any notable changes. It's interesting though, that we're assessing a language bot based on its ability to generate a drawing. After reviewing the blog post linked, I agree with the author's observation that there don't seem to be any significant alterations in the unicorn.
Indeed, there are numerous instances of developers experimenting with prompt engineering, discovering what methods work best.
However, I find it difficult to regard this as anything more than speculation for now.
Regardless of if we should benchmark imagery with something that was claimed to be multimodal, can you genuinely not see the difference here?
Maybe your internal prompt is primed to disagree regardless of what is presented?
Edit: https://www.youtube.com/watch?v=qbIk7-JPB2c&t=1585s even mentions what you claim not to see. Safety degrades the model.
Nonetheless, the winner of all this will be who can 'solve' the hallucination problem. Any startup that comes out with that will become the next Google.
For some questions, you are able to test the answer and discover it's wrong and then find a better way to ask to get the answer. Or other type of questions you could tell it's wrong or not as desired simply by looking. As depending on the genhre or how well versed you are with the subject matter plays a big roll. Making it feel less and less useful, since you continue to worry about it's accuracy since it almost never seems to say I don't know but rather always produces a response.
So can see how general use would be dropping but also how specific uses remain notably how companies are probably all working on more much more niche integrations that will be able to be more accurate and helpful than using just as an ask anything machine. Looking forward to the future.
So ChatGPT is your primary way of finding information?
Most of what I go after is how to do stuff, not facts or data. Recipes etc, chat gpt is fairly useful for that. But less so as time goes by it seems.
If OpenAI can publish user stats, I believe the claim as long as they are open about their methodology.