GPT-5 might arrive this summer as a "materially better" update to ChatGPT
arstechnica.com
arstechnica.com
> We’ll release in the coming months many different things. I think that’d be very cool. I think before we talk about a GPT-5-like model called that, or not called that, or a little bit worse or a little bit better than what you’d expect from a GPT-5, I think we have a lot of other important things to release first.
In other words, the next OpenAI GPT update(s) WILL NOT be GPT-5. Yes, we can expect improvements (because of course) but we shouldn't expect a major leap forward in AI the next couple of months from OpenAI.
When a CEO says "before we can talk about" it means "no". The language about other cool updates suggests incremental improvements or platform improvements, but not a major leap forward in terms of intelligence.
>We will release an amazing new model this year. I don’t know what we’ll call it.
And:
>So part of the reason that we deploy the way we do, we call it iterative deployment, rather than go build in secret until we got all the way to GPT-5, we decided to talk about GPT-1, 2, 3, and 4. And part of the reason there is I think AI and surprise don’t go together. And also the world, people, institutions, whatever you want to call it, need time to adapt and think about these things.
>We’ll release in the coming months many different things. I think that’d be very cool. I think before we talk about a GPT-5-like model called that, or not called that, or a little bit worse or a little bit better than what you’d expect from a GPT-5, I think we have a lot of other important things to release first.
When a CEO deflects when presented with an opportunity to hype the next major leap forward in AI that's not a good sign. He talks about "iterative development" to give "the world [...] time to adapt". He's telling people to temper their expectations.
> So part of the reason that we deploy the way we do, we call it iterative deployment, rather than go build in secret until we got all the way to GPT-5, we decided to talk about GPT-1, 2, 3, and 4. And part of the reason there is I think AI and surprise don’t go together.
This was presumably said with a straight face, which I find both impressive and unsettling.
On a sidenote, my first thought was someone was a StarTrek fan.
CLI one-liners or simple scripts are probably my main use-case. So many things I always knew were possible (and implemented in some cases in the past) but often avoided since I wasn’t sure if it would pay off that now I can do with ChatGPT. Simple things like “benchmark this endpoint and output the min/max/avg/p99 stats every 10 requests with a rolling window of 100 requests”, no it’s not hard to write that but often I just wouldn’t have to time to test out a theory so I’d skip this.
I’ve lost count (easily over 30+) times I’ve dealt with a production issue using ChatGPT to write something to better analyze the problem (parsing stuff out of logs often). A whole slew of “I wonder if…” questions that I’m able to answer in <2min vs having to spend 10-15+ if I did it all manually.
No, it’s not perfect and sometimes I need to go back and forth a few times but the overall speed up using ChatGPT is significant.
They nerfed it at Dev Day in November. It’s just no longer the magic tool it was before than anymore. It’s ok for maybe one shot problems or refactors of small functions but the version where you could in an hour work back and forth with it and build something moderately complex from scratch is dead and often some of the generations just waste time rather than solve the problem.
No it was not, and nor were its predecessors.
The only people who think GPT is "earth shattering" are rose-tinted-spectacle AI dreamers, and the world's media who love a good AI story.
Anyone sensible can see that GPT is merely the latest iteration in "AI will rule the world" nonsense. It comes round every 5–10 years, I remember most of the previous hyped up iterations, this one is no different, its just the hype is sadly sticker this time round. Its still hype though.
Just because you can, doesn't mean you should.
Companies using AI to automate "customer service" ? I encounter that regularly in my daily life these days. Its shit. The "intellligent" robot is always a useless waste of time.
I think in cases of new LLM implementations you can and should.
AI is and always will be supremely confident at being supremely wrong. Its the nature of the beast. Some things are just best left to real-life humans.
Sounds like you are also supremely confident at making predictions!
AI's are already less confident and more cautious than their previous iterations, so I'm not sure how you can confidently conclude it's entirely unsolvable.
This is not my first rodeo as they say.
I've seen it before. I've seen incredibly intelligent people before going all-in on the latest AI hype. And it always ends the same way, badly. The hype dies down and then returns again when the next iteration comes around.
I have a rather large spend across the universe of models and I think compared to a year ago, its amazing what is possible. If we continue anywhere close to this speed, it will be amazing what will be possible a year from now.
People said that a year ago and look today, still gpt-4 but now its turbo. So if we keep that pace in a year we will be at roughly gpt-4 but turbo++, better but not a giant leap.
I found that if I treat it as a program that can generate text, it's a lot easier to write the prompt that generates code that's more or less what I need.
It can do the boring stuff pretty easily, like generate boilerplate code or settings files. It can explain technical issues better than googling.
It also helps that I use a functional language with immutable data structures and the code it generates is either good or good enough to give me an idea on how to move forward.
Once you lower your expectations and use it for what it is, it's a very powerful tool.
*edit: I should specify that when I say LLMs are a collection of facts and heuristics I mean they are a collection of those things encoded as language which itself has been encoded as vectors of floats which in turn have modified the weights of the network to produce yet another encoding. I don't mean that the facts and heuristics are stored in a lookup table or as procedures.
Perhaps a system of mutually-criticising LLMs is what is needed to avoid confidently incorrect answers? Give them all a task to find errors in each others' outputs, and only return the output to the user when all of them agree on correctness.
At least most humans will know that most modern programming languages have built-in basic math functionality.
When I asked GPT-4 to write me a basic program in Go, it tried to tell me that a basic math function was not available in Go stdlib and proceeded to write some DIY horror-show of code as a substitute.
On the same occasion it also took great delight in importing obsolete libraries, and using deprecated functions.
And for the icing on the cake, the code it generated did not even compile !
Every time I corrected it, it returned with fresh code and a supremely confident assertion that "this fixes the problems identified". It was wrong, of course. Every, single, time.
All basic errors that even a Junior programmer fresh out of school would not make.
I doubt that most humans will know what those words even mean.
But pedantry aside, what I was trying to say is that people, without being corrected by others or without having self-criticism embedded in their way of thinking, make same mistakes as LLMs do. Think of an isolated programming novice who lives in his mom's basement and whose knowledge of programming comes only from stack overflow, random blogs and documentation he doesn't really understand. He'd probably give the same output as an LLM.
Considering the way LLMs work, I find them conceptually most similar to a stream-of-thought kind of thinking - which, in human cognition, is only the first step in reasoning. The other steps are observing what has been thought, detecting mistakes, fixing them, and doing so repeatedly until the thought is deemed to be correct.
I'm just saying maybe it's possible to emulate (to some level) that process of reasoning by connecting multiple LLMs together and giving them a task to criticize and correct each other.
But then again, you don't seem to be even acknowledging that point of mine. Maybe you're just angry/jaded/<insert negative emotion> at LLMs and feel the need to vent your frustration. And then again, maybe not. If you are, that's understandable, but I think that HN is not an appropriate arena for that - Twitter or 4chan would be more appropriate.
Hard to evaluate if this is a fair test without knowing what you asked of it.
Oh, right, the old "can't possibly be the AI, must be the user" claim.
The same excuse certain car manufactuers use when their car gets confused by shadows in the sun.
Give me a break.
https://chat.openai.com/share/371863ec-edbd-4454-8b19-382035...
Here is an example of the kind of scripting I do regularly with ChatGPT. I cannot speak to its capabilities with Go, but it is quite proficient at Python.
If you start getting into more esoteric edges of Python, like SharedMemory, ChatGPT quickly falls apart. Or, more relevant to your example, using carriage return to overwrite text for a progress indicator – works great, very simple. Until you try to use it through subprocess. To ChatGPT’s credit, it eventually came up with using pty, which with some massaging, I got to work.
I’m not saying it isn’t useful – far from it. It’s just that on anything modestly complicated, you sometimes have to spend more time fiddling with the prompt than if you just sat down and wrote the code. Another example that comes to mind was implementing a B+tree in pure Python. I know how they work, and wanted to see if ChatGPT could figure it out. You’d think so, right? But no, it kept getting stuck on node splits.
It doesn't mean you can turn your brain off and just mindlessly copy and paste whatever gets output, but for basics and boilerplate it's a massive time save, and I only foresee the capabilities getting better over time. It's just that comments like the one I was replying to come off as childishly throwing one's toys out of the pram because they are imperfect.
Future generation AI's are highly likely to have more capability though - I wouldn't bet the house on the current state being the future state.
If you want something real smart, what you need is a Casio calculator.
I regularly use GPT-4 for things like "write exhaustive unit tests for the following function" and it gives fantastic results. They usually need fleshed out a little, sure, but in a few seconds I get something that would have take about $45 in billable human hours.
I personally cannot wait to see the productivity gains that come from GPT-5.
> I personally cannot wait...
My point exactly.
You can get red in the face at AI the disembodied futurist concept all you want, but the tools are the real deal - people are using image generation and chatbots today to be 50%+ more productive and produce better work. These tools require skill to use, they're not just plug-n-play, but that skill pays huge dividends, and if you don't learn it you're going to be left so far behind.
I don't think we have evidence to support that.
So can we know which are which, and what are the relative value adds of these task groups? Where does an overall cost benefit analysis land, bearing in mind the high costs of AI.
If you could both speed up inference time and reduce cost of GPT to 3.5 levels, there are incredible amount of possibilities that open up that could actually help solve a lot of problems have had with trying to interface software to the real world (robotics).
4-Turbo is actually pretty amazing but its still a tad pricey for some tasks. Altman made a comment on Lex's podcast about compute and I think its true, essentially the world does not understand how much compute will be desired as the price of these things go down.
But with only the chat interface, our social media manager would have to find every new article in every discrete vertical on our site, copy all the URLs to a spreadsheet, then prompt ChatGPT for every article, wait for it to respond, then copy that response into the spreadsheet...
Presumably this would be a major part of the big material improvement we'd need to see.
I installed LibreChat on a Linode droplet, put $10 on my OpenAI account for API usage, and cancelled my ChatGPT subscription. Since neither I, nor my friends and family use ChatGPT a ton, I let them sign up on my server and the API costs are under $1 so far.
I configured the yaml file to use my own OpenAI API key for all users.
[0] https://docs.librechat.ai/install/configuration/user_auth_sy...
edit: oh, the only other configuration change that I made was to set the default OpenAI model to gpt-4-0125-preview.