AI is and always will be supremely confident at being supremely wrong. Its the nature of the beast. Some things are just best left to real-life humans.
AI is and always will be supremely confident at being supremely wrong. Its the nature of the beast. Some things are just best left to real-life humans.
Sounds like you are also supremely confident at making predictions!
AI's are already less confident and more cautious than their previous iterations, so I'm not sure how you can confidently conclude it's entirely unsolvable.
This is not my first rodeo as they say.
I've seen it before. I've seen incredibly intelligent people before going all-in on the latest AI hype. And it always ends the same way, badly. The hype dies down and then returns again when the next iteration comes around.
I have a rather large spend across the universe of models and I think compared to a year ago, its amazing what is possible. If we continue anywhere close to this speed, it will be amazing what will be possible a year from now.
People said that a year ago and look today, still gpt-4 but now its turbo. So if we keep that pace in a year we will be at roughly gpt-4 but turbo++, better but not a giant leap.
I found that if I treat it as a program that can generate text, it's a lot easier to write the prompt that generates code that's more or less what I need.
It can do the boring stuff pretty easily, like generate boilerplate code or settings files. It can explain technical issues better than googling.
It also helps that I use a functional language with immutable data structures and the code it generates is either good or good enough to give me an idea on how to move forward.
Once you lower your expectations and use it for what it is, it's a very powerful tool.
*edit: I should specify that when I say LLMs are a collection of facts and heuristics I mean they are a collection of those things encoded as language which itself has been encoded as vectors of floats which in turn have modified the weights of the network to produce yet another encoding. I don't mean that the facts and heuristics are stored in a lookup table or as procedures.
Perhaps a system of mutually-criticising LLMs is what is needed to avoid confidently incorrect answers? Give them all a task to find errors in each others' outputs, and only return the output to the user when all of them agree on correctness.
At least most humans will know that most modern programming languages have built-in basic math functionality.
When I asked GPT-4 to write me a basic program in Go, it tried to tell me that a basic math function was not available in Go stdlib and proceeded to write some DIY horror-show of code as a substitute.
On the same occasion it also took great delight in importing obsolete libraries, and using deprecated functions.
And for the icing on the cake, the code it generated did not even compile !
Every time I corrected it, it returned with fresh code and a supremely confident assertion that "this fixes the problems identified". It was wrong, of course. Every, single, time.
All basic errors that even a Junior programmer fresh out of school would not make.
Hard to evaluate if this is a fair test without knowing what you asked of it.
Oh, right, the old "can't possibly be the AI, must be the user" claim.
The same excuse certain car manufactuers use when their car gets confused by shadows in the sun.
Give me a break.
https://chat.openai.com/share/371863ec-edbd-4454-8b19-382035...
Here is an example of the kind of scripting I do regularly with ChatGPT. I cannot speak to its capabilities with Go, but it is quite proficient at Python.
If you start getting into more esoteric edges of Python, like SharedMemory, ChatGPT quickly falls apart. Or, more relevant to your example, using carriage return to overwrite text for a progress indicator – works great, very simple. Until you try to use it through subprocess. To ChatGPT’s credit, it eventually came up with using pty, which with some massaging, I got to work.
I’m not saying it isn’t useful – far from it. It’s just that on anything modestly complicated, you sometimes have to spend more time fiddling with the prompt than if you just sat down and wrote the code. Another example that comes to mind was implementing a B+tree in pure Python. I know how they work, and wanted to see if ChatGPT could figure it out. You’d think so, right? But no, it kept getting stuck on node splits.
It doesn't mean you can turn your brain off and just mindlessly copy and paste whatever gets output, but for basics and boilerplate it's a massive time save, and I only foresee the capabilities getting better over time. It's just that comments like the one I was replying to come off as childishly throwing one's toys out of the pram because they are imperfect.
Future generation AI's are highly likely to have more capability though - I wouldn't bet the house on the current state being the future state.
I doubt that most humans will know what those words even mean.
But pedantry aside, what I was trying to say is that people, without being corrected by others or without having self-criticism embedded in their way of thinking, make same mistakes as LLMs do. Think of an isolated programming novice who lives in his mom's basement and whose knowledge of programming comes only from stack overflow, random blogs and documentation he doesn't really understand. He'd probably give the same output as an LLM.
Considering the way LLMs work, I find them conceptually most similar to a stream-of-thought kind of thinking - which, in human cognition, is only the first step in reasoning. The other steps are observing what has been thought, detecting mistakes, fixing them, and doing so repeatedly until the thought is deemed to be correct.
I'm just saying maybe it's possible to emulate (to some level) that process of reasoning by connecting multiple LLMs together and giving them a task to criticize and correct each other.
But then again, you don't seem to be even acknowledging that point of mine. Maybe you're just angry/jaded/<insert negative emotion> at LLMs and feel the need to vent your frustration. And then again, maybe not. If you are, that's understandable, but I think that HN is not an appropriate arena for that - Twitter or 4chan would be more appropriate.
If you want something real smart, what you need is a Casio calculator.