But. Every single test I ran lead to functionally wrong designs from smallish memory errors in C (that hilariously ChatGPT was able to correct ND explain when pointed to) over misplaced/hallucinated methods in python APIs to completely hallucinated perl packages. It never gave me a useful answer.
I don't know if that tells us something about my work vs. your work or my approach to the model vs yours. But it definitely tells me that such a model cannot replace a developer.
Even Copilot, which is a smaller model, generates working code most of the time.
Certainly not an average developer (however that's decided in C's culture). Not making any statements on skill of anyone here, but I'd bet it could replace the very low ends of skill as it is (GPT-4 and maybe Bard). I doubt there's too many of those kinds of developers though.
It's not a guaranteed positive for all languages/people, but it's farther along being useful as an enhancement technology than it is as replacement technology. I think it's making progress on both axes, but one is clearly ahead of the other and it's still possible it stops progress until other innovations are made.
I was legitimately afraid for my job when GPT-4 was released. After using it, however, I sleep easy now.
Things like ChatGPT might get 90% of the way there very, very quickly for many use cases (similar to how lane keeping and adaptive cruise control is 90% of “self-driving”) but the remaining 10% of edge cases will require humans for a very long time (similar to how we never quite managed to get “100% self driving” because the vehicles still can’t respond reliably 100% of the time). But we still value it because the 90% works.
Once people understand that 10% limitation and once the 90% value has been gained/normalized/integrated, the hype will die down and we’ll be watching a slow grind to figure out how to “fix” the last 10% which is necessary to live up to what’s being hyped today.
Just wait until GPT-5, though. It will blow your mind... \s
I was using it to try and help me create Arma3 scenarios, which often involves using an ugly little DSL called SQF.
The last thing I asked it about was how to play a custom audio file sitting on my disk. It gave me a simple one-liner, that didn’t do anything. I asked it several questions to clarify, but it seemed confident that all I needed was this one liner to play an audio file sitting anywhere on my disk.
Eventually I said fuck it, and decided to google it myself, and it turns out it’s much more involved than just a single function call with a path as an argument, as suggested by ChatGPT. In fact you have to create an entirely separate config file with classes to represent the audio, and then you can reference that class to play the audio, not just a simple string path.
I’ve had it fail with a lot of other things, mostly trying to make computer controlled units perform deterministic, scripted actions, but it’s hard to tell if GPT-4 or the game’s engine was at fault here, because in my experience making Arma’s AI units do anything reliably is near impossible, no matter how much scripting you do.
Only the next day when I still couldn't wrap my head around why a real time processing system like be designed like this and did more google searching did I find it explicitly spelled out in the docs that a batch read counts as only a single read operation and can retrieve up to 10,000 records.
It's the first time I've been bit by this and it will definitely make me wary of any information it provides me in the future for technologies I am not intimately familiar with (at which point it's unlikely I'd need to consult it in the first place).
Technology is supposed to make us smarter. Blindly believing an AI that we know can hallucinate makes us dumb with confidence.
This keeps being my argument when people at work daydream about time and cost savings by offloading non-critical business functions to AI. I say, "Great, so it can produce 1000x more work than a person. But then what army of people are we planning to use to check those outputs?"
I'm super-impressed with the current crop of language models for their ability to so accurately simulate correctness, but their inability to understand what they don't know - because, in fact, they don't 'know' any of it in the sense that we do - makes them like very productive but completely untrustworthy employees. A junior dev who monopolizes his mentor's time through inconsistent performance is not a good hire.
Have you ever had a dumb/wrong thought in your head? I'm going to go ahead and answer yes for you, you do all the time. In fact you don't (hopefully) verbalize a stream of consciousness to other people around you. In general you think of something then reflect on what it is true/false.
This is not what LLMs do, they pitch back the first 'thought' they have, "correct" or not. This is why things like COT/TOT greatly increase the accuracy of LLM output. The problem? It requires at least an order of magnitude more processing to get an answer, and with GPU time already in high demand and expensive you don't see much of it happen.
Betting on LLMs commonly being wrong is not a safe bet at this point.
It's like reviewing an overconfident junior developer's code except you can't learn their particular weaknesses. If a developer is bad about memory leaks, you know to check their every PR for memory leaks. An LLM won't necessarily produce the same types of errors given similar prompts or even the same prompt with some period of time between invocations.
https://arxiv.org/pdf/2305.08291.pdf
https://github.com/jieyilong/tree-of-thought-puzzle-solver
In this paper, we introduce the Tree-of-Thought (ToT) framework, a novel approach aimed at improving the problem-solving capabilities of auto-regressive large language models (LLMs). The ToT technique is inspired by the human mind’s approach for solving complex reasoning tasks through trial and error. In this process, the human mind explores the solution space through a tree-like thought process, allowing for backtracking when necessary. To implement ToT as a software system, we augment an LLM with additional modules including a prompter agent, a checker module, a memory module, and a ToT controller. In order to solve a given problem, these modules engage in a multi-round conversation with the LLM. The memory module records the conversation and state history of the problem solving process, which allows the system to backtrack to the previous steps of the thought-process and explore other directions from there. To verify the effectiveness of the proposed technique, we implemented a ToT-based solver for the Sudoku Puzzle. Experimental results show that the ToT framework can significantly increase the success rate of Sudoku puzzle solving.
Before, I'd often end up going down the wrong rabbit hole when I would try to write out my problem in a google-able manner.
https://arxiv.org/abs/2304.13187
Generally, chatbots function in direct communication with humans and do not function autonomously. Thus they are dependent on not just human meaning-making but human thirst for meaning which will find it even when it isn't there.
The latter doesn't follow. They may just be dependent on good prompt design.
People who aren't used to the AIs, and don't know how to use them, get worse results. This shouldn't be a surprise; it's how every tool works.
That’s why people’s first impressions are often so wildly off. People who have the fortune of asking the right prompts for their first conversations walk away believing they’ve spoken to AGI, people who ask a less than optimal prompt (or things it is bad at like math problems) walk away not seeing what the hype is about.
Same thing with copilot. The quality of the code that it suggests depend a lot on the context, so if someone isn’t writing good comments and giving their variables proper names the suggested code will be terrible.
These systems are very capable, but using them is a skill that must be learned, and there’s little material out there to teach that skill because things move so quickly.
If we didn’t move the goalposts, we’d have declared Stockfish to be full AI, despite it only being a chess-playing program, long ago.
Push out a plugin that sets up a virtual board GPT-4 can read after each move and see if its any better.
My point is that the hyperbole about “we’ve cracked AI but they changed the goalposts” is self evidently not true. You just proved it right there: I need to add more plugins because ChatGPT does not understand what it’s doing.
It’s a potentially useful tool but that doesn’t make it intelligent.
Being able to properly use tools is a sign of intelligence in of itself.
Where is this threshold? I don’t think anyone really knows yet. That’s why I don’t see a problem with “moving the goalposts,” because doing so is the best way to help us truly understand what it means to be an intelligent life form (artificial or otherwise).
And again "reasonable" is a pretty useless metric. Asking a 'reasonable' person about any system that requires expert knowledge to understand is going to derive an unreasonable answer. This is because they'll conflate intelligence with human behavior.
This is just silly. Also, a machine killing me is still not enough. We have drone targeting systems in development right now but I don’t think you’d call them AI.
> And again "reasonable" is a pretty useless metric.
When you can’t define intelligence, it’s really the starting point.
The thing with LLMs is that they look almost there but, as many are pointing out, the method by which they make inferences is an analysis of how words fit together without understanding the meaning. This is why things like ChatGPT confidently spout absolute nonsense about topics they weren’t trained on (with much human intervention): it doesn’t know what it’s saying so it doesn’t realise it’s making stuff up.
I'm convinced that when the dust has settled and historians look back to decide on THE point in time at which we achieved AI or even AGI, that time will not be in the future, but in the past.
No, we didn’t. There was plenty of disagreement over it. Heck, the Chinese Room is very popular rejection fundamentally of the premise of it.
That aside, even if in a blind scenario (where the builders didn’t know the criteria used to test) being able to fool humans in linguistic interaction would be reasonably likely to be a good test of general intelligence, LLM’s are about as an obvious of a direct and deliberate attempt to Goodhart’s Law the Turing Test as one could imagine.
Using a particular capacity to test a more general capacity is obviously vulnerable to systems built to specialize in the tested capacity, as opposed to those that have it as a consequence of general ability.
If historians are looking back to decide the point, by definition doesn't it have to be in the past? Historians looking back to the future doesn't make sense...
The conventional narrative is that chatGPT can code as well as a junior engineer, but I feel like it’s more like a senior engineer scribbling on a whiteboard—no, it’s not perfect code, but it’s the right idea.
I've been using github copilot and I'm waving back and forth between impressed and unimpressed; it regularly generates syntactically invalid code for me. Eg calling lib functions that don't exist.
Thanks!