That's good to see from GPT-4, but the comparison seems disingenuous, and I'd expect more from someone in academia.
That's good to see from GPT-4, but the comparison seems disingenuous, and I'd expect more from someone in academia.
It's "close enough" to be worth trying to see if it hits on something useful, but also bad enough that you really need to carefully scrutinise what it comes up with, and probably need to write a much more comprehensive prompt about strategies it should and shouldn't try to apply to make it more generally usable.
My point of view on this is that this type of AI cannot reason, so it seems to me quite audacious to use it for programming. Following the AI's suggestion is like examining a PR from someone you don't know: don't they try to insert a backdoor or a subtle bug?
Even if you use it as some kind of "rubber duck", examining its suggestions - doing the necessary reasoning and the double checks - is pretty time consuming. Wouldn't the long term result be better if one spent some time to acquire a deeper knowledge of the field, rather than play the "prompt engineering" game (of which the rules might change next month)?
It feels like we are in the equivalent of programming language's "honey moon" [1] phase with this kind of technology. People will eventually get over the "wow" effect and use it for what it is good at.
"Would it be better to acquire a deeper knowledge"--for the cases described above, actually no? I already have that deeper knowledge. I don't have the attention span (both because this stuff can be boring and because I am frequently randomized) and for a lot of this stuff the gate is actually not solving the problem but the physical act of typing, or even copying and changing. Copilot, at least, absolutely does create bugs, but in my experience they're mostly bugs of the form "this actually isn't typesafe in the way that it's written" and that makes isolating and fixing them clear. (Rarely are those bugs logic bugs, just a burp in the output.)
There is a very clear future where this turns into incredible copy-paste garbage because it's not capable of refactoring, and that's a different problem. Avoiding it is why they pay us the medium bucks.
but it can answer the glass door push or pull question (see https://news.ycombinator.com/item?id=35506666), albeit a bit off (it didn't understand the mirrored writing).
But some other models had gotten the right answer.
This points to some evidence of being able to derive logical reasoning. Just because the model is "doing statistical guessing of the next word", doesn't mean this process does not produce logical reasoning as an outcome.
In one of my other responses I dug into how I'd asked it to analyse a piece of assembly generated with a compiler I've written, from a test program that didn't exist at GPT4's knowledge cutoff, and it did just fine.
Being capable of logic and reasoning is more of an illusion that arises as a result of the enormous amount of training they put in. You can make simple questions that it can't answer, so you can't claim it has logical reasoning if it fails at simple tasks. It just looks a lot like it has logical reasoning because it has been trained on text displaying logical reasoning.
Getting it to symbolically evaluate code it most certainly have never seen before has also worked well for me, as have getting to explain the hypothetical outcome of executing a piece of code written using a made up combination of two languages. Frankly, I've worked with plenty of developers who'd struggle to reason about simpler tasks than some of the stuff I've successfully thrown at GPT4, so I'd argue that if it can't reason then many humans can't either.
It also fails at some very basic things. The biggest challenge is that it's reasoning is different to humans, and it's terrible at some things we consider basic even as it can do other things at a far higher level.
> Wouldn't the long term result be better if one spent some time to acquire a deeper knowledge of the field, rather than play the "prompt engineering" game (of which the rules might change next month)?
"Prompt engineering" works best in fields you do have a deep knowledge of, where you can tell whether you're actually getting better solutions. I'm willing to put effort into that too is acquiring deeper knowledge of the field by exploring dependencies and getting an understanding of ambiguities of language we use to talk about a problem. While you can get bogged down in specific "incantations", the more important part of "prompt engineering" is getting better at recognising under-specified problems. In fact, one of the things I've toyed with is to feed it explanations of a problem and instead of asking it to solve it, I first ask it to suggest improvements and ask questions, and that will often reveal problems with the prompts with little effort.
With respect to assembler, it seems to hit one of the bad spots, in that you get far fewer textual structural hints. E.g. a label can be the start of a function or the start of a basic block, or part of a loop, or part of an if/else block, and you largely need to recognise contextual clues and know the patterns of whatever produced the asm to know which rules applies. E.g. one of the issues I ran into was that the asm I gave it was output from a compiler that used %esi in a very specific way. There was no way for GPT4 to know that from the sample it was given without "thinking it through" (see below), and it made a "guess" that many novice asm programmers also likely would've, but where someone more experienced would've stopped and demanded a confirmation of the calling convention.
That said, to the test I mentioned in the previous comment, I tried giving it the code in a fresh chat, and asked it to reason about the code, and to give me a list of information that I could provide that would make it easier to suggest optimisations.
When I explicitly asked it about the calling convention used for %esi in the code I gave it, it correctly reasoned its way through it by observing how it was treated within the given code. That's something an smart beginner at asm might spot, and an intermediate asm programmer pretty likely would spot, but the important part is that it was clearly able to explain why with references to the code.
It also when asked correctly recognised that the object was tagged (low bit used to indicate it's not a "real" pointer but a value object), and recognised vtable calling pattern that few developers without compiler experience would pick up on.
Frankly, after asking it these follow up questions my opinions on its understanding of asm is way up. It won't beat an experienced asm developer, but most developers I can ask questions of are not experienced asm developers, and even fewer knows the patterns it easily explained and reasoned about.
Apparently you haven't been following many academics on Twitter.
And delivered with sarcasm and condescension. Not good to see, but par for the course on social media.