ML code generation vs. coding by hand: what we think programming will look like
wasp-lang.dev
wasp-lang.dev
async function isPositive(text: string): Promise<boolean> {
const response = await fetch('https://text-processing.com/api/sentiment', {
method: "POST",
body: `text=${text}`,
headers: {
"Content-Type": "application/x-www-form-urlencoded",
},
});
const json = await response.json();
return json.label === "pos";
}
This code doesn't escape the text, so if the text contains the letter '&' or other characters with special meanings in form URL encoding, it will break. Moreover, these kinds of errors can cause serious security issues; probably not in this exact case, the worst an attacker could do is change the sentiment analysis language, but this class of bug in general is rife with security implications.This isn't the first time I've seen this kind of bug either -- and this class of bug is always shown by people trying to showcase how amazing Copilot is, so it seems like an inherent flaw. Is this really the future of programming? Is programming going to go from a creative endeavor to make the machine do what you want, to a job which mostly consists of reviewing and debugging auto-generated code?
Not the one I want. I'd like to see better theorem proving, proof repair, refinement and specification synthesis.
It's not generating code that seems to be the difficulty, it's generating the correct code. We need to be able to write better specifications and have our implementations verified to be correct with respect to them. The future I want is one where the specification/proof/code are shipped together.
Also, it's not actually writing the code that is the difficulty when generating code changes for maintaining existing code, but figuring out what changes need to be made to add up to the desired feature, and where in the codebase controls each of those, and possibly rediscovering important related invariants.
For that matter, the really critical part of implementing a feature isn't implementing the feature -- it's deciding exactly what feature if any needs to be built! Even if many organizations try to split that among people who do and don't have the title of programmer (eg business analyst, product manager...), it's fundamentally programming and once you have a complex enough plan, even pushing around and working on the written English feature requirements resembles programming with some of the same tradeoffs. And we all know the nature of the job is that even the lowliest line programmer in the strictest agile process will still end up having to do some decision making about small details that end up impacting the user.
Copilot is soooooo far away from even attempting to help with any of this, and the part it's starting with is not likely to resemble any of the other problems or contribute in any way to solving them. It's just the only part that might be at all tractable at the moment.
As interesting as those systems may be, they have it exactly backwards. Don't let the computer generate code and have the human check it for correctness: let the human write code and have the computer check it for correctness.
The human also needs to make sure the code does what the Ticket asked for. And sometimes not even what the Ticket asked for, but what the business and user want/need.
Honestly writing code is such a small part of a dev's job. It's figuring out how to translate the human (often non-tech) requirements into code that is also in a way that's also looking at requests that haven't been made yet.
Isn't that pretty much what ML does?
> ... I will not delve into the questions of code quality, security, legal & privacy issues, pricing, and others of similar character. ... Let’s just assume all this is sorted out and see what happens next.
Honestly, it seems pretty dangerous. Seemingly correct code generated by a neural network that we think probably understands your code but may not. Easy to just autocomplete everything, accept the code as valid because it compiles, and discover problems at run-time.
Formal methods and theorem proving on the other hand, place a strong emphasis on proof production and verification.
Instead, they decided to waste their work on that.
AWS has ML driven code analysis tools in their "CodeGuru" suite: https://aws.amazon.com/codeguru/
More often than not, once I have the tests written, the code nearly writes itself. There are edge cases when I have to look up things, but most of the time it's I'm happily TDD'ing.
A couple of years ago, we invested in using the libraries from Microsoft Code Contracts for a couple of projects. It was a really promising and interesting project. With the libraries you could follow the design-by-contract paradigm in your C# code. So you could specify pre-conditions, post-conditions and invariants. When the code was compiled you could configure the compiler to generate or not generate code for these. And next to the support for the compiler, the pre- and post-conditions and invariants were also explicitly listed in the code documentation, and there was also a static analyzer that gave some hints/warnings or reported inconsistencies at compile-time. This was a project from a research team at Microsoft and we were aware of that (and that the libraries were not officially supported), but still sad to see it go. The code was made open-source, but was never really actively maintained. [0]
Next to that, there is also the static analysis from JetBrains (ReSharper, Rider): you can use code annotations that are recognized by the IDE. It can be used for (simple) null/not-null analysis, but also more advanced stuff like indicating that a helper method returns null when its input is null (see the contract annotation). The IDE uses static analysis and then gives hints on where you can simplify code because you added null checks that are not needed, or where you should add a null-check and forgot it. I've noticed several times that this really helps and makes my code better and more stable. And I also noticed in code reviews bugs due to people ignoring warnings from this kind of analysis.
And finally, in the Roslyn compiler, when you use nullable reference types, you get the null/not-null compile-time analysis.
I wish the tools would go a lot further than this...
[0] https://www.microsoft.com/en-us/research/project/code-contra... [1] https://www.jetbrains.com/help/resharper/Reference__Code_Ann...
Thanks for making me smile. Some of us are working on exactly that kind of thing :)
Something along lines of this perhaps...
For me, this is definitely a piece of the future of coding, and it doesn't change coding from a creative endeavour. It's just the difference between "let me spend five minutes googling this canvas API pattern I used 5 years ago and forgot, stumble through 6 blogs that are garbage, find one that's ok, ah right that's how it works. Now what was I doing?" To just writing some prompt comment like "// fn to draw an image on the canvas" and then being like "ah that's how it works. Tweak tweak write write". For me the creativity is in thinking up the high level code relationships, product decisions, etc. I don't find remembering random APIs or algos I've written a million times before to be creative.
Even bugs are acceptable if the advantages of auto generated code are large enough.
My concern is that ML code generation will create very shallow mental models of code in developers' minds. If everyone in the organization is building code this way, then who has the knowledge and skills to debug it? I foresee showstopping bugs from ML-generated that bring entire organizations to a grinding halt in the future. I remember back in the day when website generators were in vogue, it became a trend for many webdevs to call their sites "handmade" if they typed all the code themselves. I predict we'll be seeing the same in the future in other domains to contrast "handwritten" source with ML generated source.
What's that famous quote? "Debugging is twice as hard as writing the code in the first place." Well where does that leave us when you don't even have to write the code in the first place? I guess then we will have to turn to ML to handle that for us as well.
I think it’s more like WYSIWYG website generators (anyone remember Cold Fusion? The biz guys LOVED that - for like a year).
I think most companies' management would still leave decisions like that in the hands of developers/development-teams. At the level of management, I think companies aren't asking for "code" but for results (a website, an application, a feature). Results with blemishes or insecurities are fine but a stream of entirely broken code wouldn't be seen as cost effective.
The problem always comes later (often after the manager has moved on and is no longer aware of the mistakes that were made and the repercussions can't follow) where the maintenance costs of poorly written initial code come back to haunt the origination for years to come.
Which I personally suspect this will end up about as well as WYSIWYG HTML editors/site generators. Which do have a niche.
That would be the ‘makes a lot of cheap crap quickly, but no one uses it for anything serious or at scale’ niche.
Because what comes out the other end/under the hood is reprocessed garbage that only works properly in a very narrow set of circumstances.
I suspect with this market correction however we’ll see a decrease in perceived value for many of these tools, as dev pay is going to be going down (or stay the same despite inflation), and dev availability is going to go up.
That the monkey is actually a computer and not an actual banana eating primate doesn't change anything there.
Any business fool that tries this will discover this very quickly - that's like claiming we don't need drivers because we have cruise control in our cars! It is not perfect but it is cost effective/works (some of the time ...)!
Do it enough times, and I’m sure it will eventually compile - might even complete a regression test suite!
For application/x-www-form-urlencoded as shown, that’d be using URLSearchParams:
fetch("https://text-processing.com/api/sentiment", {
method: "POST",
body: new URLSearchParams({ text }),
})
If you wanted multipart/form-data instead (which I expect the API to support), you’d use FormData, which sadly can’t take a record in its constructor: const body = new FormData();
body.append("text", text);
fetch("https://text-processing.com/api/sentiment", { method: "POST", body })
Note that in each case the content-type header is now superfluous: the former will give you "application/x-www-form-urlencoded;charset=UTF-8" (which is all the better for specifying the charset) and the latter a "multipart/form-data; boundary=…" string. (Spec source: https://fetch.spec.whatwg.org/#concept-bodyinit-extract.)As a fun additional aside, you’ve made a tiny error in your transcription: the URL was actually surrounded in backticks (a template literal), not single quotation marks. Given the absence of a tag or placeholders (which would justify a template literal) and the use of double-quoted strings elsewhere, both of these would be curious choices that a style linter would be very likely to complain about. So yeah, just another small point where it’s generating weird code.
Yes, many companies prefer to hire them, or can only retain them for a variety of reasons. None of this is good in any way.
Anyway, those programmers have really good odds to stop creating this kind of code given some learning. While adding Copilot as a "peer" just makes them less likely to learn, and all programmers more likely to act like them. That's not matching the status-quo, that's a very real worsening of it.
Security holes are more excusable because someone who didn't realize the above could happen maybe never tested it... given the use case, this is more like "did you even run your code?"
I think it's because people copy/paste generated code without reading it carefully. They eyball it, it makes sense, they go tweet about it.
I don't know if this predicts how people will mostly use generated code. I note however that this is probably too much code to expect CoPilot to generate correctly: about 10 LoCs is too much for a system that can generate code, but can't check it for correctness of some sort. It's better to use it for small code snippets of a couple of lines, like loops and branches etc, than to ask it to genrate entire functions. The latter is asking for trouble.
These are the examples which GitHub itself uses to demonstrate what Copilot is capable of, so it's not just a matter of people tweeting without reading through the code properly. It also suggests that the people behind Copilot do believe that one primary use-case for Copilot is to generate entire functions.
Well, OpenAI are certainly trying to sell copilot as more capable than it is, or anyway they haven't done much to explain the limitations of their systems. But they're not alone in that. I can't think of many companies with a product they sell that tell you how you _can't_ use it.
Not to excuse misleading advertisment. On the contrary.
Maybe, in some organizations, sure. However, there are still people hand-crafting things in wood, metal, and other materials even though we have machines that can do almost anything. Maybe career programmers will turn into "ML debuggers", so perhaps all of us who enjoy building things ourselves will just stop working as programmers? I certainly won't work in the world where I'm just a "debugger" for machine-generated crap.
function evil(code) {
eval(code);
}
evil("alert('hi');");Taking your comment as inspiration, I would like to add OWASP/Unit Testing thinking to Copilot. If Copilot considers your remarks and maybe others, it becomes helpful. Something like security checking on the fly, what would normally be considered by colleagues during code reviews or checks with SonaCube.
> Let’s just assume all this (code quality, security, legal & privacy issues, pricing, and others of similar character) is sorted out and see what happens next.
What could go wrong next in the systems where they employ this AI with this mindset?
Why? POST data doesn't need to be URL encoded.
For tasks like self-driving or spotting cancer in x-rays, they are producing novel result because these kinds of tasks are amenable to reinforcement. The algorithm crashed the car, or it didn't. The patient had cancer, or they didn't.
For tasks like reproducing visual images or reproducing text, it _seems_ like these algorithms are starting to get "creative", but they are not. They are still just regurgitating versions of the data they've been fed. You will never see a truly new style or work of art from DALL-E, because DALL-E will never create something new. Only new flavors of something old or new flavors of the old relationships between old things.
Assuming that it is even possible to describe novel software engineering problems in a way that a machine could understand (i.e. in some complex structured data format), software engineering is still mostly a creative field. So software engineering isn't going to performed by machine learning for the same reason that truly interesting novels or legitimately new musical styles won't be created by machine learning.
Creating something new relies on genuine creativity and new ideas and these models can only make something "new" out of something old.
If you trained an ML model on recognizing patterns of older less efficient algorithms and swapping in new ones, that could be useful. It probably isn’t that much harder of a problem than knowledge graph building and objecg recognition.
Also, having an ML model trained to recognize certain types of problem patterns in whatever source product/model/codebase or whatever and then spit out code to solve them is certainly not a theoretically impossible or low value task.
That said, ML as an art or science is no where near being able to do that usefully near as I can tell.
Hell even basic chatbots still suck.
Hell, even getting Hibernate to usefully reverse engineer a pretty normal schema still sucks.
Machine Learning algorithms are never better than the data they're trained-on. But they can easily be "worse".
Specifically, an ML algorithm trained on one data set can have a hard time operating on a different data set with some similarities and some fundamental differences. This is related to algorithms generally not having an idea how accurate their predictions/guesses are. This in turn relates to current ML as being something like statistical prediction with worries about bias tossed-out (which isn't to dismiss it but to illuminate it's limitations).
For tasks like self-driving or spotting cancer in x-rays, they are producing novel result because these kinds of tasks are amenable to reinforcement. The algorithm crashed the car, or it didn't. The patient had cancer, or they didn't.
Ironically, both those applications have been failures so-far. Self-driving is far more complex than a binary crash-or-not scenario (usually the road is predictable but just about anything can wander into the roadway occasionally and you need to deal with this "long tail"). It also requires extreme safety to reach human levels. Diagnosis by image has problems of consistence, of doctors considering more than just an image and other things possibly not understood.
Computers exceeded human chess capacities without any AI as such in 1999 or so. Computer exceeded human addition and multiplication capacities as soon as they were built - that's what they were built for.
To be pedantic, that was without any machine learning. Deep Blue was absolutely considered an AI system at the time. Tree search as used in Deep Blue is also still central to Alpha* family of game-playing systems.
Alphago basically generated its own training data (hence, "exceeding" is not correct) because the domain was amenable to a specific form of reinforcement learning. Programming isn't such a domain.
Alphago is proof that ML systems can exceed human abilities in at least some domains. Other domains, like identifying animal breeds/species by sight, I believe ML models are also super human. It's just not correct to say models are limited by their training data to something less than human quality. We have multiple counterexamples already.
Well, this is obviously incorrect because the training data included all the games it played against itself (good luck replicating that in the domain of programming though).
We were not talking about "exceeding human abilities in at least some domains", we were talking about whether ML (DL really) systems were able to "go beyond" their training data - but your deflection and going off point is noted.
If you're trying to argue that the model can't be better than it is, i.e. the last training data it produces, then that's true but tautologically so. ML has conquered some domains, like Go or Chess or birdwatching, and is making progress in others - like programming. There is no reason to assume ML systems can't exceed human ability at programming or any other task.
More deflection and going off point.
>is making progress in others - like programming
I already explained that programming is not amenable to the approach that allowed alphago to learn from the games it played against itself.
This is already getting pointless and tedious, should in the future know better than to respond to people that are obviously in the "scale is all you need" nonsense camp, nothing meaningful comes out of it.
It’s formalized statistics translated to code by community.
Regex and grep are just as “black box magic” to those who read a blog post on how to scan some text files.
ML is statistics on arrays, whereas Unix tools are stats for a Unix system.
It’s the same abstraction at its core, new implementation.
if (personInFrontOfCar) { halt(); }
The reason, of course, is the self-driving algorithms are all done with machine learning. The model is trained with a corpus of information and the result is a "black box." The programmer could not insert the above `if` statement if they wanted to, they could not modify the source code of the model as if it was normal business logic. Instead, they would have to re-train the model on new data.This is what people mean when they say it is a "black box." There is no source code to edit, only training data that outputs an as-is model. It's similar to getting binary blobs, for eg Linux drivers.
In contrast, I could fork regex if I wanted and add an `if` statement in the middle of it, and re-compile it. It is pure source code. Same with grep.
`if personInFrontOfCar:`
(using object detection, masking, 3D pose estimation, et/or cetera)
And actually defined policies on how to handle certain situations given that inference are possible, in the case specifically of self driving vehicles.
My assumptions come from when I'd asked someone at NeurIPS in 2019 about what reinforcement learning methods they use as they said "I don't think any self driving companies are using reinforcement learning, it is too risky". I don't mean to imply this is hard to find information either, reading a few papers would be all it takes to clear up the degree to which most self driving policies are "controllable" in the way you describe, or at least to what degree.
My main point is that I think it is the inferring of the environment (current state) rather than the chosen policy at each time step which is more of an error prone black box.
> Trying to solve the lack of progress in unifying computer architectures needs to be the next step in computer science. I just don't understand how or why it needs to be polluted by the unnecessary added hidden complexity of GCC. You're just shifting your assembly language generation into a black box and its magic output.
(I don't actually know how to write assembly, so pardon any technical faux pas. But the comparison seemed apt. ML as an intent->code compiler just seems like an inevitable step forward.)
When I code something out, I have to first know which problem to solve, and why it’s a problem. I then have to understand the problem in minute detail. This usually involves a slew of human factors, desires, and interfacing with idiosyncratic systems that evolved a certain way because of human factors and desires.
With the current way of ML code generation, it seems that there will always need to be a critical mass of coders who produce code on their own, to be able to serve as new input for the ML models, to be able to learn about new use cases, new APIs, new programming languages, and so on.
ML code generation may serve as a multiplier, but it raises questions about the creation and flow of new knowledge and best practices.
The day before Copilot launched, if someone had told me code generation at the fairly decent quality Copilot achieves was already possible I probably wouldn't believe it. I could happily rattle off a couple of arguments for why it might be better than autocomplete but could never write a full function. Then it did.
Who can say how far it's come since then? I think only the Copilot team knows. I wish we could hear from them, or some ML experts who might know?
Program Synthesis
https://www.semanticscholar.org/paper/Program-Synthesis-Gulw...
Copilot and similar systems (like DeepMind's AlphaCode) are a step backwards compared with earlier systems most of which are capable of not only generating code, but generating correct programs, where correctness is measured with respect to a specification, which can be complete or incomplete (as in input/output examples etc).
Systems like Copilot instead use a large language model to generate code by completing an initial "prompt" and have no way to control what kind of code is generated or whether it does what the programmer expects it to do. The task of checking correctness is left to the user. The result is code that may or may not contain bugs to an extent difficult to predict beforehand.
Specifically about Copilot, it's based on Codex, a large language model that's a variant of GPT-3 fine-tuned on all of github. In formal testing, OpenAI found that Codex could generate the correct program about 30% of the time, which is pretty low:
Evaluating Large Language Models Trained on Code
https://arxiv.org/abs/2107.03374
Interestingly, if Codex was allowed 100 "guesses", the correct program was in there 75% of the time, instead. But that only goes to show that these kinds of systems are code firehoses without any ability to tell a correct program from garbage.
I remember hearing “a million monkeys with a million type writers writing for a million years could write Shakespeare”. Sure, but how would they know it? How could they recognize when they have achieved their goal?
ML will lower the bar for the trivial tedious things and let people believe they are more capable than they are. Even with ML you will have to know what to ask and what the answer should be. That is always the hard part.
I guess if it did one thing it was reinforce something I already knew about my job - writing code is always the easy part.
I think this is a misconception. It’s true for these first prototype code generation tools, but there’s no reason to think that in the future these models won’t be adapted to modify/maintain code too.
Consider the instructions "take a list, add the number three and five, return the list"
This compiles to both
f(list):
3+5;
return list;
and f(list):
list.add(3&5);
return list;
and f(list):
list.add(3);
list.add(5);
return list;
Decoding this type of vague descriptions is something human programmers struggle with, often resulting in a discussion with whoever wrote the requirements about just what the requirements say (sometimes they don't know and it needs to be worked out).[three code segments]
Did you intentionally do an off-by-one error?
It's as they say: Time flies like an arrow, fruit flies like a banana.
If there was high-level intent (e.g. something like spec language) that was preserved and versioned, that might be possible I guess.
Again, always good to remember the thing in the background: This is not a philosophical debate, it's a legal one. Intellectual property, for better or worse, is a made-up concept by humans to try to encourage more creation.
*for the internet, it would simply transfer literally all the power to the youtubes and spotifys.
The point is that this massaging is a lot faster than writing everything from 0 in most cases. It doesn't have to be bug free. It doesn't have to be perfect quality. It doesn't have to do "the hard parts" of coding. It just has to be good enough to be sufficiently faster for the developer than starting from 0, and it is.
Perhaps the best approach could use ML within the constraints provided by a type system. Does research like this exist?
- the average developer tenure now is very short
- most software companies just want features, features and more features due to market pressure
- writing new systems is much more popular than maintaining or upgrading old ones.
So maybe this is the future. But it's kind of a sad one.
"Let's assume magic ponies exist, are commonplace and would love to give us all magic rainbow rides"
The top comment on this post is complaining about a bug in Copilot generated code. 1 million times better Copilot won't do that.
Perhaps there is research on providing mathematical guarantees and constraints on complex models like neural networks? If not, it feels like it would be harder to give a model a high degree of control. Although embedding the model in a system that did provide guarantees (e.g. using a compiler) might be a pragmatic solution?
Yeah man, I'm not sure about the latter. Not sure...
> allows us to write even less code and care about fewer implementation details
Remember Bjarne Stroustrup: “I Did It For You All…”? More code, more complexity — more job security.