The End of Programming
cacm.acm.org
cacm.acm.org
> Matt Welsh (mdw@mdw.la) is the CEO and co-founder of Fixie.ai, a recently founded startup developing AI capabilities to support software development teams. He was previously a professor of computer science at Harvard University, a director of engineering at Google, an engineering lead at Apple, and the SVP of Engineering at OctoML. He received his Ph.D. from UC Berkeley back in the days when AI was still not playing chess very well.
Don't the arguments stand on their own, independent of who made them?
Like sure, it’s in his interest to portray AI as being powerful. But this article felt pretty candid about what the effects of that power could be.
Seriously, this is the text from the home page of Fixie.ai
"We're setting out to change the way the world builds software, using AI as a foundation. We're founded by a team from Google and Apple with expertise in AI, systems, and the web. We're funded by Zetta Venture Partners, SignalFire, Bloomberg Beta, and others. We're hiring for multiple roles."
Yes, he is a highly competent and experienced individual, but so were legions leading up to the AI winter. Can't place much faith in someone who has something to sell with much as confidence as a ChatGPT response - an academic no less - in regards to the near-term future of this space.
I’m not yet sold on the idea that it will replace coding. It’s like autonomous cars. 99% is far far too low. But it’s also not to be dismissed. A brilliant tool with many use cases.
I think we have invented the Enterprise “ship’s computer.” It won’t answer “how do I solve this problem?” But it probably will answer things like, “if X and Y and Z, what might W be?” We get to be the big brains standing around a terminal! I can’t wait to have a holodeck where I can construct a world just by asking for things.
Which is also the problem with any AI. If the AI can program, then not only does it have to write code; it will also have to read existing code and make adjustments to it, based on request. This is the real challenge. "Hey, I have this 1MM line software, please change feature X to do A on top of B but keep C working". If the AI can do THAT, well THEN that's the end of programming by humans and then I'm gonna escape to a small island ASAP because the singularity will be soon after.
As with many things, AI will excel at the simple tasks that just require pattern matching and generation to work. Nothing more but also nothing less.
People get carried away by how great AIs are at spitting out complex sentences from a huge training set that look like a human would have written them. Winter v2 is coming.
Besides, there's no training set for the specific domain problem I'm currently solving. It is - by definition - not generic. The problem domain is not bubble sort or fizzbuzz.
The halting problem doesn't say anything about whether a program could prove things about another program, just that it can't do it for all programs.
Lots of programs obviously do or don't halt, and you can prove it in a couple lines ("there's an infinite loop here and every execution will enter it" or "all the loops are bounded by constants" etc). There's nothing stopping an AI from noticing things like that just as much as humans can. For more complex programs it obviously gets much harder, but it's very hard for humans to prove things about programs too!
Unless we are not talking about correctness.
And we are not talking about trivial cases either.
But I agree with your sentiment: to get AI to write the code you want, you will need an engineer to work with it.
I do expect that future AI code generation models will learn more about a specific code base over time, partially overcoming domain specific knowledge gaps.
Like, most developers aren’t writing proofs that a series converges to pi or building Apollo modules where everything has to work perfectly. I’m not gonna act like what we do is easy, because it’s not. However I’ll never forget the day I wrote a WebGL shader in GLSL with Copilot’s help, and then clicked over from a GLSL file to a JS file and Copilot immediately wrote the code to create a THREE.ShaderMaterial, import my shader (with the right file path) and filled out the uniform names and types perfectly on the first try. It took under 2 seconds, and at the datacenter maybe a fistful a watts, and it left me fucking speechless.
I’ll admit it’s pretty weird that these models can write working code in a multi-language codebase but can’t add two four digit numbers. They’re almost savant-like the way they fail at basic tasks while excelling at more complex ones. But given that they are in fact excelling at complex tasks, I’m worried about what could happen if, say, every enterprise dev shop laid off 60% of their headcount and kept the remaining 40% just to do code review. That would be very bad for everyone in the industry.
It's not practical to do that sort of thing alone, as there's a lot of tedious nonsense involved. Maybe it will be more practical with an AI assistant.
Whether current AIs are more or less somewhat close to the basics of how humans might be approaching code writing is more of a philosophical point, but a "translation room" view of the human mind might on occasion make me feel like I'm just working off patterns from my data sets. :-)
The halting problem applies to humans as well...
A method of producing statistical correlations is a product of a reasoning mind and could be thought of as a subset of reasoning (if by "subset" we assume "everything produced by a reasoning mind via an act of reasoning"), but in order to recognise this "subset" another reasoning mind should firstly internalise the notions of statistics that are external phenomena to an act of reasoning itself. And "the act of reasoning" isn't proven to be "just something that produces correlations", it's more than that and nobody knows what exactly it is. Otherwise AI would be solved long time ago.
Maybe it's cheating, but after all, this is the only way humans can do universal computation- we can't hold an infinite tape in our head either, and neither can a CPU, we have to give it sufficient scratch space to act as the tape.
Alternatively, perhaps there's a (very large) neural net that can prove things about Turing machines that aren't too large. It only has finite input and finite output (it's not Turing complete) but it can prove stuff about smallish Turing machines, providing the proofs aren't too long. That seems reasonable, because that's what humans do when we prove stuff about Turing machines! Perhaps neural nets could never actually do this, either they're fundamentally not capable or we never work out how to actually find one that does, but it seems possible?
You take symbol A and derive another symbol from it or combine multiple symbols to form new symbols.
You can approximate this through statistics but sometimes approximations aren't enough.
Have it recite the halting problem
In a new chat have it write a python program that tells you if a given string containing a python program will halt.
It's great at hacking, and completely incapable of engineering.
Like, I see a lot of people saying things like “It’s just pattern-matching, it still needs an engineer to guide it”. That’s true of this version of Github Copilot. If it improves at the same clip as some of these other large models, I suspect it will become much much less true in a short time horizon.
I would be unsurprised if GPT-4 or 5 suddenly becomes really good at math and logic problems. And long before that I expect AI to be hard to distinguish from a human in terms of day-to-day executive function stuff.
I'm old-fashioned about AI, over decade ago I was fascinated by Minimax heuristic in a game as a bot. It looked really smart I can't beat them in a deepest tree the machine can support (less pruning). My point is that the key is [UI -> "heuristic" + [business]logic -> product] is the way they should utilize these new neural network based AIs. Don't spit the product out directly because the product currently is looking like "heuristic", not a usable thing in business. Try putting logic in front and behind, especially put it far away from the final product so that the result will look like a legit magic.
We build endless higher level abstractions ontop of each other in programming, this is just another one.
I'm not bullish on AI actually understanding something in the near future and it'll rather continue to be something more akin to mimickery, albeit amazingly expressive and accurate.
I think this is rather going to become an amazing tool to help reduce repeating already solved problems. But humans would still be needed to plumb it together and adjust it to meet some final need.
If the AI can do the whole thing then whatever you're trying to create probably already exists and there's unlikely to be a need for it in my opinion.
https://arstechnica.com/gaming/2011/01/skynet-meets-the-swar...
> [Ben Weber] set about organizing a tournament for StarCraft AI agents to compete against each other, hoping to kick-start progress and raise interest.
> The announcement for the tournament was made in November of 2009, and the word soon went out on gaming websites and blogs: the 2010 Artificial Intelligence and Interactive Digital Entertainment (AIIDE) Conference, to be held in October 2010 at Stanford University, would host the first ever StarCraft AI competition.
[...]
> the only way to really test and improve the agent would be to play against skilled human players. Flush with pride that the agent could defeat the built-in AI, we played a game during the class against John Blitzer, a post-doc in Dan’s group who played ranked ladder matches on International Cyber Cup (iCCup).
> It was a disaster.
[...]
> Manually iterating through parameters and making adjustments would take far too long, however.
> Instead, we let the Overmind learn to fight on its own.
> In Norse mythology, Valhalla is a paradise where warriors’ souls engage in eternal battle. Using StarCraft’s map editor, we built Valhalla for the Overmind, where it could repeatedly and automatically run through different combat scenarios. By running repeated trials in Valhalla and varying the potential field strengths, the agent learned the best combination of parameters for each kind of engagement.
[...]
> Recruiting Oriol as our “coach” helped us apply the final touches. Oriol had played StarCraft at the pro level before retiring and turning to a life of science, and he joined the team as our coach, designated opponent, and in-house StarCraft expert.
> With a high-level human expert to test against and all of the algorithms in place, the agent progressed rapidly in the last few weeks, culminating in that first victory against Oriol mere days before the final submission.
[...]
https://www.theverge.com/2019/10/30/20939147/deepmind-google...
> Like OpenAI, DeepMind trains its AI agents against versions of themselves and at an accelerated pace, so that the agents can clock hundreds of years of play time in the span of a few months. That has allowed this type of software to stand on equal footing with some of the most talented human players of Go and, now, much more sophisticated games like Starcraft [2] and Dota [2].
Note that they are lucky there to have a controlled environment, meaning that experimentation is cheap, with clear goals - something that is not always the case in "real" life !
I dont fully agree with this. A lot of folks in systems land have mechanical sympathy and deeply think about memory, IO and processors. Things are mostly built upon underlying abstractions. With AI becoming mainstream, some of the abstractions will be pushed down and some might evolve further.
You think 50% of SWEs understand the physics of transistor design???
I can do that, at least for the basic layouts, but my own lack of (clear) understanding would be in the "middle" of the stack : starting up from logic gates, and bottom from scripting programming languages.
These requirements need to be precise, unambiguous, and complete.
The AI could help the requirements writer to “fill in the gaps,” but the main onus is still on the author of the requirements.
As mentioned, we don’t program in machine code, anymore. Maybe the result of an AI-assisted construction would be machine code, but it would take an AI to test, debug, and maintain it.
I know that every C-Suite denizen has been dreaming of getting rid of “annoying engineers,” for my entire career, but that won’t happen, as those requirements will look a lot like … code … and I guarantee that the C-Suiters will have zero patience for writing it.
We could spend years in a lopsided state where groups of investors fund an AI that operates on an investment thesis, delivers commands to humans who manage physical labor in areas that have been tough to automate (like the remaining Amazon warehouse jobs), and handles on its own all of the work that would normally be done by office employees.
There's no way AI is going to take the place of the customer, since the AI doesn't know what the customer needs, nor will it take the place of the analyst since even the smartest AI can't deal with a customer who is unable to clearly articulate their requirements. Hence the need for human analysts.
The AI might be able to help the programmer turn the specs into code (see Copilot) but it will always be hamstrung by not fully understanding (as a human would) the actual requirements.
Then I just need to procure the right assault weapons. Then I will be (at least my bunker will be) unstoppable. No need to hire mercenaries anymore.
Maybe one day it will get even better, that I can have my own attack units like robot dogs/personal tanks equipped with insane amounts of assault weapons and javelins. Then I can mount an attack against anything. A person, an organization, a city, a small government. A true one man army, with AI controlling everything.
But there is also a great danger of the Killer Joke, which results in instant death of anyone, who hears that joke. A malicious AI can re-invent the Killer Joke, and exterminate the humanity.
You know, I would love for all business app development to die in a wretched fire of scum and villainy. Not that it’s bad, but it’s probably some of the most mundane work that programmers could do. The people giving out busy work or bullshit jobs won’t be able to affect actual people. They’ll just have AI do it, which is great!
Traditional software will continue to be chosen when we want predictable, unbiased, mechanical execution of instructions. There are many areas where this is preferred, and I don't see that changing. Mechanical and later silicon calculation devices are invaluable for their speed, but the greatest benefit is that they are predictable and consistent: they do not make errors unless the design is in error.
AI, machine learning, and other training/learning-based technologies also have many useful and tantalizing applications. For applications such as those that enhance productivity, provide entertainment (e.g. art and music), or autonomously perform tasks where mistakes can be tolerated, these training/learning-based technologies will reap great things.
However, for many applications we don't want a complex device, whose behavior, while it can be ostensibly tested, cannot be completely understood and examined to be provably correct. Or, whose faulty action cannot be definitively reproduced and root-caused after a mishap. Or, whose 'black-box' can be infected or influenced by bad actors in a manner that is undetectable.
I don't ever want to see a radiation dosing machine that is clever, an industrial control process that is expected to be trained to infer its own decisions where injury or life is at stake, nor do I wish to argue with a machine to open my pod bay door.
Alternatively, perhaps legal precedent will just establish the degree to which machines are allowed to make mistakes, and if they make fewer than a human, we will just accept the cost/benefit of injury, loss of life, or evil as 'practical', and move on. 'Actuary Shrugged'?
The most ominous prospect is if humanity fails to evolve past war and conflict faster than this technology's destructive capability. Maybe Fermi will get his answer.
Context matters a ton and a lot of programming is understanding context and requirements and goals and needs and economics and those, while trainable, will suffer from the slowness and lack of richness of the I/O between the real world and the model (this interface is not improving nearly as fast as the models themselves).
[1] https://bartoszmilewski.com/2020/02/24/math-is-your-insuranc...
[2] https://eli.thegreenplace.net/2022/asimov-programming-and-th...
Fast forward to today, the new programmers I meet don't truly understand what the code does. If a problem pops up, the first action is googling it for hours. Very few people have this high level structured thinking ability to filter out the noise to distill a problem.
I believe Ai will accelerate this, where programmers will know even less about what goes on in the program and struggle more when stuff doesn't go as expected. Over the years, google became less useful as all the SEO spam took over. I noticed how people couldn't come up with solutions on their own, as a google search yielded no results.
Now I read these articles every other month about another tool, framework or article announcing the end of programming as we know it. Nothing ever happened and truly experienced developers became even more valuable.
If we cross this fine line between aiding and replacing developers, we set ourselves and the next generation up for a bad time in my opinion...
"We are no longer particularly in the business of writing software to perform specific tasks. We now teach the software how to learn, and in the primary bonding process it molds itself around the task to be performed. The feedback loop never really ends, so a tenth year polysentience can be a priceless jewel or a psychotic wreck, but it is the primary bonding process--the childhood, if you will--that has the most far-reaching repercussions."
Bad'l Ron, Wakener, Morgan Polysoft
Accompanies the Digital Sentience technology
Bonus :https://www.youtube.com/watch?v=-aUcvswVJ58&t=3852s
"'Abort, Retry, Fail?' was the phrase some wormdog scrawled next to the door of the Edit Universe project room. And when the new dataspinners started working, fabricating their worlds on the huge organic comp systems, we'd remind them: if you see this message, always choose 'Retry.'"
Bad'l Ron, Wakener, Morgan Polysoft
Accompanies the Matter Editation technologyAnd then it would also be handy for it to design a programming language for both readability and writability that leveraged those libraries.
Throwback to Keynes forecasting that we'd all be working 15 hour work weeks.
The point I'm trying to make is that for certain technical concepts (mathematics, programming, etc.), it's not the technical detail that must be remembered. Sure, you forget exactly the syntactical details of implementation, but you do remember (assuming you understood it at the start):
1 - How do I traverse the tree? 2 - Is there any ordering required among the nodes so that my traversal is 'correct'? 3 - How do I maintain such an ordering and establish correctness after my new node is added? 4 - Can I make any improvements so that locating the position of the to-be-added node doesn't take forever?
You don't begin to touch on these higher level "computationally minded" ideas until you can understand the primitive action of adding a node. And once you've grasped it, you forget the primitive details, only the essential concept remains.
The author seems to suggest that the next generation doesn't need to make an attempt to understand the fundamentals. I argue that unless you're a genius, you need to first add a node to a binary tree before labeling yourself a computational thinker.
In the past "programmers" used to write raw assembly. Now the compilers do this, and programmers write source code for the compiler. In the future AI may write the source code, but we will still need to write higher-level specifications, which will likely be much more detailed than "develop a mail client app" or "develop an FPS".
Even today, business managers who have programmers to do all of the actual coding, need to write detailed specifications (and when they write bad specifications, they get bad products); those specifications are in a sense, "code". But even those specifications are not detailed enough: when refinement is fast and cheap (with AI doing the coding) you really want to be able to customize the UI, add various features, properly handle various edge cases, etc.
or
A committee is formed whose only purpose is to hit the AIs that have become too big with an oversized wrench and kick it back into a madmax style desert, so it can evolve in a different way.
Maybe the key is testing and coding can be seen as adversarial actions, and there is benefit in separating them. If the ai or code generator or whatever you want to call it writes the code and test, I'm even less likely to trust it.
You know, kind of like those who said you won't need to know what a relational database is if you use ORMs.
What is impossible to do then is something like convex optimization models for guidance of rockets. You need to be able to mathematically prove that the output of the program will always converge to the solution. You have to assume that the inputs to the algorithm could get scrambled by ionizing radiation on one cycle and that the next cycle it'll recover completely. You want hard mathematical proofs. You don't want a black box that has been trained a whole bunch and might have some sharp hidden edge condition that you'll never know exists until just the right input hits it.
I also see it failing a lot. There were lots of promises made about 4GL languages that didn't turn out to be true. They made for great demonstrations but as soon as you push past the boundaries a bit things get tough and you need an actual programmer to make things work.
The guys that came and demonstrated PowerBuilder[0] back in the nineties had an absolutely stellar demo, but the devs who were given PowerBuilder to work on at my workplace quickly got mired in details that brought them undone. I feel like the same thing will happen with anything non-trivial that AI generates. The re-work will be more effort than just building whatever it was from scratch.
You could hook the program up to a python interpreter and approximate that, but then you're still generating code, not doing the task directly. In order to have an AI that will do the task directly, we would need to train a different model, not just plug the one we have directly in to the task.
It won't replace a janitor because it has no hands but then again, if the model's only claim to replacing programmers is that it produces text, then the article is pretty weak.
I have experimented a little with coding and ChatGPT and it seems impressive at first looks.
I had a funny social interaction while on a hike this morning with three friends who are all semi-retired videographers. They seemed enthusiastic enough to hear about automation for software development but when I brought up my continued joy at automatic photo/video/added-music mixups in Apple Photos (‘memories’) they didn’t like that. One friend categorically stated that AIs will never be able to adequately edit video, etc. in postproduction. I think he is wrong but I let it slide. I travel a lot and have a ton of digital media assets and I so very much appreciate the automatically created mixups. I watch every mixup and flag about half to keep forever.
Especially if you combine the text/instruct model with the coding model and then give specific instructions, it is able to complete many simple coding tasks without me opening an editor.
Right now I am focused on something like Codepen but with English specifications only.
I believe that I should immediately start leveraging this type of tool in my other projects. Similar to the way you would use a calculator or Google Translate.
I believe over the next few years the models will continue to get better and also start to incorporate visual understanding and better reasoning. The point at which it I would consider it poor software engineering to write programs manually is rapidly approaching for many domains.
It will take longer in the fields where neural networks are banned, like government-related software.
Consider this toot as well https://phpc.social/@andrewfeeney/109466122845775778
tbh sounds like a lot of people, myself included
"If you want a career you should move away from programming as soon as possible, those days are over, we're already testing automatic programming systems, it's a matter of months before we use them in production."
It was 23 years ago.
I think that the current state of AI will be sufficient to build excellent assistant tools, and I can see some productivity increase in some areas.
But for anything more advanced, let's say I am not worried.
With that being said, we've seen incredible increases in the ability of AI, even in the last 5 years. We're approaching human-level ability in NLP quite fast - and will surpass it soon. I (personally) don't believe that the past 23 years will be a great indicator for the next 23 in terms of AI. Some of the driving forces in the past 10/15 years are: hardware (we started using GPUs), Transformers, and massive - and I mean massive - increases in training data, fueling a nonlinear progression.
It is not just AI "explainability" but techniques for exploration and fact-finding. For eg: could an AI system prove or disprove P=NP.
It is a little premature to dismiss human endeavor while such large gaps in our epistemic knowledge exist.
AI progress is astounding and the results impressive, but they still seem to be 'fuzzy', almost like an instantiated dream generated by the input. Sure, as an assistant it, is powerful, but it still seems to be missing something to be able to provide a complete abstraction between people and code.
I imagine it’s not just about having API examples, you would want real codebases and use cases for training.
Presumably the same type of thing could work for code.
I would extrapolate from copilot, since software does not scale linearly. 100LOC << 1000LOC << 100000 LOC.
The AI would probably stop at the range of 100-500 LOC.
- We'll want human engineers reviewing and testing the code generated by the AI.
- We'll want human engineers building a better, higher quality training sets
- I suspect AI will run into the same issues that humans do when making changes to complex existing code bases.
- We'll want better languages and libraries for the AI to program in
- This may all be well and good for standard types of applications, but I wonder how well it will work for novel types of computing
- None of this will solve the problem of needing to clarify our ideas (https://xkcd.com/568/)