So yes, the codebase is growing fast and with it the tools to manage it.
So yes, the codebase is growing fast and with it the tools to manage it.
AI writing software will be a exponential explosion in software complexity.
AI would very well create its own programming language to be more efficient for the AI to code in that we have no hopes of understanding. Imagine that AI started to output a large SaaS app written today in Python in Assembly because for the AI the extra cognitive overhead of using Assembly doesn't exist. At first we might resist and tell it to use a language we understand, but then as time goes on we grow more comfortable that it does the "right" thing and down the road people are just generated raw Assembly using AI without really understanding what the code is doing, only looking to see if the code behaves the way they expect.
Imagine entire codebases spun up in seconds with so many lines of code, a single person would never have hopes of understanding everything, needing to rely on AI to summarize and explain the code for them. Now imagine that massive code base being iterated and worked on for a decade over the life of a company.
AI could bring Terabyte sized code bases in a decade or so.
I’d worry more about generated code from an AI that doesn’t fully understand the codebase. That would be as bad as letting the junior devs run loose.
What you are describing is an traceable transformation of code in which there are several intermediate layers that people can inspect and understand. They can inspect the exact rules for that transformation. The process is repeatable and verifiable.
What I am describing is a black-box stochastic generation of low level code in which there is no higher level representation anymore. AI generating Assembly not by a set of rules, but using statistics. There will be no individual layers to unwind or inspect, because for AI it doesn't need them. Our separation of concerns was built for our human brains and limiting complexity of projects to our understanding.
Also, you can "get good" at reading assembly, but that doesn't matter if the AI can output a custom OS from scratch and a custom VM to execute the program it wrote to solve your use case. It will be so impossibly complex that it would be the equivalent studying protein folding.
Instead people will just trust the AI.
It also won't help you if the code base the AI produces for a SaaS app is a million lines of assembly.
Instead of having different layers of OS, compiler, high level language, an AI will just be able to produce one layer. because after decades of trusting the AI to write our code, why wouldn't it?
The current gen of AI outputting code that in human-centric programming languages will be a blip in the history of AI. As it advances, it can just skip that step.
Its will be orders of magnitude more complex and opaque than anything we have today.
No, this is not merely anti-AI sentiment or a complaint about how the AIs of today aren't really as good as they are being presented. For the purposes of this argument I'm assuming AIs will work exactly as advertised some point in the relatively near future and you really can just AI-prompt your way through a non-trivial problem.
The reason why we have way too much code isn't that code is too hard, it is that it is too easy. Modern cars are complicated, but their complication is restricted by the limitations of being in three-dimensional space, being something that has to be physically manufactured where each individual part is itself something that a person may have spent a person-month or three on, the failure rates of physical parts and how those rates multiply if you keep piling on the complexity, and so on and so forth. The result is that a car has tens of thousands of parts, but it can't have trillions of parts. It doesn't matter how many engineers you throw at the problem, nothing like a modern car could be that complicated. It would simply be guaranteed that too many parts would fail. (To be that complicated it would have to be a different kind of thing, like some sort of biological organism, which are themselves rather complicated. But we humans don't make those.)
We have monstrous code bases because they are easy to add to. We can add a new module and test that module in a degree of isolation no physical engineer could dream of. We have analysis tools with no physical equivalent. We don't even design the raw parts, which would be the compiled assembly, we have piles and piles of tools to take things much simpler to understand ("source code") and turn them into the final parts rather than working with the final parts directly. It is easy to be cynical about things like browser quality and I see a lot of fashionable poses struck about this on HN, but the reality is there is nothing like Chrome in the physical world. It is the software world whose artifacts are approaching biological levels of complexity. This isn't because we software engineers are special as people. Engineering tasks by their nature tend to rise to the maximum level of complexity the engineers can deal with, and we all have basically the same intellectual prowess. But we software engineers have unparalleled tools to deal with our complexity, so we get orders of magnitude farther down the path of creating it, and then having to deal with it.
Thus, when you understand that our code bases are complex because creating more code is easy, not because it is hard, you can then easily see that any paradigm shift that makes it even easier to create code means that we will have even more of it, not less. If I can type "Dear GPT-12, I don't like my web browser. Please write a standards-compliant web browser except the tabs stack in 3D in my VR environment instead.", and a few hours later the world actually has another 150 millions lines of code and a new browser (I'm assuming more standards the browser will have to implement in the meantime) to deal with, that's not going to contribute to simplifying the code world.
It will not be a situation where you "can" use AI tools. It will become a situation where you will have no ability to do anything else, and the code will naturally grow to the limits of what AI can deal with, until it then exceeds even those tools as much as our current code bases are exceeding ours.
Of course, then the question is, what will we get for that complexity? I don't know. But you'd better hope it doesn't break because your only option to fix it will basically be "Oh AI genie, please fix this code." And I do think I can guarantee a lot of quirkiness in the result, where things just randomly don't work and it becomes even beyond the AI's abilities to understand why.
This is an interesting analogy to make because that complexity, much like the complexity in large software systems, is emergent from many agents interacting in small ways without top down coordination. So perhaps the right low level behaviors from software engineers would allow for monstrous complexity whilst still being effective.