Gen AI will increase demand for software engineers
roarepally.com
roarepally.com
There's still massive hype and group think about LLMs and what people project they will soon be able to do (or even what they can currently do - Dwarkesh's insider friends surprised it fails on Chollet's simple ARC tests!). A large part of this seems to be based on how new the tech is, and how capable it is, but people seem to be forgetting that the capability is based on brute force training time gradient descent. There is no runtime gradient descent or any alternative runtime learning method ... just limited in-context learning by example.
If you're thinking over the next 3-5 years, then AI is a novelty in most cases, a useful tool in some cases, and something truly disruptive in a fairly narrow set of cases.
If you expand the time frame to 10-20 years, then AI is poised to become something that truly upsets most of the economy. Obviously, nothing can be said for certain, there could be black swan events, etc. But on the current trajectory of cheaper, more powerful compute and better understanding of the engineering around getting useful results out of AI, you'd be foolish to think any white collar profession won't be deeply affected by the technological shift that's coming.
This feels like a lot of "20 years away" type predictions. There needs to be some sort of step change in the technology, and perhaps more than one or on more than one front to make this possible but its far enough away that 20 years feels like a safe bet.
When people say this about LLM based AIs, I wonder how they imagine text prediction making those step changes. Is your thinking that text prediction is generalization enough that doing more of it better, faster, and on scaled hardware is sufficient to hit a tipping point?
LLMs dramatically reduce the price of bullshit. It follows that they'll likely reduce the value of bullshit too, but we don't know to what extent or how that'll affect the organization of the economy until entrepreneurs generate vast bullshit factories and counterparties start adjusting their behavior to adapt to it. A huge swath of the economy runs on bullshit, from media to PR to sales to venture capital to politics. How these industries are structured will likely change when writing vast quantities of truthy-enough-to-be-convincing text becomes insanely cheap.
Definitions vary, but all definitions I have heard seem to boil down to one thing: AI is little more than a set of fancy programming languages. Which means you still always need a human to be the programmer. Nothing really can change for the white collar so long as you still have to provide instruction. That's what white collar people do.
"AGI" draws a distinction – where the machines start to be able to act as completely independent agents without needing a human to give instruction. Is this, perhaps, what you are envisioning in 10-20 years?
AI today is not even close to replicating the brain’s ability to perform multiple tasks together without suddenly hallucinating a totally different task at multiple points in the project. LLMs are good at regurgitating text, I would hope so after they just memorized all available text data. I can feel the chill of the next AI winter right around the corner, personally.
This is an interesting framing that I think exposes a key difference between people who view AI dominance as inevitable and those who don't. You view the progress of AI slowing down as a black swan event, something that's extremely unlikely. I, on the other hand, view it as inevitable.
Technology has never progressed along a single axis linearly or exponentially for very long. Technology progresses in bursts—there's a major breakthrough that leads to a huge amount of interest and progress which leads to rapid growth in a particular area until it saturates. After a few years we start to see diminishing returns as the area opened by the breakthrough becomes fully explored, and eventually progress becomes very slow indeed until the next major breakthrough.
Given this general pattern of progress in technology, it would be very foolish indeed to predict the next 20 years' trajectory of AI less than 2 years after ChatGPT was released.
The data is tapped out for the most part. They already probably acquired 80% of the available text training data.
Compute is now heavily cost constrained. This is obvious because even GPTs responses have gotten worse over the last 6 months, especially with image generation tasks. Likely because they’re throttling compute.
Model structure: OpenAI has admitted that they cannot understand why chatGPT has generated a given output. If you can’t explain your model you can’t improve it either. And more compute won’t help you that much here if I had to guess.
It's perhaps relevant to note that the reason estimates have always been off seems to be that people have always thought that the technology of the day was sufficient, and that it was just a matter of applying it. So far we've had procedural programming (SHRLDU), expert systems (CYC), general-purpose problem solvers (SOAR), PDP/connectionist approaches, modern neural nets, LLMs...
I do think we're getting closer, and that neural nets will get us there (maybe not with gradient descent as the learning method though). There have been lots of valuable lessons and insights learned from various neural net architecture, not least of which is the confirmation by LLMs that prediction seems to be the basis (or at least one of major pillars) of intelligence.
And yet, I can't help but feel that there is still a long way to go ...
I'm guessing that solving run-time learning may well require a different approach, and it's not clear that reasoning (in general form - ability to dynamically synthesize a problem-specific solution) can be just added to LLMs either (e.g. by adding tree search). There are also other missing components such as working memory that seem simpler to solve.
Coming up with brand new architectures and learning approaches is likely to take time. There have been attempts to find alternatives to gradient descent, but none very successful despite a lot of effort.
Perhaps it's just a reflection of people chasing the low hanging fruit (and as Chollet says, LLMs "sucking all the oxygen out of the room"), but architectural advance post-LLM has been minimal. In 7 years we've basically just gone from transformer paper to big pre-trained transformers.
Even when workable architectural approaches to run-time learning and reasoning have been developed, they will also need to be scaled up (another 7 years?), and will also be competing with LLMs for mindshare and dev. resources as long as scaling LLMs continues to be seen as profitable.
The timescale for coming up with new architectures and approaches is hard to predict. AGI prediction timeframes have always been wrong, and the transformer was really one of history's accidental discoveries. Who'd have guessed that a better seq-to-seq model would create such capabilities!
If I had to guess, I'd say human-level AGI (human-level in terms of both capability and generality) is still 15-20 years away at least. 7 years to go from small transformers to big transformers doesn't make me optimistic that architectural innovation is going to happen very quickly, and anyways this is an unpredictable research problem, not an engineering one.
For example, yesterday I asked GPT-4o to write multiple alternate endings to the short story "The Last Equation". They weren't dramatically compelling, but they were logical and functional.
How is that not problem solving? And so help me, before anyone tells me it's just stringing together the next most likely tokens - I don't care. Clearly that is at least a primitive form of intelligence. Actually it's not even apparent to me that that isn't exactly what human intelligence is doing...
So, intelligence exists on a spectrum - some things are easier to predict given a set of learnt facts and methods than others. The easiest things to predict (the most basic form of intelligence) is "next time will be the same as last time", which is basically memorization and pattern matching, which is mostly what LLMs are able to do thanks to brute-force pattern/rule extraction via gradient descent.
Going beyond "next time will be the same as last time" is where reasoning comes in - where you have the tools (experience) to solve a problem, but it requires a problem-specific decomposition into sub-problems and trial-and-error planning/testing to apply learnt techniques to make progress on the problem...
Certainly a lot of human behavior (applied intelligence) is of the shallow "system 1" pattern matching variety, but I think this is over stated. Not only is "system 2" problem-solving needed for on-the-job training, but I think we're using it all the time when we're doing anything more than reacting to the current situation in mindless fashion.
So, sure, LLMs have limited intelligence, but it's only "system 1" shallow intelligence, gestalt pattern recognition, based on training-time gradient descent learning. What they are missing is run-time "system 2" problem-solving.
Anyway, I actually am more towards this is a hype cycle that will pass, it's just the article itself didn't really drive home the argument in my opinion.
You nailed it. This kind of thing should be saved for a personal blog and not printed in a scientific journal.
I would argue that generative AI in the hands of an already skilled developer is equivalent to the mill scenario. One person's expertise leveraged by way of this tool can allow for output that is orders of magnitude greater. The mere inertia of the "machine" encourages more productivity than otherwise. It is almost as if it is pulling you forward when you use it properly. Knowing when not to use it is also very important.
The concerns over hallucinatory output do not seem as relevant to me in this scenario of highly-skilled milling machine operator. You expose yourself to something for 10k+ hours and you will immediately recognize when it does something funny. Novices will cause trouble in any domain regardless of quirks in their tools. I don't think lack of experience should be conflated with hallucinations when evaluating the impact these tools might have at large.
They'll fetch you all the latest libs, best practices, code samples, etc.
But they have no idea how to apply it to the business case at hand, how to secure the code, what real world edge cases to protect against.
And half the time their code is suboptimal or doesn't quite compile ... and wait till they try complex async ...
It's a force-multiplier for good/experienced engineers, but it's going to devastate the market for white-collar sweatshops.
And for testing to find those inconsistencies.
All the solutions I've heard up until now is "we just need more data" which isn't particularly sustainable or something a SWE would be in charge of doing
AI is exponentially improving and there is no reason to believe that this won’t continue.
Not for the last two years at least. Over the last year it's been obvious that there are diminishing returns already...
We increase the computational power by a factor of ten, and see a factor of 1.5 improvement.
The theoretical research may have improved exponentially. Implementations available for use seem to be on the decline – at best, stagnant. I find less and less utility as time goes on. Things LLMs shined at a couple of years ago now produce garbage. And for creative work, the output is much too formulaic, which was all well and good initially while still novel, but one has to keep pushing new boundaries and the current crop of tools really seems to struggle with that.
I would also disagree with the idea that LLM tech is "exponentially improving" now. There was the initial release of ChatGPT which was an enormous step forward, but since then its been small iterations on the same fundamental technology. We are already seeing significantly diminishing returns in our ability to improve these models. Most available training data has already has already been fed in, 10xing the compute results in only marginal improvements in performance, etc.
And with that experience I am advising my grandkids to look at trades for their future careers. Because there is zero doubt in my mind that nearly every job that requires a college education today is going to be gone in 20 years.
I am not one to talk in absolutes but the basic circumstances of current LLMs have hallucination at their core and can likely not be guaranteed fact based. It's just an expression of likelihood where you rely on your training data to contain truths.
This is why OpenAI gives more weight to Wikipedia than Reddit comments - the model contains no reasoning mechanism.
It ought to be treated more as a fallible human rather than a superintelligent lookup resource based on fact checked data.
AI a threat to humanity, my ass.
Perhaps that is what he is trying to get at? I am sure my communication skills are poor, but I found utility in using these tools a year or so ago. Whatever gibberish I was able to give it often produced a good result. These days I can't seem to get anything usable out of them. It does seem, like the parent suggests, that they have gotten worse.
It may very well be that said tools are no worse, if even better, where communication ability is stronger. However, if we have chosen to optimize these systems for those who are great at communicating, at the cost of those who are not, that doesn't help with the topic at hand. The non-software engineers – implying Average Joe – are almost certainly not going to be great technical writers.
That definitely stops some developers from considering it, others are happy to use just a standard editor without intellisense, etc.
Things that formerly were prohibitively expensive are now cheap enough to be doable. Which increases the market for software engineers that can get things done with the help of AI and other tools. These won't be the type of software engineers that specialize in things that should be automated but the type of engineers that can build a lot of stuff that formerly would have required them to delegate a lot of work to other engineers.
This is how the software developer community has historically grown actually. It used to be that software engineers were faffing about with punch cards, assembly code, etc. Working months to produce a few kb worth of software. No libraries or anything. It all ran on bare metal. These days what engineers build is the tip of the iceberg because it all runs on frameworks, libraries, operating systems, etc. that do all of the heavy lifting. Which makes them way more productive than their colleagues from half a century ago.
Thankfully, AI is still not there yet to make "professional" (as most people would define it) apps—unless maybe when people start feeding my book contents into GenAI, ha! (I guess it's a compliment?)
Also, AI is like DB. Sure, you can use DBeaver or something to access it "raw", but 90% of us programmers are still hired to write a CRUD wrapper. Same thing with AI. It's more useful to have an app supercharged with AI wherever it makes sense.
That’s like being against IDEs or git.
It’s like writers who were brought up on typewriters being against word processors.
The way to think about the impact on employment is comparative advantage. https://www.investopedia.com/terms/c/comparativeadvantage.as... Even if AIs are better than humans at every programming task, there is still a limited amount of compute in the world and unlimited amount of potential work to do, so there will always be tasks available for humans to work on. If the cost of writing code becomes cheaper, there will be tasks to automate that aren't worth the effort to automate now, that will _become_ worth the effort to automate in the future.
Some people are against a technology which isn't mature enough. I don't know about word processors, but take digital photography for example. Some people stuck with film photography even after digital became cool because it was superior in some regards (such as dynamic range). Once digital improved in every way over film, those people moved too.
Another reason is job security. Why would someone support something that will leave him homeless and destitute just because the software will be superior?
Personally, I think that software engineering jobs will go up while AI assisted programming is perfected (10-20 years), and afterwards there will be a huge decline (10x shrinkage) once the tools won't need people anymore.
If not, your competition will.
With some additional architectural changes we might start seeing some sparks of real human stupidity.
Also the latency is much lower. (Also I'm sort of mocking LLM's here)
The exponential graph people are posting is just to pump the stock price.
The set of benchmarks used to measure LLMs do not include things that require significant run-time learning or reasoning (e.g. Chollet's ARC test, although frankly that's a pretty simple test).
This exponential progress will flatten once LLMs can no longer improve along that LLM-benchmark axis (maybe they score 100% on all LLM benchmarks), but this "100% achieved on all tests" won't mean "mission accomplished - AGI achieved!", it'll just mean it's time for a new set of more challenging benchmarks, one step closer to what humans are capable of.