Assuming that AI is helping developers to write more code, it could mean:
* there are fewer developers
* developers are working less
* the efficiency gains are resulting in more products being created rather than existing products being improved
* AI isn't widely enough adopted or used to make enough of a difference
* the benefits are too recent to be measured
If you accept that premise the only conclusion left is that it has made developer's work much less but their bosses haven't noticed yet.
But you would still expect that some people are working for themselves and continue to work the same amount of hours so should be producing 10 apps a year instead of 1. Are there any examples of that?
Actual productivity takes time, and actual products take work. Talk is cheap, show me the code.
India have had massive layoffs in tech sector.
Europe and US will follow suit once companies layoff dead weight in outsourcing locations first.
I've found that for all but the smallest and tightest of teams, there definitely is a correlation... an inverse correlation.
Often the difference is between finishing within an order of magnitude of the predicted time versus not finishing at all. But the highest quality product is almost always finished first.
There are exceptions, and AFAIK, those are all for very small projects. In my experience, high-quality takes about a week to be paid back. So if you have something smaller that you'll throw away after a use, you may want to cut corners.
Another aspect of quality is "polish". A team that can get a UI in front of QA twice in a development cycle instead of only once will benefit from more fault-finding.
I think the more likely outcome would either be it still taking the same time to deliver, with extra fluff in the middle, or the time simply shrinks. One thing is still evaluated and shipped, but slight faster.
the efficiency gains are resulting in more
products being created rather than existing
products being improved
This has been perhaps the only constant in the history of this industry. As software tooling and hardware get better, we never really feel or see the gains because companies and individual developers are pressed to do more.If a tool makes my job 2x easier, then I'm simply expected to have 2x more output. Not 2x "better" output.
Software is easier than ever to write, but now the average AAA video game is 100+ GB and most popular software is browser-based.
This chart illustrates it better: https://x.com/feketegy/status/1809173358279311672/photo/1
I'm also not sure about "better"; I find Copilot is a good wingman for writing things like shell scripts, CMD.EXE scripts, powershell scripts and python scripts that do simple things. Even there I find it confuses forward slashes and backslashes sometimes so I often have to do a little debugging. Copilot can help me figure out how to use obscure (to me) features of PostgreSQL in JooQ. It will also argue with me and take factually wrong positions such as telling me that there is no zero-argument version of Optional.orElseThrows() which there is.
My big concern with all this stuff is it is pushing devs away from the uncomfortable part of coding. When you approach a new area of a codebase you just have to sit in discomfort and step through the code and experiment with it until your brain gains context. Sure you can tweak something on the surface to fix a bug, but to really gain full understanding of it can take days or even weeks.
Patience is one of a developers most important tools, whether its spending 30m uninterruptedly working to achieve "flow" state, or slowly stepping through a complex piece of code to see why it is failing. I worry AI is going to instill a rapid reward system that makes future devs goal just to rush to make it work. They won't have much interest in really gaining mastery over what they work on.
I take your point with respect to understanding large existing codebases. AIs don't yet seem to deal well with strategic large-context thinking. But things are changing at such a furious pace, so that may change sometime much sooner than either of us expect. But you are right. At the present moment in time, AIs are not particularly good at that.
But overall, I am much more optimistic than you are. I think AIs will save us time, so that we can spend more time on issues of overall structure that make "spending months to understand a codebase" a thing of the very distant past. I have certainly worked on codebases that require "months to understand". But I would (perhaps naively) like to think that's a symptom of a disease that codebases shouldn't have, and that modern codebases usually don't have anymore (but sometimes do).
Wow.
AI, like every technical advance in history, will be deemed "better" if fewer people make more money from it.
This has nothing to do with you, mister insignificant user...
In some ways it's like taking over a project written by someone else that's "80% done". You're locked into their design, get to analyze all that code 1 bug report at a time, get frustrated by obvious mistakes, confused (and mislead) when they relied on some clever side effect.
The quality of life in this maintainence mode depends enormously on the quality of the original coder. "Why did you choose this over that?" Is a common question I have for earlier devs. The AI answer is the least satisfying "it seemed like a good probability at the time".
IME writing the app from scratch to "done" is 10% of the lifecycle of the code. It's the other 90%, spanning over decades, where the quality (or lack thereof) reveals itself.
Personally I'm finding AI useful as a tool. Would I want to be the human fixing AI bugs? (From human bug reports which are pretty vague?) I'm not so sure about that.
https://greaterdanorequalto.com/ai-code-generation-as-an-age...
If these are the people who are coming into the industry, LLMs will not help them. There is no desire to learn or understand or dig even remotely beneath the surface.
Anyway, about this:
> It’s in the interest of mainstream employers to treat programmers as a fungible resource.
It's in their interest to make programmers fungible. It's self-delusional to think they succeeded.
I don't really mind whether you have some innate desire to learn something, or if you're doing because its your job and your job pays you. As long as in the end you suck it up and do it.
The difference (broad-brushing here) between someone who is a dev because that's their passion and someone who is a dev for the pay is the quality of their work. As our industry matured and gathered more "in it for the money" types, the overall product quality has been declining.
I think that's a shame. Perhaps inevitable, but a shame nonetheless.
Sure, we’re not at the “hook it up to prod codebase and ask it to make features from scratch” phase, but we didn’t have this 2 years ago either. But writing it off as completely useless? Nah.
I’m more of a process oriented person, who more or less cares about code quality. And as of now, it kinda sucks for it. However if your main goal is just the result, it’s delivering good stuff.
If your main goal is the _short-term_ result, sure, it's delivering. In the longer term though, I still believe it to be dangerous.
Even if we had a magic box that results in perfect code coming out every time for a given feature description, that doesn't mean the feature itself is good or well thought out.
That's why it's way easier for indie developers to deliver high-quality software—their incentives are directly aligned with the user.
The issue is that the people writing the UAT generally have no incentive to write good UAT. They want to ship. They don't want to get code blocked because it doesn't meet requirements. Amusingly and to get back to the core of the discussion, AI is generally pretty good at helping write good UAT.
Code AI tools today absolutely crush at creating proof of concept apps. You can test your idea and get market validation in days vs months.
They are getting better at medium/large codebases, but still have a ways to go before being super useful and it translating to a huge increase in productivity. Currently it's really good for helping with the menial tasks (creating docs, unit tests, understanding and onboarding) but not quite there yet when it comes to integrating gen ai code in large codebases, but it's only a matter of time.
I run DevRel at Sourcegraph and our AI coding assistant, Cody, is used by tons of individuals, small business, and large enterprises. I get to talk to a ton of customers and see how their adoption of AI is going. And it's certainly increasing and developers are finding a ton of value.
GitHub Copilot, ChatGPT and Phind are all a bit like this for me - they both lower the barrier of entry and save me a lot of time for trivial algorithms and boilerplate code, in addition to helping me find things better than search engines sometimes do, especially when given a look at the code that I'm working with.
It might not be an order of magnitude difference in my case, but things that wouldn't have happened with the higher barrier of entry are now happening and that's quite the difference in of itself! I'm cautiously optimistic about LLMs and other forms of "AI". If nothing else, so far we basically have a more versatile form of IntelliSense, even if it's not always going to output correct code.
I wonder if some day it'll be feasible to feed in the entirety of a larger codebase and reason about it better than people who only know a part of it could.
That's the real issue for me. I remember learning programming and I either had not so good intelligence (Codeblocks, IDLE, Netbeans) or none at all (notepad++,...). This forces me to either follow the book attentively (and hunting down errata) or read the manual and getting explanations from forums or friends. When you're a beginner, uou need a good source of truth, not something that can be subtly wrong.
Throw away apps are the easiest to make. No scaling, no bug fixing, no long term maintenance considerations, no consequences for poor architecture, no need to consider data models, you just write some shit. Bravo.
Recently I used Cody to rebuild the entire video processing pipeline and made it much more efficient and scalable, and I actually learned a ton about ffmpeg by pair programming with the AI. Now I'm building additional features into this app, mostly w/ just iterative prompting or chat-oriented programming to replace 3rd party services that are still in this pipeline and it's been a blast.
I've also used AI tools to really brush up on frameworks, like Laravel, that I haven't touched in a while, and it's been a great experience. Also started building a game w/ Godot and found AI super helpful there in walking me step by step. So for me it's been great.
Everything old is new again
If we are strictly speaking about the cutting-edge models like OpenAI's o1, their context is getting smaller, not larger.
Even in a large code-base, you don't have to have the full context of every single file for every single question, can usually get away with half a dozen files, it's figuring out which ones to provide to the LLM to get the best response.
I will agree that they're very useful for boilerplate tasks particularly around deployment (cloudformation, github actions, etc.)
Personally, I have seen a ton of change in the last 4-5 months, so if you haven't tried these tools recently, I encourage you to try them today and see what's possible.
The real explosion of great software will happen in 3-5 years. AI is huge for the beginning of projects. You know _what_ you want the app to do but you don't know _how_. That's where AI adds huge value. People are now starting new projects with AI help, and they are building foundations of codebases that will be much more maintainable and sustainable as development continues compared to the current suite of software products we interact with today.
I'm not following. How are these codebases going to be more maintainable and sustainable if the developers are committing code they don't even understand?
You can also generate tests more efficiently, meaning you can get better test coverage cheaper. This leads to better maintainability as well, as you know more quickly when you've broken things with a change.
Because it's not helping them code better. It might be faster, but the quality in my experience is worse. Then the user is trying to verify or troubleshoot code they didn't write.
The bigger issue is garbage in, garbage out at the requirements level. The business hardly ever documents their business system before turning it into a technical system. How can we create a system to meet requirements that nobody knows and didn't have a chance to really think about while writing the code?
I would go even further and say that that relationship between the two is weak, but also very peculiar: bad code can ruin a good product; but good code alone says very little, if anything at all, about the quality of a product.
It’s like mobile cameras, they reduced barrier of entry to photography, but you didn’t end up getting better photos necessarily. Instead you had more of them.
I like this idea, any new technology that empowers people simply increases volume, there's still an uncaptured element of strategy, design, and taste that doesn't improve with the tools themselves getting better.
You can be above-average in your favourite programming language, but suck at the system or library that will be required in your next project.
AI will help you get up to speed.
2. Productivity gains don't go directly to products getting better. Individual developers may choose to realize some gains by spending more time with their family. Of course the company will claw that back but it takes time. And when they do, some of the gains may instead be realized by higher profit margins rather than better products and it will take time for consumers try to claw that back using their market choices.
3. Companies have lots of moving parts and a speed they're used to going; it will take time to adjust if one part goes a little faster.
4. LLM-assistants help a lot with getting up to speed in a new field or making stuff from scratch, and a lot less for a skilled team who already knows all the product code and surrounding tools. So "products you use regularly" benefit the least.
unless you're being paid to clean up the mess
it's like outsourcing on steroids
It's the difference between Euclid and modern notation, with AI programming being like Euclidean notation and current programming languages being the modern notation:
"if a first magnitude and a third are equal multiples of a second and a fourth, and a fifth and a sixth are equal multiples of the second and fourth, then the first magnitude and fifth, being added together, and the third and sixth, being added together, will also be equal multiples of the second and the fourth, respectively."
versus
a(x + y) = ax + by
If AI programming can find a better way to express the problems we're trying to solve, then yes, it could work. It would become a matter of "how well the compiler works". The current proposals which use natural language as the notation is not better than what we have.
The vast majority of issues is missed edge cases between what the user wants and expects, and the design and function of the software
Higher productivity would in theory allow programmers more opportunities to address more issues
Programs don’t exist in a vacuum, users and other actors need to interact with them
Whether or not these LLMs result in increased productivity with the same or better quality is a more pertinent question
1. The products you use are not developed by people using LLMs.
2. The products you use may be using LLMs in development, but only recently so you'll see a delay before any improvement.
3. The products you use are using it, and maybe it's helping with quality, but not anywhere that users care about or notice.
4. The products you use are using it, and it's not helping with quality, just churning out more code.
For a simple example, consider a would be program that takes 100 tasks of 16 hours each to build a program with a quality of 75%. With AI, those tasks can take an average of 12 hours each, meaning the software can be delivered faster. Unless someone purposefully invests the saved time into improving the program, you'll end up with the same 75% quality program faster.
Now what if AI makes the code slightly worse, leading the quality to drop to 70%, but some of the savings are used to improve quality, bringing it back up to 75%? Same outcome of the product not being any better to the end user.
Even if the code is higher quality, how much of that 25% of missing quality is the result of bad code verses bad designs or a mismatch between what customer wants and what those designing the project think the customer wants? Even a perfect AI that solves all bugs won't improve that.
In short, programming better can mean many different things, some of which might translate to a better or worse product, but with no consistency.
Usually a mix of poor design choices and hostile design.
Neither of these directly correlate to the quality of the code, how quickly it was created, how many bugs it has, etc.
If everyone working in tech (or any sort of programming related project in general) was an expert level programmer with decades of experience, neither of these things would be noticeably better. They'd still create software that's miserable to use because of bad design, and we'd still have companies trying to scam the users by making basic functionality hard to use (see cookie notices, unsubscribe processes, etc).
So at least some software is getting better.
There are use cases where LLMs help with coding (and they are growing as the things get better), but even if they could do as well as an experienced engineer at doing a first draft (which, debatable in any setting and falls off sharply once it's not a highly mainstream setting), a first draft is almost never a high-quality artifact.
They can also be used to get a sort of minimum viable diff that represents a liability to the codebase and those who maintain and depend on it, to do this with very little effort and therefore impose the negative externalities on someone else. Anecdotally this seems to be a distressingly common use case. I'm more than a little concerned that software quality is about to take an abrupt turn for the worse in aggregate.
More broadly, if you're anything like most people I know, the products you're using are getting better all the time... at making money for the companies that build them. Consumer Internet profits and/or valuations are at something like an all time high. All that lag and jank and spam and shit? That's not easy code to write or simple infrastructure to operate. That's full-metal-jacket monetization at great effort and expense.
If I’m investigating a new field or trying out a new language or library, relative to my own experience, then it’s quite common for an AI code generator to use idioms or libraries I hadn’t yet come across. That alone sometimes saves me a useful amount of time doing research.
However, it’s almost all breadth and very little depth. The quality of the generated code is rarely better than something a junior-to-mid-level developer might have written. It needs to be reviewed and corrected with similar diligence.
Similarly, the quality of a generated review of existing code or of generated supporting assets like test cases or documentation is often superficial and error-prone. I rarely find it an overall win to use current AI-based tools for these things instead of existing tools that can’t do as much but are consistent and reliable at what they do do.
So I wouldn’t necessarily expect current AI tools to help me code better, only sometimes a bit faster, and that mostly in new areas I’m exploring rather than areas where I’m doing professional work that is going to get shipped in the near future.
Developers using AI will get mostly average solutions faster but exceptional ones will be obviously rare. And, crucially if the idea itself is average or bad there isn't much an elegant coding solution will do for the idea.
I think this ultimately is the divide between the hype and reality of how AI will impact products. If you just give a product manager the keys to do all the coding as no code "prompt engineer", more than likely will lead to further enshitification of features in products with unmaintainable code bases. At the current state, understanding algorithms and thinking computationally is a requirement to improve a code base.
The hopes of having a "build me a $1 billion app" prompt capability, or "improve my shitty app" are too long horizon and subjective requests to bypass the hardships of product ideation and iteration to have the LLM deliver on the requests. It's not magic, it's probability. Averages are the end goal here, not excellence.
If we arrive at a point where LLMs translate general prompts into idealistic versions that are more like version 100 of the idea while still capturing the user's intent, then we will see these improvements. Otherwise it's copy pasta on steroids, and done mindlessly, will mostly lead to enshitification rather than improvements.
Also, in my experience, AI is helping people who don't know how to code to write code they don't understand and can't support. In my experience, the people who are already making products aren't getting much benefits from AI (not yet anyway).
Does it make the end product better? Not really: I would have gotten there with a function written by me or some LLM. But like everything I've been asked to do my professional career, it allows me to do more with less. More dumb functions in less time.
Further, non-coders being able to become equivalent to a junior developer is a huge leap.
What active developers do with AI remains to be seen. It really could 20x the average developer, but it doesn't seem like a huge chunk of developers are really using AI in a way that it's the rage on the developer level broadly.
Maybe that's why Cursor going "viral" on youtube seems different when it was known to some, and not others.
You may as well ask why the advent of StackOverflow didn't massively increase app quality, it's the same target audience.
It doesn't how effective software development is if you're not doing anything to improve the discovery and planning process.
As far as I know, most of us do research with AI to get ideas and find pros and cons but we are still the ones mostly driving the logic with AI filling in the function level blocks.
I imagine for first pass prototypes, AI will greatly accelerate the process. But getting to the fine details and getting things done well will still take same amount of time.
AI-guided coding will help code up “good enough” implementations, which is great for research and testing ideas but not for production.
Of course this all depends on what you mean by "better", or for whom the product is better.
AI, like every technical advance in history, will be deemed "better" if fewer people make more money from it.
This has nothing to do with you, mister insignificant user...
For most part, it will be same quality or lower at a cheaper cost delivered faster.
User research, UX improvements, feature ideation and creation, etc, are all the same as they have always been. Getting the code out faster doesn't help if its in service of a bad feature.
And for much of that year there were a lot of questions around the ownership of AI generated code as well as information security. In fact, most of those questions are still not satisfactorily solved. So I'm not so surprised that we haven't seen massive results "yet"
I don't believe they offer creative solutions, just a faster way to refer. So it's still the responsibility of developers bring their creativity to the process.
Sources:
https://openrouter.ai/models/anthropic/claude-3.5-sonnet:bet... - claude-dev: 5 billion tokens this week
https://openrouter.ai/models/anthropic/claude-3.5-sonnet/app... - aider - 325m tokens this week
Also see aider's leaderboard for models: https://aider.chat/docs/leaderboards/
I know my hobby stuff is getting better. I generally dont make front ends for my projects, but genai has helped me build widgets for other front ends, and css/js frontends for a lot of my nonsense.
I'm sure there are hypotheses. But it's also likely that things are simply not improving.
Like, there's little reason to think that LLMs are helping people code better.
In other words, drugs are better than the internet, and we could all learn a thing or two from drugs, to make our products better!
Copilot mostly helps people code faster, or with less knowledge required. I'd expect output quality to go down, not up.
And it's like anything in capitalism. Companies could choose higher quality, but instead they do whatever gives the highest profit. which is usually adequate quality at low cost.
I don't think it is?
30+ years of experience here and I haven't seen any AI coding examples that makes me want to use AI for coding.
It might happen one day but so far nope.
Whether AI is used to write code is irrelevant to a product getting “better”. AI copilots can be used to bootstrap early stage concepts which might be unpolished, or can be used to add polish by writing bug fixes.
I neither subscribe to the mania around AI, nor do I think it will enshittify products. I believe it is just another tool that we can use.
No longer are developers bound by physical media, or have to force clients to troubleshoot, manually download updates on their website and install them.
... and as others said, LLMs impact is greater on junior developers, and at that, more on speed than quality. For experienced developers, the impact is greater on speed.
I have no data, only sense to make.
Besides, products have been getting worse for a while. Enshittification is a potent force, and even if AI was axiomatically helping people code faster and better, enshittification might still lead them to add in annoyances, privacy risks, et cetera.
First, let's address the tooling side. While the current crop of "code completion tools" built out of or around LLMs are quite capable in their own right, they're not exactly "free thinkers" like we can be. Rather, their output is limited by a combination of training data, the model itself, and - increasingly - the user's ability to put their ideas into a prompt that can generate the desired output consistently. So there's already a huge hurdle just on the tooling side to overcome before we can begin "improving", one tied just as much to the capabilities of the product as the capabilities of the end user. I would argue that this is the most immediate hurdle to cross if we want to see meaningful improvements to code as a whole.
In addition to that immediate hurdle, there's three more issues on the tooling front:
* The existing training data is largely bad, bloated, or insecure code samples (generally from publicly-available social media and repositories), because code security and efficiency are only relatively recent prerogatives of large development companies or outfits as they seek to dodge lawsuits (security) and increase margins (efficiency)
* LLMs aren't very good at teaching a user how to think better about a problem, only making them better at phrasing their prompt to get closer to a possible solution
* LLMs are stuck in a predictive framework that mandates an answer for the customer, as opposed to a human who is able to say "I don't know" and going off to learn more about that thing they're stuck on.
Ultimately, the tooling is helping novice or entry-level developers and hobbyists write better code, but only because the models were trained on code from more senior or professional developers that was also shared publicly. Senior developers and above may find utility in writing faster code with LLMs, but aren't nearly as likely to write better code as a result of the tooling, at least from my subjective reasoning.
Now let's switch to the business side of things, which I already touched on above. Businesses haven't been interested in secure or efficient code until very recently, as we began bumping up against the limits of physical hardware in x86-64 land and lawsuits for failures became more of an existential threat. This means a lot of the code from public samples fits the "done is better than good" mantra of modern business practices, rather than being an improvement to prior releases; even if a business has taken the time to create more secure or efficient code, they likely haven't shared it as it's a core part of their competitive advantage or product line. This will take years, maybe a decade before the LLM training sets have enough "superior" data to outscore the "inferior" training set data, during which time the status quo - barring a literal revolution in computing - is likely to remain.
Admittedly all of this is my subjective POV from infrastructure-world, and could be way off base; YMMV, buyer beware, caveat emptor, etc.
But, the rate of change in this area is breathtaking. I reasonably expect my AI to improve in the coming months, or even weeks. And I find it difficult to keep up with which AIs are best for generating code at any given moment. There may be AIs that are good at reviewing multi-million-line code bases for security flaws. But I am not currently using one at the present time.
What I do know: my AI coding partner this year is writing code that is more accurate and more stylish than any AIs were producing this time last year. The code that's being produced is often strategically brilliant -- elegant, concise, only very occasionally using hard-coded constants instead of including the correct headers, and almost completely absent of "hallucinations". And I'm using it to regularly generate code in three different languages (C++ for the app server, typescript for the web client, java for the Android client application).
I've frequently found myself adopting coding conventions that my AI has shown me. I particularly like
namespace fs = std::filesystem;
And the solution it came up with for writing a std:filebuf implementation stills leaves me speechless. I've done that a few times over my career, and the solution the AI uses is infinitely superior to anything I've ever written -- not something I've EVER seen, but clearly the horrible, never documented way the original authors of the iostream libraries MEANT people to do it, which provides substantial advantages over the way I've been doing it. And absolutely nowhere to be found in the first 30 page of google searches, or among the strangely variously broken and obsolete fragments of code on StackExchange.But my current AI often falls short when it comes to strategic thinking. Functional decomposition is often odd. I often have to refactor code that my AI generates -- sometimes by coaching it through refactoring, and sometimes doing it myself when I move the generated code into production code. But that may change next week. Who knows?
Have I used it for debugging existing code? A couple of times. I'm not currently seeing a huge productivity boost in this area.
Today, I coached it to write me a bash shell script to generate a graph of a key development metric using gnuplot. Bash: as a former Windows programmer, bash still terrifies me. Gnuplot: a documentation set that could be called unforgiveable if you were feeling particularly generous. Took me about 20 minutes. "Change the font of the title, please". "Take input from this program, which produces a column of ISO 8601 dates, and an integer value". "Rotate the date labels 90 degrees anti-clockwise" (It mistakenly rotated them clockwise. The only flaw in an otherwise fantastic performance). Etc. It took me about 20 minutes to do what would have taken me a couple of hours. I wouldn't have done it if I didn't have an AI at my disposal.
Context: Using Claude 3.5, 40 years of very senior Windows development experience, but only about 3 years of Linux development experience.
1. Look for products that don't have (c)opyright. Any product still using that or licenses is going to evolve too slow and go extinct.
2. Look for products built on revolutionary simpler stacks like PPS.
I thought this was going to be an essay and not just a Tweet, so I did record a long winded response, which I think contains a lot of relevant info: https://news.pub/?try=https://www.youtube.com/embed/KhDvFNef...