So the comparison is not only "built with and without LLM" but "would you even build this if you didn't have the LLM?". The gap in productivity in this case is much more wide.
So the comparison is not only "built with and without LLM" but "would you even build this if you didn't have the LLM?". The gap in productivity in this case is much more wide.
This can be a negative multiplier: code I thought I wanted that gets immediately abandoned is a net-negative if no one else wants it (lets face it, this is the safest default posture for software of unknown providence).
In isolation, instant-abandonware takes up hdd space, burns dependabot's CPU-cycles, and wastes human attention when appearing in search results. In aggregate, it floods the zone with a deluge of forks with imperceptible differences between them, based on nit-picks, legitimate stand-out products will have a much harder time going forward.
Some things are already very useful as just a one shot. I just made a quick app to help me pack for a trip, it updated forecasts every day, let me know when rain entered the forecast at one of my stops and gave me a checklist that helped me quell my travel anxiety.
The greatest thing that LLMs have done is allow many to achieve things that they couldn't have before. I'm incredibly disinterested in "I can do the same thing I was already doing x% faster"
it's truly fascinating how many positive descriptions of AI gesture at emotional management. I think that's the killer feature of this technology -- it makes people feel good, capable, reassured -- without the risk and vulnerability of interacting with another human.
I use Google Weather for my forecasts btw, no need to vibecode an app for that
I can keep track of my expenses on a napkin but i'd much rather use a spreadsheet or dedicated app especially when that app is effectively free.
I think you misunderstood my point: by "legitimate stand-out products", I meant exactly that, with no connotation of commercialization. Maybe you can agree that having a high signal-to-noise ratio for (open source) projects is a desirable goal?
> Some things are already very useful as just a one shot.
I agree. I too have made or forked about a dozen apps and tools in the past few months. It would be dishonest not to consider the flipside, that this software is overfitted to the needs of a single person. Further, this hyper-bespoke software typically feature-complete within moments of the final prompt, and I have,on occasion, completely forgot about the tool/app I spent a weekend created, it clearly wasn't worth the effort I put in.
> The greatest thing that LLMs have done is allow many to achieve things that they couldn't have before.
Let's not pretend there isn't a cost to this.
The word "product" implies commercial.
I disagree; but see where you're coming from. I'm chuckling at the irony of my word-choice: I initially had used "project" but nixed it because of its frequent association with Open Source. I instead opted for "product" as a broader term. For the sake of clarity, my original comment is referring to commercial and non-commercial software projects/products.
Also it teaches you that what you think you need and want is not what you need and want. This is why you are not using it
It’s not obviously true. A higher number of attempts, a larger talent pool, typically doesn’t change the average much (or it might even make the average go down), but tends to produce higher peak outcomes.
We see this everywhere (science, startups, sports, chess, etc).
If you want the best spreadsheet, game, or whatever app you want, you’re only interested in the few highest peaks.
So you do actually get better signal to noise with a larger wasteland of discarded attempts. The higher peaks make it easier to filter out the noise.
The goal you’re intrinsically motivated by seems different than this. That seems to be the whole disagreement.
obviously not ? negative result is still a result, just like in science. It adds new information ("approach X does not work" / "is useless") which is the only thing that matters
Unless you only run code that's protected by God. That's probably not a bad policy if you can verify it.
edit: Oh right, there's an OS for that https://en.wikipedia.org/wiki/TempleOS
If I were running around saying "providence" instead of "provenance" I would want someone to tell me.
Here's an example: my office has a few cars for employees to use as rentals. The number is small enough that it would never be worth any serious software dev to build a tool to manage, but large enough that its a moderate amount of work for someone to manage the requests/getting supervisor approvals/schedule changes due to breakdowns.
AI one-shot that guy a tool. Its now dead simple, he's got a calendar, automatic emails going to people's supervisors with click-here-to-approve links, rescheduling options, fleet management. It doesn't even look bad.
Who cares if its using some un-backed-up sqlite database in the backend, has some placeholder tab for a feature he changed his mind about, or violates the DRY principles a bunch or uses some inferior authentication mechanism. Its an in house tool, isn't mission critical, and it makes his life significantly easier.
Basically everyone is now a few prompts away from their own bespoke tools, and only they will be able to judge the benefit thereof.
Edit: to tie this more directly to the article, I would argue that this is an example of "infinity-x" coding, because the user was in fact not capable of coding a solution on their own without AI.
In my business, I haven't found much area to use code. It's a pub, and we've long been low-tech. Cash register, no POS. I have a little code surrounding my own processes, but mostly it's manual. Hand-entering numbers in my spreadsheet, etc.
But what's interesting to me is that now I can probably program an esp32, or create a small mobile app for a mounted android tablet. I was a web dev in the past, and programming hardware was outside my skillset without dedicating some serious time to learning. Mobile I just always avoided--mostly the same reason.
Anyway, I've got some CYDs on my desk, we'll see what I can make with 'em. I want a kitchen ticketing system instead of the old hand-written ticket stubs, for starters.
However, without an agent running its own experiments on a cloud GPU, would I realistically have invested my limited work hours and tried evaluating 10 different models, each with 10 different tuned parameters, to solve my specific use case?
Or would I have tried 1-2 models and spent my time trying to optimize those models?
I think there is some merit to the spray and pray approach when one is in the exploration phase of the solution space.
Also, on more than one occasion now, I have had fable halve the inference latency of a model simply because the original implementation from an academic included unnecessary GPU-to-CPU-to-GPU transfers or similarly inefficient operations. Those optimizations came at essentially 0 time cost to me and I can verify that the outputs are byte-identical. Pretty sweet!
E.g. software that generates these models that I can print
https://wiki.roshangeorge.dev/w/Blog/2026-06-30/Modeling_a_W...
https://wiki.roshangeorge.dev/w/Blog/2025-12-01/Grounding_Yo...
Or blog post authoring software
https://wiki.roshangeorge.dev/w/Blog/2026-04-25/The_rise_of_...
There were so many things that no one will ever study and won’t give humanity any benefit but I use everyday to make my life better. That’s enough. The value far exceeds $200/mo. I’m getting it for cheap and now that I have my GPUs and my models they can’t even take it from me in the future if they wanted, haha!
LLMs allow for human flourishing on a massive scale. One of the best inventions to occur in my life. Up there with the Internet/Web. Truly a marvelous time.
Agreed. I haven’t been this excited by computers since I got broadband DSL in 1998.
This has led to 3 parallel pieces of adjacent work that each speed up our build by quite a drastic margin. When combined, this is a massive improvement. None of this would have happened in the old days, as the research itself takes a long time to babysit and a lot of options to check.
So I very much agree - the activation energy can be a lot lower on some kinds of tasks, and some of those get big returns for small inputs. It's not all like that, but part of the game is identifying when you can spot those high return efforts.
dev A knows exactly what the program should do and how to verify the AI output
dev B thinks they know what they are doing but are actually misguided by bad psycophantic AI output they have incorrectly verified.
both work on product C
This is basically replicating the plight of the solo open source dev, writ large. Individual programmers have long built the thing they've cared about on their own time (essentially "for free" because, despite kindergarten economics theory, a programmer cannot usually monetize a marginal hour). And it usually goes that the project never gets adopted anywhere. It might acrue more features and total man-hour effort than most of what FAANG does in open source to drown out the solo devs. But the market will decide that "no organizational buy-in" is a signal the project doesn't matter. Other developers will decide, "if he could do it, so could I" and also not adopt.
Same exact thing is happening and will continue to happen with all these generated "but we wouldn't have done it otherwise" projects. It's just very, very unlikely to go anywhere.
That which took very little effort to create will receive very little effort to promote.
Its like making a jig in woodworking. The measure of the jig's success is not whether it gets re-used or widespread adoption, its whether it made it easier to achieve some actual objective. Because the jig is a means to some other end.
Lots of these "we wouldn't have done it otherwise" applications are means, not ends.
- Vibe code a bunch of small projects that we couldn't justify ROI before
- ???
- ProfitIt's a braindead simple program that mostly hooks together pre-existing functionality, it just so happened that none of the widely available apps had the specific mix of features I wanted. I could probably have done it myself in a week if I took time off my non-coding day job to figure out Swift and AppKit. But I wasn't going to do that. I’m psyched. I hate web apps and now I can just write my own for all the little things I use every day.
Without AI they might have first spent more time validating the idea was worth it.
The thing with constraints is that you focus of the thing with high value first. So you focus on the most promising ideas first or choose experiments that can get rid of most ideas. Instead of trying to validate each ideas and generate what is most likely noise to the decision process.
Like if I ever hire an assistant, I want like one to three options that are closely aligned to my needs, not a bible size report on 42 choices.
Seems optimistic
I'm doing analysis on stuff that we previously simply couldn't do in my company, it would take way too much time or effort, and we didn't have the manpower.