Assuming 10x on the speed of dev, Is the vscode repo a decent example? Recently they've been all in on AI augmented development so i'm thinking they'd be a reasonable subject?
How do you isolate out what counts as the "development" part of their delivery cycle (is that the dev inner loop, does that show up in frequency of commits then?) to measure it and see if it's running 10x?
https://github.com/microsoft/vscode/graphs/contributors?from...
But from the POV of say, a young startup company looking for PMF and navigating the ambiguities involved with trying to figure out what is the "right thing" to build that will appeal/delight/convince-people-to-pay --> being 10X faster at shipping 80% done projects, is actually incredibly, unfathomably valuable of a superpower. And it is also rationally the "right thing" to do, to make lots of cheap bets and fail/learn fast.
I find that many folks on my team (I am a manager/leader of small-to-mid size eng org), struggle with accepting the nuances of knowing the difference between different projects (where same team may need to do both kinds of work, all the time and in parallel):
- "Hey, the company needs, and you and I both agree, that this situation calls for building/renovating a skyscraper --> please design a fucking strong/safe/reliable skyscraper and don't take any shortcuts, this requires 'real' engineering"
- Vs, "Hey, the company isn't sure what it needs, and neither you nor I know any better either, so let's try a bunch of different shacks/sheds/treehouses/whatever, until we find something that has traction / makes us money (and it's okay if the shed collapses -- so long as the business knows this too, that it wasn't meant to be a load-bearing, skyscraper-esque thing anyways)"
I won't get into the rabbit hole of talking about dealing with bad business leaders, who want a skyscraper but expect to pay the price of a shack/shed. Let's assume that we are talking about the type of companies (maybe the minority) that are reasonable enough to know and acknowledge the difference. Then what is the game-theoretic/rational thing for them to do, and how does this 10X idea express itself? That's where my argument is coming from.
AI is not delivering 10x shareholder value, anywhere. Software developers have quite the level of hubris about how important they are to companies. Yes our work is very complex and takes a certain mindset to do it well. It takes a lot of other roles to have a successful business, many of those roles will use AI to help draft slide decks, emails, etc. and that's the limit for them.
Look at recent companies doing layoffs claiming its because of AI, like CloudFlare and Coinbase, do their reported financials paint the picture that they are crushing it with AI? No, its net losses into the $100's of millions.
A bit facetious, but I'd expect Nvidia and the like providing the "AI equipment" to have a 10× share value at least…
One of the latest things I made with Claude was a tool that allowed me to move a bunch of very low traffic Cloud Run services to a single VPS without losing any of the Cloud Run benefits such as easy Docker-based deployment and automatic certificate provisioning. I thought about making something like that for quite some time, and Claude finally made it possible, which makes me quite happy.
The fun thing here is that no other soul genuinely cares about it, or any other code I might publish. The code, especially AI generated, is so cheap that if anyone wants to repeat my steps to get rid of Cloud Run services, they will probably vibe-code their own tool instead of figuring out how to use mine, just like I did that instead of spending time on learning Dokku or similar solutions.
So, yes, 10x and more, but no one cares about the result, which makes the whole 10x measurement less useful.
But I'm with hansvm - I haven't actually seen anyone plausibly maintain 10x. 10x is different from getting people past their activation cost.
It's not a matter of preference
It's when they practically ignore the rabbit holes where it's suspect. I'm definitely seeing speed ups. I troubleshot a linux system yesterday with minimal effort using a local llm. It likely would have taken me a few hours to locate all the docs & testing procedures. the llm did it with only a few prompts. To ensure it did it correctly, I had to interrogate it a few times before letting it proceed.
Humans make really bad scientists, and it takes a lot of effort to properly catalog and provide statistics for these things.
There is an improvement, but I doubt any random dev can give a real estimate since before LLMs they couldnt really give you a real estimate anyway. I do know when I encounter a bug now, debugging is almost immediately possible.
I build things I never would have. My tooling is better and more robust than ever. I verify and test my work better than ever. I fix more bugs than I used to simply because no one needs to care if it fits into a cycle. I explore and solve more problems in more parts of the application, even if I don’t write code. I take better care of our infrastructure. Performance goes up, bugs go down, AWS resources scale back, costs go down. I’ve paid for my AI usage in scaled back resources several times over at this point.
It might not be 10x but it’s a significant multiple.
1. I would not have attempted this without AI assistance because it's a big project.
2. I have built a functional program that I am able to use for real work in a handful of weeks, working part time on this (like literally a few hours per day prompting Claude and Kimi).
3. Had I decided to do this without AI assistance it would have been months of work.
The Turing Test used to matter until it didn't (does anyone even talk about it? was there a big news conference when it was solved?). Likewise every time it becomes easier to ship software, the bar will be pushed higher by sceptics. Ultimately the gatekeeping is going to become meaningless as software becomes "too cheap to meter".
https://github.com/KeibiSoft/KeibiDrop
It took me 2 years ago around 2k hours to build a cross platform FUSE vault, without using AI assisted tools.
The pain was debugging through logs and system traces. And understanding how things work.
Now managed to ship this one much faster, as an after hours project. Started it in may 2025, and around end of November 2025 started using claude on it.
Just by dumping logs into claude, and explaining the attack vector for the problems, saved me the FML moments of grindings walls of syscalls on 3 platforms.
I would say much easier to progress, and ship with the same rigour, minimize my time, focus and brain power involvement such that I can put the energy somewhere else.
Trying to fix syntax errors in strong interpolation on a 5-minute-delay loop is hell.
So my agent just listens for green checks and no PR comments and loops until those conditions are met.
I disbelieve this works in anything other than a toy codebase (or an incredibly fine-grained microservice).
The 70% is amazing! But a 30% failure rate requires intense supervision.
Might tend to deviate and waste time, needs guiding once in a while, and to check what is it spewing out, point it in the correct direction.
If I had to output the code myself, would take around 8 hours of constant writing to get around 1k LoC of code. For FUSE level tricky stuff, I might need to spend 3 weeks for 10 LoC. Very easy to burnout and build pain.
Complete frontend + backend + database.
Yes, it is an internal app, but it works and everyone loves it.
Does that count as an example?
(Also I absolutely expect him to need help at some point, but so far it has taken his project from absolutely impossible to 3 weeks of work in between work, renovating his house and being a dad for the first time so I was very impressed.)
The danger is not however that only that people write their own tools for calculations and capacity planning etc.
The danger is people make useful stuff that is very fine as long it is just an internal tool, but then someone add credentials to other systems so it can access and maybe even update stuff and it gets exposed to third parties etc and all of a sudden we have a major data breach going on.
And if you just mean your friend taught himself programming on his own, well that is actually very cool, I did too back in the 90ies and so did many others here.
My point is that it is now possible to vibe code a full application from frontend to backend today and still not be able to understand a line of TypeScript or anything else.
Direct github link: https://github.com/open-noodle/gallery
Nothing wrong with forks though.
We're now in an era where LoC is easy and design is hard[1]. Starting with an existing project means using an existing design, where someone else has already made many/most of the difficult decisions.
10Xing code without caring about design/UX/DX is trivial. Literally anybody with a token budget can do it. But they probably won't ship a good project. Not with current frontier models.
[1]: design has always been hard. But now it's even more difficult because of code veloocity and because LLMs are happier to work with bad code than humans. It's never been easier to go deep into rabbit holes without noticing a single issue.
The main thing they dont realize is: 1. These are mostly superficial changes. 2. The only thing they 10xed is their ability to "start" on something. 3. They have not produced actual value. Their project/fork is just a version they think they prefer. But It is less maintainable, and less robust/useful for others due to its specificity.
My observations is that consistently these arguments are made by: inexperienced devs who simply dont understand what it takes to produce value in the real world.
LLMs CAN 10x you (in very specific areas like prototyping), IF you understand how to deliver this value, but that is the hard part. It has always been the hard part.
https://opennoodle.de/roadmap/
Look at what I built, these are not all simple designs.
And this is the project I care most about. I review every single line of code the AI outputs. I push back on everything it does. I reword every commit message. I make every effort to understand how everything works before committing it to master. I cook branches for days and I don't merge until I think it's perfect.
Still gave me an ~8x improvement.
The latest development: shaped objects, like Self and V8. I asked Claude about it and it just implemented it in like 10 minutes. Boom, instant ~20% speed improvement. I read the code and it turned out to be almost obnoxiously simple. It basically converts hash tables into arrays internally, and deoptimizes back to a hash table if anyone deletes keys from it. I'm still reeling from the sheer absurdity of it.
We decided to integrate our SaaS into Microsoft Business Central and NetSuite as plugins into those systems. BC has its own programming language, called AL, that has a lot of idiosyncrasies from any other language I've worked with. And NetSuite plugins are written in SuiteScript, which is a custom JS runtime with a ton of APIs to learn.
In the "before", it would've taken 5 developers a year or more to build those integrations. I did both by myself in well under a year. Thank you Claude.
I've always been a backend engineer, never front end. And almost every team I've been on has lacked any front end skills at all, so all our tools end up being a mash of scripts, maybe sometimes an API.
Now we are all front end engineers creating UIs for things we could never do before, and this starts API first development, so the CLI + UI are just calling APIs. Nothing new here, but this used to be what teams do, now a single person does it.
Now with AI, I can easily create a nice looking front-heavy web app.
See this for example: https://github.com/erwan/sovereign-cards-database
I would never have bothered doing that without AI.
In areas I'm more familiar with, like back-end software, it's maybe more of the 2x or 3x.
iOS submissions are way up.
What I havent seen is existing open source people develop 10x faster.
Or, anything vibe coded get popular which wasnt just a gimmick.
Just a tsunami of slop.