“Vibe Coding” vs. Reality
cendyne.dev
cendyne.dev
For example, here are YC partners quoting a company in a batch claiming "100x speedup" in coding performance compared to the previous month:
https://www.youtube.com/watch?v=IACHfKmZMr8&t=1837s
You can tell this claim is false, because that level of productivity increase would be glaringly obvious to an outside observer; it wouldn't need to be self-reported.
A YC summer batch is 84 days culminating in Demo Day. So a 100x speed improvement would be like a team spending less than 1 day of coding and ending up with something that's on par with Demo Day in terms of functionality. Maybe the design would be wrong, but that wrong design would be just as fully-featured as a Demo Day app.
So if 100x were true, the partners in that video would be talking about how the new batch dynamic is "They get breakfast with a customer, learn something new, have an epiphany, and then later the same day they have their entire app rewritten based on what they learned, and that scratch-rewrite is already at a Demo Day level of functionality." The partners aren't talking about that dynamic because it's not happening. So clearly the self-reported 100x is inaccurate.
Even 10x would result in partners saying "Whoa, in this batch people have a Demo Day-quality app in production by the end of week 1 instead of week 12." The partners have a huge sample size on how much teams get done in what time period, so it would be glaringly obvious to them if this batch were shipping 10x as fast as previous batches.
That external observation would be the headline if it were what the partners were actually seeing. Since that's not the headline, it's clearly not what they're seeing, so 10x can't be the number either.
Also multiples “up” versus “down” are not symmetric. Airplanes are around 10x faster than cars, but that doesn’t mean I’ll be getting to work in 60 seconds.
Note that real engineers helps you come up with new features and test your product and all that and not just add code, adding code was never a bottleneck on just about any problem ever.
So Amdahl's law applies in terms of the time to add code, not in terms of engineers, they don't do the work of 100 engineers, at best they take 100x less time to add lines of code when they know what they wanna make. But that isn't particularly game changing, as adding code is not the hard or even time consuming part.
The reality in which people like me get to do work for US/UK for 4x the salary relative to equivalent work locally, and some of this work is actually cleaning up after folks elsewhere, who being cheapest labor available still got 4x their local salary for this work, and the total is still 4x cheaper than what the US/UK company would pay locally? :).
(I'm only half-joking; in a previous life, I worked on a project with this exact development history.)
The outsourcing market is alive and kicking, and offers a whole spectrum of quality and price. The more to the east of US you are, the easier it is to see :).
> The whole discussion around LLM coding agents feels indistinguishable.
Nah, the difference here is, in outsourcing-to-LLMs scenario, there are no people who do the work and benefit from favorable salary/costs-of-living ratio.
The local QA person could run through 100 or so scenarios a day. The offshore people could do 2 a day. They never improved. The offshore people who are tops aren't cheap.
QA is a whole other story, too. Outsourcing QA is stupid, but even more stupid and short-sighted is not having QA in the first place, and that unfortunately is becoming a norm.
There's lots of false economy going with jobs, too. Getting rid of QA may save you salaries, but the work doesn't disappear - it just gets dumped on everyone else, and now you're distracting much more expensive engineers (software or otherwise), who do a much worse job at it (not being dedicated specialists) and cost more. On the net, I doubt it's ever saving companies any money, but the positives are easy to count, while negatives are diffused and hard to track beyond overall feeling that "somehow, everything takes longer than it should, and comes out worse than it should, who knows why?".
Yes, they were moving in that direction. They centralized the QA team over a suite of probably 15 products, which means no one has any expertise. The QA VP would get mad when QA found bugs because "the dev's were supposed to find all the bugs and the QA was just supposed to just certify the release." The amount of people who don't understand how software dev works in high positions is mind boggling.
They ended up firing all the devs except me and this other guy who gave zero shits and wanted to be a manager. "We" maintained 3 products. Two were pretty standard web apps but one was a full blown decision support system (rules engine) that only I knew. I quit after a few months of killing myself. They paid me a whole lot of money a few years later when they were trying to add features to get a very lucrative government contract.
Are we talking about the reality where the size of the global software outsourcing market is $618 billion and growing? https://groovetechnology.com/blog/software-development/outso...
The thing is how you research, what you expect to get out of it, and what you're willing to pay. There was absolutely a gold rush on bottom-dollar development by cheap overseas developers by management who had no idea how software really worked, and thought they could build a business on cheap offshore development. These were software farms staffed by unappreciated undertrained people from diploma mills. I saw truly shocking things. I saw code written entirely with gotos instead of loops, because the developer had never learned how to write a while loop. The companies spent way more in the long run trying to iterate, ask for changes, and ultimately having to hire higher quality talent for much more money.
I agree with the GP. The LLMs will get better, maybe people will learn that you need an LLM with a skilled developer, or maybe the agents will get good enough to fully drive themselves properly, or maybe just good enough that a non-technical pilot can get good work out of them. Right now, "vibe coding" is largely non-technical people making messy, unmaintainable, insecure code. Some of these are programmers and non-programmers just playing around, but some people are trying to build money-making businesses off this, and it does feel like a very similar situation.
And who cares if it is true? So far programming is one of very few professions when person can set themselves for life in a relatively short period of time. When / if it is gone there will be something else. I have few friends who'd switched to be a handyman. They are doing great from what I see.
If you specifically optimize for it. Most people don't - they specialize and expect to be in their line of work for decades.
> When / if it is gone there will be something else.
There will be something else for young people who are just starting. If you're 20 years into a career and then your line of work disappears overnight, of course you can switch to something else - and enjoy your entry-level salary while competing for jobs with people who are 20+ years younger than you and have no meaningful costs or obligations yet.
Also, on what you're optimizing for. If, like many, you're looking for a job that makes sense, for instance (e.g. working to develop technologies or research that you think can change the world for the better), you're probably never going to strike that particular gold.
That reality in fact never materialized.
I think we will see something similar with LLMs. There will be areas where it will deliver cheaper and faster. There will be areas where it will deliver nothing but disaster. It'll change the industry but not eat it alive. The folks talking about fully autonomous coding on the near horizon are dreaming.
The first one is definitely happening with the LLM bubble where companies really want to pretend that the hard part of the job isn’t understanding what to build and how to do so maintainability.
The second one is going to be more interesting: I expect LLMs to put downward pressure on wages in a lot of places but also for smarter companies to realize that nothing short of true AGI is going to replace the need for people who can actually understand what the customer needs. If I’m right, this will swing the pendulum back towards specialists again – the seagull guys who come in, declare that their favorite framework will solve everything, and leave are more vulnerable to being replaced by an LLM than someone who knows how to code but is also bringing actual business-relevant experience and judgement which an LLM can’t have.
Oh no, now I love LLMs.
But that's just a continuous variant of the discrete-sounding claim that programming will get eaten by AI soon. After all, the "actual business-relevant experience and judgement which an LLM can’t have" is mostly not related to programming - and the better LLMs get at coding, the less value will the programming parts of the skillset have; take it to the limit, and it's just saying the managers and sales people will stay, while software developers will be gone.
Basically, I’m saying people should stop expecting to get six figures for being able to run create-react-app and deploy a container. The analytical and social parts of the job are where I predict LLMs to make fewer inroads because they require non-generic understanding.
This is not to say outsourcing can't ever work, but the situations where it does work are much rarer than what every outsourcing vendor would like you to believe.
I bet a previous client's attempt at outsourcing (well into the 6 figures now) is included in that number... yet the expensive onshore devs outsourcing was supposed to replace are still there 2 years later except now they have to also babysit the offshore idiots and fix their messes.
But hey, the vendor got paid, the idiot executive who fell for their pitch wouldn't want to lose face, so it all gets handwaved away as a continuing success and more money gets thrown into the dumpster fire.
The need for companies to get more value out of less spend won't disappear, nor will the comparative advantage of companies in lower CoL areas of the globe. That's two fundamental incentives on both sides that are aligned, driving the market to find the lowest-energy path from here to there. It'll get there, even if it ends up looking strange (like, idk., maybe cutting out management intermediaries but involving a middleman acting as insurance).
That is, if LLMs won't leapfrog it all and end software dev outsourcing before it started to work well.
We're not there yet, but I don't see anything preventing us from getting there in ~5 years.
(Remember: 5 years ago, SOTA in this space was letting a genetic algorithm poke at an AST and hopefully maybe arrive at a trivial program solving a small algorithmic problem.)
Like some of the other responses, I'm baffled by your comment. Have you not seen what's happened in the past 5 years or so?
Yes, there was an outsourcing craze to India after the .com bubble burst in the early 00s that largely failed - the timezone, cultural differences, and lack of good infrastructure support made it fail.
The past 2 companies I've worked for offshored the majority of their software engineering work, and there was no quality difference compared to American devs. The offshore locations were Latin America and Europe, so plenty of timezone overlap. The companies are fully remote, so what difference does it make if the dev is in your same city or a thousand miles away?
I think offshoring has absolutely put downward pressure on US dev salaries in the past couple years.
There's at least the crucial difference that you had devs in both western countries and traditional "third world", where the 90s view of offshoring was throwing whole processes abroad and only keep "heads" in-house while the remote teams/companies would deal with all the execution, making it inherently difficult to deal with production monitoring.
PS: to your point offshoring to India has become more common but Indian companies are also not that cheap, so we're past the initial framework. Perhaps the same way outsourcing production to China used to be about sweatshops, when it can now be about unrivaled expertise at a cost.
It's about those agents being (mis)used in the very specific blind faith approach of "vibe coding", not least due to the hype merchants and grifters picking up the phrase and running with it shorn of the original cautionary notes about it being useful for bringing a bit of fun back into non-serious coding.
Criticizing the idea (and conflating it with the wider field of LLM coding agents) without understanding that original context is not really any better.
Vibe-coding <> LLM coding agents, which - when used properly - are brilliant for use in serious code and are here to stay.
We already have examples of a model finding more performant sorts [0], given the right incentives and time, and the right system for optimizing (LLMs trained on “average code” probably aren’t it) the computer will best us at creating things for the computer.
Is “vibe coding” real today? Not in my experience, with even Claude code. My hand has to be firmly on the tiller, using my experience and skill to correct its mistakes and guide it. But I can see the current trajectory of improvement, and I’m sure it’ll get there.
[0] https://deepmind.google/discover/blog/alphadev-discovers-fas...
I don’t see this. Board games are fundamentally different from software development problems. The latter have imperfect information, unknown requirements and constraints, fuzzy success criteria, and more.
Not OP, but yes. But that means you don't need a dev, just someone who knows how to spec correctly in English/Jira, right? Is that likely a dev who moved on to PM? Very likely in 2025.
For better or worse, the future I imagine is a Jira plugin or MCP server that can read a project, the LLM IDE client then asks questions to fill in the blanks... and out comes the app.
For many years this will require a human in the loop. But will that human need to know the intricacies of the latest frontend framework? Less and less as time goes on.
Disclaimer: predicting the future is hard.
Regarding "the latest frontend framework" the whole situation is a bit mysterious to me, because somehow everyone keep spending millions of man-hours on yet another react contraption where a static HTML would be enough. From the user perspective all this stuff brings no value, 80% of frontend stuff could have been automated long time ago or just not done at all, yet we keep reinventing the wheel. I don't see how LLMs can change the situation because there was clearly no demand for improved productivity there before.
Not according to your very specific stakeholder demands / environment/naming/data tables/data protection requirements, otherwise you would just use a library.
Those might seem like trivial differences but plenty of things go wrong there, plenty enough that you can't just use a library instead of a programmer, and then they are plenty enough of such errors that vibe coding will also cause issues.
And how is this different from just calling the libraries in the right way to make it adhere to stakeholder requirements?
The statement isn't "its impossible to get an AI to print the code for a right program", but "the work and skills you need to get an AI to print the right program is as much or more than to do it yourself.". That seems to be true for all but trivial programs. Here trivial means you can download a git repo and change some variables to get that result.
In that you need humans that can understand stakeholder requirements, constraints of the domain, and limits of existing software, so that they can write the necessary glue to make everything work.
Thing is, LLMs know more about every domain than any non-expert (and for most software, "domain experts" are just non-domain-expert programmers who self-learn enough of it to make the project work), and they can understand what stakeholders say better than other humans can, at least superficially. I expect it won't take long until LLMs can do all of this better than an average professional.
(Yes, it may take a long time before an LLM can replace a senior Googler. But that's not the point. It's enough for an LLM to replace an average code monkey churning out mobile apps for brands, to have a huge chunk of the industry disappear overnight.)
Someone else made an “exquisite corpse” drawing game.
And another, a way to annotate medical images.
I think of all these things as functional prototypes. It’s obviously not engineering. But it is pretty magical —
I have to say, on they basis of your comment I just decided to try Cursor, and I'm sorry to report it immediately disappointed me.
First thing it did was it told me it found a syntax error in code that compiles perfectly. It went ahead and added a closing brace, telling me that "I've fixed the issue by properly closing the method with its brace. The method now has valid syntax and should compile correctly".
It was never broken! This seems like such a regression of tooling; we have parsers for the purpose of finding wrong syntax, and they don't just make up things like missing braces that aren't missing, and then break your code by inserting them. This seems like developing in Kafka's nightmares.
Should I even bother continuing to evaluate Cursor, when it can take perfectly valid and correct code, and immediately make it worse? What other nightmares will it reveal to me, and do I even want to know? I'm kind of astounded at how bad that first impression was, couldn't have been worse really.
If you let an LLM bootstrap your project, you get tech debt from the get go. It's costing you more time than it saves.
1. Lovable
2. Bolt.new
3. Replit agent
Or actually, I’d put Claude #1. You can make all kinds of cool stuff, just in their UI.
No you didn’t.
Sure, things like Roller Coaster Tycoon exist, but but writing in a compiled language is so much faster, easier, and more broadly accessible than writing in assembly that compilers took over.
“[Developer Chris] Sawyer wrote 99% of the code for RollerCoaster Tycoon in x86 assembly language for the Microsoft Macro Assembler, with the remaining one percent written in C.” - Wikipedia
What a lunatic.
LLMs will greatly increase code production, will they also increase debuggability to match?
At this point, I think you can consider vibe-coding with an LLM to be pretty equivalent to using a fairly junior developer with access to stack overflow. It's going to make a lot of mistakes, it's going to make a lot of questionable decisions. Sometimes it will be able to fix its mistakes, sometimes it will spin its wheels and never fix it. It may make a big ball of mud. The entire project may fail.
Two-three years from now, who knows. I'm pretty sure you still won't see 100% reliability in translating English->code, but you also don't see that even with senior developers.
I don’t have a link handy, but someone has already set up a service that’s “JS time travel debugger + LLM that knows how to use it”
Pretty sure it was a Show HN recently.
This is also implied in the idea of "vibe coding" - don't bother understanding the code or debugging it yourself; if it doesn't work the way you like, just say it and have the model fix it until it gets it right (or you run out of money).
To throw away code you have to understand it, so no its the opposite. Code you don't understand is the hardest to get rid of, so it stays the longest in your codebase.
> if there's a bug and it can't be trivially solved, trash that bit of the code and write it again
How do you know where the bug is if you don't understand the code? There is no known algorithm to take a bug description and return the place in the code the bug is, otherwise bug fixing would be trivial.
Edit: Not to mention that in real production systems your bugs will corrupt the database, and if you haven't set up a logging system etc you will likely not realize for a while forcing you to do a rollback to a very old state losing so much data. You wont last long doing that.
No, you don't.
> How do you know where the bug is if you don't understand the code? There is no known algorithm to take a bug description and return the place in the code the bug is, otherwise bug fixing would be trivial.
Yes, there is.
At worst, the bug is somewhere in the entire project. But you probably have a more narrow idea where the bug is, or when it was introduced. "In module X", "In feature Y", "In the last N days/weeks". Not to mention, for most bugs, `git bisect` is enough to precisely narrow down the problematic change, and doing that doesn't actually require understanding anything about the code.
It all boils down to costs. Even if it takes AI a whole day and 1000 attempts to do what would be a relatively simple fix for an experienced developer, if those 1000 attempts cost less than the developer's work-hours, the business will soon learn to prefer AI over people. If and when we get to that point, is mostly just a function of LLM performance and cost. If they get cheap enough, it'll make as much sense to have human developers fix bugs in code, as it makes sense for you to mend holes in your socks instead of buying them in bulk on-line.
> Not to mention that in real production systems your bugs will corrupt the database, and if you haven't set up a logging system etc you will likely not realize for a while forcing you to do a rollback to a very old state losing so much data.
This really depends on the kind of system you're doing, and the kind of data you're storing.
When I was in undergrad, I knew a few guys who approached every problem by pasting in snippets from Stack Overflow and tutorial sites until the code “worked”. Did not end well…
Look, I understand the sentiment. I too want the code to be done properly, well-engineered and thought through. But recall, this is not our job. We are Professionals, and by modern definition, a Professional does whatever is best for the business. And the business doesn't care about the product - it cares about the product's ability to earn them money. If generating shit code, then throwing it away and generating anew at any sign of bug (or spec change) gets cheap enough, this is what the business will want to do. This is what being Professional will mean.
(You can imagine I don't hold "professionalism" in a very high regard.)
See also: basic goods in meatspace. In developed economies, people generally don't repair clothes anymore, and increasingly rarely anyone bothers with repairing appliances. It's cheaper to just throw it away and buy a new one, than to try and repair it. Hell, even construction and remodeling these days involves a lot more of "affix it here permanently; if you need to move it, just smash it and install a new one" approach.
Why wouldn't the same eventually happen with code?
There is a reason that humans developed analytical problem solving as an alternative to trial and error: when you can do it, it’s more effective, and safer.
That’s not to say that disposable code doesn’t have some interesting implications! One of them is that experimentation becomes a lot cheaper, so it’s faster to navigate the search space of possible solutions to a problem. But just taking a “solution” directly from an LLM without validating its correctness is fundamentally unserious and will be punished by reality sooner or later.
If you don't understand what the code is doing, removing it sounds like a recipe for disaster.
In the past, if you'd tear your shirt, you'd spend time mending it, or pay someone to do it for you. Today, you just throw it in the trash and buy a new one, as it's much cheaper and faster.
Think of any other goods we don't repair anymore. Regardless of their internal complexity and beauty of engineering, and no matter how small the defect is, if it's cheaper to replace it wholesale than to repair it, people end up replacing it.
There's no reason to believe the same won't happen to code.
You're right, there's no reason to believe the same won't happen to code, but there's also no reason to believe it won't similarly end in all kinds of problems that come back to bite us down the line when the goodtimes are over.
With debugging, there have been multiple cases where I was stuck on something, I described the problem in detail and o1 gave me an insightful explanation of what I didn't understand.
It's not magic, but mostly what makes it not magic is it can't respond to very detailed questions. But if I can get it the relevant information into a small space it can draw connections I can't.
>and the right system for optimizing (LLMs trained on “average code” probably aren’t it
For creating website (not apps) it absolutely is there. This is just the first rung on the ladder though. It’s not doing Linux kernel development yet, but that time will come eventually. In between are all the other rungs. AI will climb them one by one.
Grok3 has been amazing, but it keeps on wanting me to do dark mode for an app I’m creating. When I finally said yes, it wasn’t able to get dark ode to work, despite multiple tries. Was funny tbh.
The article is pointing out real limitations of vibe coding today (which you appear to agree with). It does suggest AI coding won’t be viable in the future. You should probably update your comment to say something like, “spot on”.
Sure, the tools aren't perfect, so there's some art to using them now - which the author of TFA seems to be unaware of. Take for example:
> You cannot ask these tools today to develop a performant React application. You cannot ask these tools to implement a secure user registration flow. It will choose to execute functions like is user registered on the client instead of the server.
Of course you can ask them to do it. You can literally ask them to "write better code" (yes, with this exact phrase; see [0]), and you'll get better code. More performant, or more secure - it depends on specifics of the case. Or, you can ask them specifically to focus on security or performance, and you will typically get improvements on those axis.
That's today. Next year, people will know how to prompt the models away from most common failure modes, and the models themselves will be further trained to avoid those same failure modes. In bringing specific problems up, TFA isn't making a convincing argument against future of "vibe coding" - it's literally helping in making it happen.
--
That wasn't the requirement though. He said he wanted a performant or secure app, not "more performant" or "more secure". "more" is trivial, but actually getting to a good state is not.
To actually make a larger program secure or performant you need a unified higher level architecture that is adhered to everywhere, vibe coding can't get you that. It can do some micro optimizations as you said, but it can't do these macro contexts and architecture. You can ask it for suggestions for such architectures, but you can't make it implement a full scale large app with all components using it.
You seemed to have missed the part of the article that clearly lays out the recent progression of AI coding. You’re also refuting the arguments the article makes against the future of vibe coding, but the article doesn’t make any arguments about the future of vibe coding!
We all just end up talking past each other if no one is actually talking about the same thing.
I think you're wrong. I think AI is going to stagnate and only the surrounding tooling will improve, but not enough to get us to the promised land. Arguably we already are seeing that happen. To see the supposed "this is how all code is written now" world AI proponents keep declaring, we're going to need to see improvements to the current AI where they can operate on context windows two orders of magnitude bigger than they are now, while costs for doing so also drop accordingly. Maybe that can be done, I'm betting it can't.
So much technology these days seems to be “get it to 80% so we can demo and cash out” but 80% isn’t just an arbitrary number - it seems to be the point and which the remainder of the work to finish is either very hard or (I suspect) impossible.
It's not available most anywhere. I don't know what exactly what the threshold should be, but it should be usable by most people in first world countries to make the claim "we have it". We don't have it.
We have a fair number of offshore resources that are used for dev. They developers are fully integrated into the team, are in all the stand-ups, and substitute for the usual role of junior programmers. They don't get the grunt-work shoveled on them, they get the same work as everyone else, they're just expected to not be as fast.
In 6 months 2 out of 4 of them been sacked, and surprise, not because we could replace their work with LLM output, but because their use of LLMs was so unrestrained and scattershot the pull requests they submitted had become nightmares. One thing mentioned in the article about unit test creation was something we saw as well. Perhaps this is partly due to working an existing code base where the LLM loses some of its advantage, and certainly some of it was cultural in that progress was thought more important than actual manageable code. The two sacked fellows where told, literally, from my own mouth, multiple times, "You cannot just ask Copilot to write you code, paste the entire thing into Visual Studio with no thought of what has changed, with the end goal of just compiling and meeting the single set of acceptance criteria on your story. You're breaking other things and introducing bugs." It went on deaf ears, and now they're gone. They were nice people, I didn't know how to get through to them, but they were convinced the LLMs were the way to go.
I use LLMs to help write code every day, and I wouldn't want to be without it, but I'm fairly surgical about it. Most of the time Copilot gives you a page of say, React code, or EF Core queries, you have to be really careful about anything you didn't explicitly ask for. Honestly, there is a time savings, but there is not a quality increase. The benefit is subverted by the time it takes to figure out how to ask correctly, the time to vet the output, and the time to fix the little tiny insidious bugs it can introduce.
So, don't go vibe coding and lose your job, is something to think about. I have to admit that it has worn me down meeting these interesting people from far-flung locations only to watch them flounder and get let go.
The company has been relatively ambivalent about the usage of code assistant AI, but during PR reviews it has become very apparent that its seen widespread adoption among the outsourced dev teams purely because of code duplication. Our company has a fairly large number of repositories and bespoke libs for utility type functionality.
In the past, a programmer might have internally said to themselves, "There's no way that somebody hasn't already written this stupid function X or method Y", and they'd take a few minutes to search or reach out to see if it exists within an organization.
Instead during some of the recent code reviews, there has been a huge uptick in core functionality that is very obviously being spit out by the LLM. At best its just extra unnecessary code. At worst it will introduce new bugs since our custom functions often handle business domain specific edge cases that an LLM simply wouldn't know about.
Now non-programmers can fumble forward to a working demo. But junior engineers are walking around with a loaded weapon - if they are not learning from how AI solves a problem or using it like StackOverflow to answer specific questions, they are blowing up their own careers.
The future is we are all product engineers or domain experts. No one is going to want an army of React engineers in 2 years.
This would have been clear from Karpathy's full statement:
> It's not too bad for throwaway weekend projects, but still quite amusing.
> "Vibe Coding" might get you 80% the way to a functioning concept. But to produce something reliable, secure, and worth spending money on, you’ll need experienced humans to do the hard work not possible with today’s models.
The problem is that 80% of the job is a proof of concept at best. 80% is effectively a QA walking into a bar[1].
So there is still a productivity gain for senior developers.
This is exactly the kind of semi-mechanical, low added value work that would greatly benefit from automation, and they really fell on their faces. I really benefit from these models on greenfield tasks where I can delegate minor drudge work, but in this case I honestly think they actually increased the difficulty.
It's the CRUD thing all over again.
And so long as you have some decent-to-solid understanding of coding and testing (this is non-trivial, I've been coding for professionally for ~20 years) then you can direct the machine to put up decent guardrails first, and then you can kinda go nuts and let shit grow, prune it back, repeat.
Basically, if you know what code/tests ought to look and act like, then you can significantly reduce the negative externalities of having LLMs do your coding for you.
Just for fun, I asked the AI assistant in IntelliJ (free trial) to write most of the tests for me. I was actually blown away. The tests were largely really good. In many cases they were more thorough than I would have bothered with. Even when it didn't manage to write good tests on its own, the AI complete as I was writing them myself was incredibly useful. Most of the time it would predict the line I was going to write next, just press tab to accept.
I did have to review the tests, and there were a few minor mistakes I had to correct. The entire process took about 4 hours - and I was trying it out for the first time.
So this is a massive time saver for me, and lets me take on coding in future I would simply not have had the patience to complete.
This has to be a troll no?
They're basically advertising they have a poorly coded, insecure app.
What it excels at: - Boilerplate code that's been written 1000x, which can saps your time and enthusiasm for the meaty problems beyond that.
- Complex DSA work. It has been demonstrated millions of times in training material.
- Simple and tedious tasks like making dummy data for tests and struct literals.
- Tightly scoped refactors.
Where does it falter?
- Mapping your product/business to the code or abstractions needed. I think this is where junior devs struggle to leverage it.
- Doing large scale multi-file refactors without proper specifics, guidance, and context. It also can't write a huge project from scratch. Humans are still need to fit the pieces all together or provide guidance. I think this gap closes soon.
Code quality simply isn't a problem IME. If it didn't one-shot your dream abstraction, you probably weren't specific enough in the prompt. Most human-written code is also junk, so pointing out a minor gaffes isn't really a dunk on AI. It's still a massive productivity booster if wielded by even a half-competent engineer.
To pile on: if a large part of our job is purely mechanical, then there is a bigger problem with our engineering processes and AI can't fix that.
It is! And AI is fixing precisely that. What businesses actually care about (well, 99% of them where code is written) is shipping fast and solving the immediate problem, NOT code quality and craft. It goes against what I want to believe as an engineer. Most problems are not new, they are not hard, they are not sensitive. You will need to start with a good understanding of the business need. It's not that the AI can't code to this. I will often stub out an abstraction, explain inputs/outputs in detail, provide sample data etc. That's all. There are frighteningly few showstopper problems with AI coding at this point, and it's moving so quickly.
We're not at the point where non-engineers are capable engineers with AI, but if you are an engineer not using AI extensively, you are being lapped.
I don't think AI is really fixing business problems, though. I think it's only fixing developer problems. And unfortunately nobody really cares about that except for developers.
I just find it sad that instead of focusing on improving how we build things and reducing the need for so much mindless, tedious, repetious, mechanical work, we're content to just build bad things faster with AI and call it a win.
The AI is doing precisely that: reducing the mindless, tedious, repetitious, mechanical work. And what "vibe coding" wants you to embrace is treating high-level code as if it were compiled assembly: an implementation detail you never want to look at or care about if you can help it.
Yes, in some sense AI isn't fixing anything, because all that "mindless, tedious, repetitious, mechanical" code still exists, it's just autogenerated. I too wish we could've first eliminated the need for that entirely. But we didn't, because most programmers and the industry at large still don't understand where the problem is in the first place. They can't see we've long reached Pareto frontier in our programming languages, that we're being limited by the default paradigm of working directly on plaintext codebase that's a single source of truth.
So yeah, in this sense, LLMs aren't fixing anything - they're just an abstraction layer on top of our exhausted coding paradigm.
This is also what puts many companies out of business and create huge security issues. If AI is not fixing this but making it worse, then that's not improving software engineering.
Those companies you mention just overdid it. Like with everything else on the market, there's a limit to how much value/quality you can optimize away before the end result stops being fit for purpose. However, existence of this limit doesn't stop companies from racing to the very edge of it.
> and create huge security issues.
Security is mostly a solved problem.
Yes, it truly is - at least from the business point of view.
Nobody except attackers and infosec people cares about the mathematical and technical details, or whether your stack or coding practice is secure enough. Not the customers, as they neither understand any of this, nor could do anything about it even if they did. Not the companies, since they manage it at a higher level of abstraction. Whatever holes and vulnerabilities the AI coding introduces, the industry will account for it. Some headlines will be made, some stocks will move, and nothing will change.
FWIW, I don't like either of these things. I'm an engineer in my heart, so it pains me to be constantly reminded that our work is merely means to an end, and matters only to the extent it can't be substituted by some alternative.
Most tech savvy places will avoid it, most good programmers will never encounter it. A bunch or us will make a career out of fixing the mess it makes after it explodes.
My first real job was doing just that at a broker trader which lost 10m on a trade made by an Excel spreadsheet that used a stale yahoo finance API to get exchange rates.
Similar energy
Everybody can vibe cook. - Vercel
Really this whole industry is on another fucking planet. I hate it but it’s so easy to make money.
https://news.ycombinator.com/item?id=43446695
Judging by the comments, most people couldn't even tell it was satire, which goes to show how absurd the hype is right now (and probably why it was buried).
You learn something new every day. Some days that thing does not piss you off. Today is not that day.
"ever since I started to share how I built my SaaS using Cursor"
random thing are happening,
maxed out usage on api keys,
people bypassing the subscription,
creating random shit on db
as you know
I'm not technical so
this is taking me longer that usual
to figure out
- (leo, 2025)And most of linked in (as the algo show me!) is basically this HN post in 100 words or the polar opposite saying how software engineers have had their chips.
It sounds like a euphemism for:
I don't want to work hard.
I don't care about the details.
I don't want to learn new things.
I want somebody else to do my homework.
I don't want to put in the effort it takes to succeed.
I cheated my way through school instead of learning from classes.
I've always had everything handed to me on a silver platter, and I expect that to continue.
I want to put more effort into yapping on Twitter and getting followers that actually doing any work.
The situation seems to reflect the issue that Kernighan's Law refers to, which is that debugging code is twice as hard (perhaps more?) as writing it in the first place. I imagine debugging AI-generated code might be even harder.
Or maybe the numbers looked jumbled and they really didn’t enjoy STEM classes?
You sound arrogant.
Have you ever paid for artwork? You may not have said “I’m not artistic,” but the same criticisms apply in reverse.
Not wanting to learn new things or understand how things work or put in any effort to succeed because you think you can just cheat your way through life is a totally different thing.
And that's what the difference between using AI programming tools responsibly and "vibe coding" is.
A new tool released that basically lets you design a GUI through guided voice prompts. The 50yr old school teacher can “vibe code” all she wants.
The real problem arises when someone makes false claims they’re a developer when they’re just managing AI generated code. Making it your primary interest and then cheating to make it seem true.
It took me 6 months when I was 10 years old to become technical and create a website. If a 10 years old can do it in 6 months, an adult can do it in 1 month if they work hard.
In the right contexts, I find LLMs can speed up my work a lot. But it's nowhere close to being able to replace what I do.
In time the AI will be good enough design whole applications in this vibe-code-y way... But all of the examples I've seen so far indicate that even the best publicly available models aren't there. It seems like every example I've seen has the developer bickering with the ai about something it just won't get right - often wasting more time than they were slightly more hands on. Until the tech gets over that I'll stick to it being the "junior developer I give a uml diagram to so they can figure out the messy parts".
I'm increasingly worried that that's not the same bar employers will have.
A lot of kids are going to enroll college to study CS, computer engineering, software engineering, etc. - and will not finish their degrees until 3-5 years. They might just find themselves redundant (junior positions, that is)
Despite developing LLMs for years I haven't actually used them much in day-to-day work, but asking Claude 3.7 Sonnet my coding questions has been a superior experience to just Googling them (particularly if there are specific functional requirements/constraints)
Python, and other high level languages made a lot of development much faster, but it never lead to reduced engineering needs. Cloud made deploying services massively easier, and as a result we actually have a lot more people working in infrastructure.
Faster development mostly leads to expanded economic viability for new types of software. The real question is what becomes economically feasible if development costs are halved.
"In 1865, the English economist William Stanley Jevons observed that technological improvements that increased the efficiency of coal use led to the increased consumption of coal in a wide range of industries. He argued that, contrary to common intuition, technological progress could not be relied upon to reduce fuel consumption."
(Unless that company somehow has no substantial engineering backlog, which I've yet to encounter anywhere I've ever worked.)
But there's a lot of coders in industries whose core business isn't technology. Knapheide for example, a truck outfitting company where my brother codes. I'd imagine in those companies, being able to do the same work with fewer engineers and less cost would lead to fewer hires. Technology isn't their core product and they aren't being held back by software.
At the moment, a truck outfitting company building a customer CRM optimized for their workflow is an absurd idea: they would need a team of a dozen developers working for a year before they could even get a feel for if it was a feasible project or not.
Add LLM assistance and maybe a team of three developers could get to an initial working version in three months.
At that point, companies that had previously ruled out custom software development entirely may find that it makes sense for them - growing the demand for software engineers as a whole.
That's such a economical fallacy that I'd expect the HN crowd to have understood this ages ago.
Compare the average productivity of somebody working in a car factory 80 years ago with somebody today. How many person-hours did it take then and how many does it take today to manufacture a car? Did the number of jobs between then and now shrink by that factor? To the contrary. The car industry had an incredible boom.
Efficiency increase does not imply job loss since the market size is not static. If cost is reduced then things are suddenly viable which weren't before and market size can explode. In the end you can end up with more jobs. Not always, obviously, but there are more examples than you can count which show that.
But let's assume you have true, fully general AI. Further assume that it can do human-level cognition for $2/hour, and it's roughly as smart as a Stanford grad.
So once the AI takes your job, it goes on to take your new job, and the job after that, and the job after that. It is smarter and cheaper than the average human, after all.
This scenario goes one of three ways, depending on who controls the AI:
1. We all become fabulously wealthy and no longer need to work at all. (I have trouble visualizing exactly how we get this outcome.)
2. A handful of billionaires and politicians control the AI. They don't need the rest of us.
3. The AI controls itself, in which case most economic benefits and power go to the AI.
The last historical analog of this was the Neanderthals, who were unable (for whatever reason) to compete with humans.
So the most important question is, how close actually are we to this scenario? Is impossible? A century away? Or something that will happen in the next decade?
Very strong assumption and very narrow setting that is one of the counter examples.
AI researchers in the 80s already told you that AI is around the corner in the next 5 years. Didn't happen. I wouldn't hold my breath this time either.
"AI" is a misnomer. LLMs are not "intelligence". They are a lossy compression algorithm of everything that was put into their training set. Pretty good at that, but that's essentially it.
People laugh at coders like we are the only manual loom operators when everyone's job, even PotUS can replaced by the AI we can dream will exist.
My thoughts: buy SPX so you own a sliver of the new overlords.
Devs are just coping so hard around LLMs its hard to watch. OTOH the few engineers who have embraced it are excelling.
Our code is better and more robust than it's ever been. Our rate of user-reported bugs have dropped more than 50% since we started "vibe" coding 6-ish months ago.
I could certainly see a possible reduction in engineering team size, but going from 9 to 2 makes me question how much of that reduction was a result of over-hiring in the first place.
Product people love the idea of being able to fire their dev teams, but I'm not sure they understand the implications (some of wich may not become clear for years).
It's interesting that you describe yourself as a developer now.
Because just three months ago, in your first post [1] to HN, you said:
> I'm somewhat non-technical but I've been using Claude to hack MVPs together for months now.
Sure: you might feel as though you have now 10x'ed yourself. But, quite honestly, when the reality is that just a few months back you self-described as "somewhat non-technical", it's clear that (a) you're at such an early stage in your learning and understanding of tech, as a developer, that it's relatively easy to experience bigs gains, and (b) you can't actually have much of an objective measure on this, because you are in fact quite new to the field.
I read a lot of your other comments. To me, even before I had confirmation that you were actually "somewhat non-technical", and fairly new to the field — effectively a junior developer by any real measure — this was already quite apparent to me.
Based upon having been a developer for some decades myself already: I can generally spot those that talk-the-talk — and similarly: I can generally spot those who have non-trivial / deeper experience with various fields of tech.
Powering-up with AI tooling doesn't remedy that. Even if it might seem otherwise from your "somewhat non-technical"-but-newly-empowered position.
Good luck with your coding endeavours though, and with your evangelism.
I have no doubts at all that the world is changing — including how software is developed. But I see your posts for what they are.
you've been a developer for some decades which is why your reality is threatened that your craft is increasingly becoming irrelevant so you had to snoop my profile to find some confirmation that your reality doesn't get shattered
this is nothing new of course. obnoxious neckbeard engineers who don't understand where the world is going have existed since the unix debates on irc. you'll find plenty of people who agree with you on mastodon lol.
Hahaha - no, that's really not accurate at all. On lots of levels. The ability to read another user's comments is there so that anyone who chooses can actually get a better understanding of who they're talking with, and what that person is about. One doesn't have to feel threatened at all to want to use it, one simply has to be intellectually curious, and interested to find out more...
There's no need to try and portray it as a negative, and make out there's something afoot which isn't actually taking place.
Anyone who's been here on HN for any significant amount of time knows exactly what that feature is for — as well as when it might be best to use it. And people absolutely will use it.
It helps separate the wheat from the chaff.
— Please do try and take care that your wide-of-the-mark unnecessary put-downs and name calling don't violate the HN guidelines! (Just for your own good!)
Plenty of room for you to 10x many times over then.
Over the years, I’ve met plenty of folk who have dabbled with software development, before deciding it wasn’t for them - then pivoting to something less technical.
Nah, I don’t feel threatened at all by AI. My job is secure. Tools change, sure. But there’s plenty of years left in software development for sufficiently skilled humans. No matter what a junior-level dev / AI evangelist might claim.
I’ll be cleaning up and properly re-implementing the MVPs that less knowledgeable folk are throwing together, slap dash. For a long while yet. And doing other stuff that AI simply can’t do properly - and quite honestly is quite far from doing.
Your rhetoric betrays your knowledge, and your bravado and insults can’t make up for that in any way.
It’s easy to get enchanted by current generative AI, and believe it far more capable than it is. Particularly if not overly skilled in whatever domain. Particularly if one doesn’t have much of a grasp on how generative AI actually works. Good luck with that.
But that’s why I call it out: yes, exactly, it degrades the conversation when someone is preaching about a new tech, and how it’s gonna change development, and claiming they’re a developer themselves - while not being upfront about the fact that they’ve not actually got much real-world experience as a developer at all in general.
And this kind of thing should always be called out when spotted. It’s just plain disingenuous at the end of the day.
I’ve probably been contracted to fix more broken projects (by devs who royally messed up), than the count of MVPs this person has made, or indeed the number of months they’ve been coding.
But at the end of the day, these kinds of folk simply make us more experienced folk more valuable to those that need a professional service in a bail-out scenario. I’ve got decades of real-world coding experience, and a healthy list of successfully published / deployed projects, including some fairly big clients over the years. My CV speaks volumes, particularly when contrast against someone with little experience in the field of software development. I’ve seen languages and tooling come and go. I’ve headed teams and worked solo. I’ve witnessed plenty of folk like this in my time. It’s certainly not my first rodeo!
Unfortunate that someone chose to downvote me, as opposed to engaging me in conversation as to why my view might perhaps be incorrect or maybe shortsighted - as per the HN guidelines. But no real surprise - I guess that in itself is quite telling here.
Karma points might come and go sometimes, but whatever: I’ve been posting on HN (and other sites) for years, on and off. I’ve no need to try and portray myself as something I’m not, nor portray myself to have skills or experience that I don’t have. I generally post to share my knowledge and experience, because real-world experience adds up over time.
We just didn't need them anymore because LLMs and tools like cursor are insanely good.
Totally believable bro
> I've replaced a product team which had 9 devs 2 years ago with 2 devs with AI
Sure but is everything else the same after that? I doubt it.
it's ok, all craftsmen who got automated away once thought they were special. you're not the first you won't be the last. you will likely be unemployed soon though.
So why wouldn't you keep them? If you're able to produce even more with AI enabled engineers, why downsize? To me, it sounds like a startup's dream to be able to output more without increasing headcount.
This is a real problem that I have experienced on and off. It's getting to the point where everyone on my team is actively looking for alternatives. Generally, I've found Cursor works correctly after business hours. But, it's increasingly giving absolutely useless responses during business hours.
-----
That being said, I agree with many of the author's observations. However, for me, it's not really a breaker. It's not much different than working with an intern or junior engineer. If you ask them to do too much all at once, they come up with bad solutions. Plus, they have a tendency to make "dumb" decisions.
For me, I've found solutions for nearly all of the listed issue. Much of it comes down to being diligent during code review (like you should). For example, the Typescript issue, I come back later to have it fix it.
Specs are the one that still baffles me. It's absolutely terrible at writing proper specs. In particular, it falls into a really bad cycle whenever there are errors. I don't have a solution for this one.
Interestingly the times I've experienced the most weirdness were during extremely not normal business hours (from the California perspective). For 3 nights in a row last week, I found myself coding at/after 2:30am during what were apparently periods of excessive load on Claude Sonnet. When asking Cursor to do things, it would fail, tell me about the high load, and encourage me to try again soon. Well, I just kept clicking the button over and over again, thinking it would eventually be able to handle the request properly, and otherwise continue presenting the error. Not the case!
Incorrect/hilarious things Cursor/Claude did at points during those nights:
- repeat the inquiry back to me in full, then do nothing at all after that
- confidently assert it had located the bug I was looking for, then direct me towards the entire codebase
- assert that it had done what I asked, and request that I approve the changes it wanted to make to my code, which were... nothing, none whatsoever
- (possibly the most hilarious) begin to answer questions in borderline leetspeak, randomly substituting numbers in place of letters in words, before eventually devolving into total gibberish
Mostly just annoying due to the wasted time, though it's possible the entertainment value negated it. I don't expect miracles from Cursor to begin with, nor do I give it wide latitude to change very much in my projects, so the risk of damage wasn't really any worse then than at any other time. Of course, I am not a team working against deadlines on critical projects, just a guy screwing around at 2:30am.
That's morning business hours in Europe and afternoon in Asia.
It’s consistently better during that time.
The issues with Claude Plays Pokemon (an overview here: https://arstechnica.com/ai/2025/03/why-anthropics-claude-sti... ) is essentially due to the 200k context window being finite, which is why it has to use an intermediate notepad. In the case of coding assistants like Cursor, the "notepad" is self-documenting with the code itself, sometimes literally with excessive code comments. The functional constraints of code are also more defined both implicitly and optionally explicitly: For Pokemon Red, the 90's game design doesn't often give instructions on where to go for the next objective, which is why the run is effectively over after getting Lt. Surge's badge as the game becomes very nonlinear.
Although, both vibe coding and Claude Plays Pokemon both rely on significant amounts of optimism around the capabilities around LLMs.
For more esoteric fast changing languages/frameworks it has me chasing my tail in a chain of code updates where each fix breaks something in the n-1th, or n-2th version. Sometimes it's deprecated code, or it halucinates functions that would be valid if your were using a a different language of framework. And sometimes simple coding errors.
But it will get better, a lot better.
The main benefit is that it will let a invested non programmer client build a functional framework prototype and then combine that with a list missing features that a more skilled programmer can flesh out to a first cut solution.
For the first time we 'might' get better requirements with an actual working model instead of having the implementor doing most of the requirements as a first pass from a high level hand wavy requirement. I think we're going to see some amazing tools for this.
What I don't see it doing is creating original algorithms to solve things being done for the first time.
I see statements like this a lot when talking about AI in general. People seem to think it is a foregone conclusion that no limit to LLM model improvement and capability exists. What causes you to believe this and what evidence do you have to back it up?
Now, about your comprehension skills, where is there any mention on my part of there being 'no limit'? In fact I go as far as to speculate on at least one.
Also it is important to be able to review the code because it could be the case that it looks mostly correct but has some subtle errors in it that can mislead you. For example I was trying a couple of different ways of computing some indices that have a bunch of variables and one way had a mask that made no sense involved. "Vibe coding" without being able to check the work of an LLM is almost certain to go poorly, IOW.
With this developers need to level up to architects, get more domain knowledge.
[1] it's unbelievable what a difference in quality 1 year made for chat gpt
If we are the top of an s-curve, the recent samples on that curve would be below the trend line, not above it.
I'm also stricken by the superficiality of analysis like "oh it's just probabilities" from so many devs; might as well say "it's magnets".
You can call it an implementation detail but it's like both a wheel and a wing can take your over some distance but the difference between them is staggering. Wheel will never send you flying (normally)
The good thing about vive coding is it avoids the software development lifecycle completely from the user perspective in a platform that has an integrated SDLC which means from defining idea to ensure visibility in changes to a runtime where the user can see it. In my mind, modifying without a hassle in a controlled environment is what users look for. Software development assisted by AI will be a thing for engineers but Vive coding is aimed for users outside of engineering. I sadly see only a handful of companies being able to pull this off.
But 80% of the functionality is only 10% of the work. The last 20% of the functionality remains and will require 90% of the work.
Wake me up when someone vibe codes a Chrome replacement, or an iOS replacement, or MS Office...
Except we know this won't happen anytime soon because we all know vibe coding isn't very useful beyond toy projects that leverage complex libraries written by actual developers.
If you are working at a place where that quality level is standard -- and let's face it, a large number of companies produce average or below-average quality code (by definition) -- then using an LLM assistant isn't that bad. At least if such an assistant doesn't have some extra flaws beyond producing the best summary of its training data, which is exactly what an LLM does. It actually justifiably replaces developers in such an average-or-below place. But if you are aiming for the top end of the quality scale then there is no way this can be achieved by LLM output. Purely on principle.
This shouldn't even be a controversial opinion. I'm quite surprised every time this is questioned or even just debated.
I think "one shot ready for production code" is what AI cannot do yet. Which is why I am not worried for another 12 months at least :)
Strongly agree with the article, and happy to see so many lucid people, comments and articles on HN that thoroughly deconstruct the "vibe coding" illusion.
Also, Andrej Karpathy really disappointed pushing such brittle BS as a revolution.
I wrote this because I was worried that "vibe coding" was being misinterpreted to mean "any time an LLM outputs code", as opposed to the intended definition of code where you deliberately don't review the code and see how far you can get.
1) Cursor has been crashing several times an hour for me recently.
2) Cursor seems to ignore .cursorrules files. I'm using the json format that's supposed to let you filter on file name patterns (although how that works for cross-cutting agent stuff I don't know).
3) Cursor is obsessed with making sketchy iffy defensive code checking for the most recent symptom and trying to guess and shart its way out of it instead of addressing the real problem? And it's extremely hard to talk it out of doing that, I have to keep reminding it and admonishing it to cut it the fuck out, fail instead of mitigate, address the root cause not the symptoms, and stop trying to close the barn door after all the horses have escaped. It's as of it was only trained on Stack Overflow and PHP manual page discussions.
Moments ago just saw this ad at the front of HN-
https://www.ycombinator.com/companies/domu-technology-inc/jo...
I honestly am not sure if this ad is a joke. I assume not, which is hilarious. Put in you 12-16 hour days for hilariously bad pay, and your onboarding will be doing one of the most pathetic, deadbeat jobs possible which is making collection calls. And your "vibe coding" is to use voice agents to...make collection calls.
Must be pretty grim pickings if this trash is getting advertised on here.
They are 3 different things, and neither of the first two represent anything more than a subset of the capabilities of the latter.
If you don't like LLMs that's cool, but at least take some time to understand the context here.
>OH SHIT you're right! We have duplicate mouse handling:
>HOLY SHIT. Let me analyze what's happening:
>HOLY SHIT. NOW I get it. We're calculating the SAME THING in TWO different places:
>You're absolutely right - I fucked up by removing the loading material completely. Looking at the diffs in shame, here's what needs to be fixed:
> There's a trend on social media where many repeat Andrej Karpathy's words: "give in to the vibes, embrace exponentials, and forget that the code even exists." This belief — like many flawed takes humanity holds — comes from laziness, inexperience, and self-deluding imagination.
I'm going to go ahead and give the author the benefit of the doubt that they aren't literally saying Andrej Karpathy is "lazy and inexperienced", because that claim is obviously absurd.
In general though, I think the author is missing the actual point Karpathy was making! Let's look at his detailed criticisms for the typescript agent run, for example:
> Regularly clones TypeScript interfaces instead of exporting the original and importing it.
> Reinvents components all the time with the same structure without searching the code base for an existing copy of that component.
These are only problems for human codebases. You're not vibing if you are expecting agents to write code the way humans would.
Duplicating interfaces and implementations is inefficient, and would be a nightmare, in a human codebase. But, the code will still work! So if an AI agent is managing the codebase, who cares if it duplicates things all the time?
Maybe it'll see that it did that later and decide to consolidate things, maybe it won't. It doesn't affect the actual outcome of the code, unless you actually look at the code as a human, which is not "vibe coding."
> When told to fix styles with precise details, it will alter the wrong component entirely.
> When told specifically where there are many duplicated components and instructed to refactor, will only refactor the first instance of that component in the file instead of all instances in all files.
> When told to refactor code, fails to search for the breaks it caused even when told to do so.
You're thinking about the code again, gotta stop doing that if you actually want to ~vibe code~. Refactoring code isn't a thing when you're vibe coding, English is your programming language now, the Typescript (or w/e language) is the assembly. You wouldn't spend much time observing the assembly output of your compiler (especially for web dev), so why are you observing the code output of your agent?
If you don't want to vibe code, that's fine, nobody is forcing you to. But if you're going to do it, grade it on the metric that Andrej was actually claiming: that you can get working results on a lot of software projects today by telling coding agents to make some code do something, and then just keep running it with "fix this bug" until it works, and it'll often get to a working result.
He never claimed that the code outputted would be beautiful, from a human perspective, or well formatted, or well architected, or efficient.
Also this article is immensely distracting.
To put it simply, it doesn’t matter if AI does 80% of the work if that last 20% takes 5x longer. As long as you need a human in the loop who understands the code, that human is going to have to spend the normal amount of time understanding the problem.
It seems to me that LLM output creates a similar situation.
a) More knowledge than the most senior developers b) Can work at ridiculous throughput 24/7
Please provide evidence to back up your claim here: papers, studies, etc.
Otherwise this claim can really only be dismissed as being exaggerated / untrue / naive.
But we have to endure these tedious self-congratulatory "mwa ha well it's still not as good as my code" posts.
No shit. Nobody is saying AI can write a web browser or a compiler or even many far simpler things.
But it can do some very simple things like making basic websites. And sure it gets a lot of stuff wrong and you have to correct it, or fix it yourself. But it's still usually faster than doing everything manually.
This post feels like complaining about cruise control because it isn't level 5 autonomy. Nobody should use it because it doesn't do everything perfectly!
It's nothing like that, because cruise control works reliably. There is never a situation where cruise control randomly starts going 90mph or 10mph while I have it set to 60mph. LLMs on the other hand...
This is why I disagree with people who argue (as you did) "it really does speed up simple tasks". No it doesn't, because even for simple tasks I have to check its work every time. In less than the time it takes me to do that, I could've written the code myself. So these tools slow me down, they don't speed me up.
This hasn't been my experience at all. At worst you skim the code and think "nah that's total nonsense, I'll write it myself from scratch", but that only takes a few seconds. So at worst it wastes a few seconds.
Usually though it spits out a load of stuff, which definitely requires fixing up and tweaking, but is usually way faster than doing it all.
Obviously it depends on the domain too. I wouldn't ask it to write a device driver or something UVM or whatever. But a website interface? Sure. "Spawn a process in C and capture its stdout"? Definitely. There's no way you are doing that faster by hand.
https://twitter.com/leojr94_/status/1901560276488511759
https://twitter.com/leojr94_/status/1902537756674318347
Good thing this guy wasn't in charge of an actual business, because if he was it would have been killed overnight.
However, on small tasks and bug fixes, it often fixes the bug before I've even root caused it. It's amazing when I can focus on throwing it information about the bug then have it think in the background while I continue researching. In a surprising number of simpler cases, it one-shots the fix and eliminates any need to root cause (this is a bit easier when it's a feature you understand intimately).
The cycle of tool and framework re-skilling is constant in industry, and those trying to fight the wave always lose. And this one is a tidal wave. UPDATE YOUR SKILLS FOLKS!
Of course, I haven't seen a single vibe-coded thing that I'd want to spend money paying for yet, but that's probably more reflective of the difficulty of making something people want than whether or not you use vibe coding to do it.
Seems like someone is quite bitter about new stuff.
Or would you argue that NFTs actually did live up to the BS that was ascribed to them in some circles during their hype?
But "some circles" ascribe BS to any new technology.
This article is a bit more balanced, though, and clearly isn't criticising use of AI in programming, but specifically the "Jesus take the wheel" style of vibe coding. It's the same old "if you write code as cleverly as you possibly can, you are not smart enough to debug it", but to the next level, where people are writing code that they aren't even smart enough to read.