The problem with "vibe coding"
dylanbeattie.net
dylanbeattie.net
I've changed my outlook on vibe coding after seeing the results of this wholesome vibe coding: https://www.youtube.com/watch?v=4pJUXocn7aE .
This guy didn't just make a small program, he made a product that is attuned to some very stringent requirements. If you gave this project to experienced software developers it would be heavy on process and light on impact, take forever to make, and stupid expensive.
The code is probably garbage. You can tell the UI is weird and jumpy. He's using timed delays to manipulate the browser with keyboard keys. The whole thing would be absolutely shredded here. But it's beautiful. It's crafted for function over form with duct tape and bubble gum but it's the difference between his brother being locked helplessly alone in his body, or being able to express himself and have some level of autonomy. It's just awesome.
LLMs rule for this kind of stuff. Solve your problems and move on. The code is not so precious. You'll learn many things along the way and will be able to iterate quickly with a tireless and semi competent but world knowledgeable assistant.
I would strongly disagree on this count. Hacker ethos, at least as traditionally understood, emphasizes grokking and elegance.
I think you may have an understanding of that word that is not very common.
But that does not relate to "hacking", and most certainly does not relate to "hacker ethos". And while "hacking" is sort of ambiguous, I don't recall ever seeing "hacker ethos" defined by quick and dirty solutions.
The hacker's ethos is about freedom of information, freedom of access (apropos to this video about vibe coding accessibility tools), community, collaboration and so on.
I've never seen or heard "elegance" associated with hackers, hacking, the hacking ethos etc. If I were to go further, elegance as it is used here has some anti-hacker connotations, i.e. it is a gatekeeping term to separate "true" coders from "amateurs" which is very anti-hacker.
Anyway, I don't think I'm conflating these terms, but it could be true that my understanding of the hacker ethos is wrong.
This is where I would strongly disagree. Are you familiar with shibboleths like 1337?
I've had some free time recently so I've been trying to use the various AI tools to do a greenfield project with Rails. I know the stack intimately, so it's been a useful exercise as a check against hype. While I can occasionally trust the tooling to make big changes, most of the time I find that even the small changes are wrong in some important, fundamental way that requires re-writing everything at a later time. Asking the models to fix whatever is broken usually results in an endless cycle of (very junior-like!) random changes, and eventually I just have to go in and figure out the fundamental restructure that fixes the issue.
Would this be good enough for someone writing a one-time-use, bespoke piece of software? Maybe. Would they be able to maintain this code, without having a truly senior-level ability to work with the code itself? No [1].
I'm less fearful of LLM taking experienced programmers' jobs [2], because I fundamentally think this kind of stuff plays in the same end of the market as knockoff designer goods, Ikea furniture, and the like. It makes custom software available to an entire customer segment that couldn't have afforded it before, but nobody is going to confuse it with a product made by a team of skilled professionals [3,4].
[1] But to a really inexperienced developer, it looks the same.
[2] I am more fearful of clueless business types, who think this technology obviates the need for programmers.
[3] ...but will unskilled professionals be under pressure? Maybe!
[4] ...and what about skilled people, using the same tools? My current hypothesis is they're beneficial, but ask me again in six months.
[0] https://manuel.kiessling.net/2025/03/31/how-seasoned-develop...
1°) Know how to diagnose and fix critical problems by themselves, because there will be bugs in production for which AI won't be any help
2°) Write maintainable code, AI doesn't care at all of maintainability of code while devs should (must imho) consider pasting AI code as a merge/pull request.
For 2°) IIRC, I read somewhere that more and more code published on i.e. github is AI-generated. While this means the quality of code is decreasing, i.e. Copilot is also trained on more and more % of AI-generated code. So the quality of AI-generated code is decreasing globally.
As time passes, I think skilled devs will become more skilled and unskilled will become more unskilled.
The cutoff date for Claude 3.5 is April 2024, that's currently slightly over one year ago.
Things like a loss of context about the nature of some installed dependency or something.
You really have to know what you're doing and read carefully to avoid glossing over things that'll catch up to you later.
This approach is different than blind trust in the output from a place of ignorance.
firefox("find cats") -> gimp ("carton filter cats")-> gimp("compose into partypic") -> google drive ("upload to cat folder")
The day of only a few people around the factory floor to babysit the robots will come, but lets keep celebrating the day they start unloading them from the delivery trucks for installation.
Even before AI much of software development was managing overseas teams.
>If you gave this project to experienced software developers it would be heavy on process and light on impact, take forever to make, and stupid expensive.
No, it wouldn't. Every single software developer knows how to make programs which "work well enough". This claim is totally ridiculous, it also is not a "product" in any meaningful sense.
What the article is pointing out is that you can not sell software like this. It is not a "product" at all, it is a hobby project. And you can not develop a hobby project like you would develop a commercial product.
Just think for a second what it would take to sell this.
And for the rest of the companies that embrace the 30%~ efficiency spike, it will just accelerate our work goals faster.
I like to use it on stuff that we wanna do to enhance the UX but rarely sees the light of day. Plus my wrists have never felt so good since replacing boring syntax choirs with LLMs.
The author goes out of their way to play up the toy aspect of the LLMs when driven by the inexperienced. No mention is made of them being used by experienced developers, but I get the impression he feels they aren't useful to the experienced.
I'm just playing with a small client/server tool (rsync but for block devices), and vibe coding allowed me to figure out some rough edges of my design document, and experiment with "what would it look like built with threads? As async code? As procedural?) I would never have done that before because of the sheer amount of time it would take, I'd have picked what I thought was best and then lived with the sunk cost falacy until I had no other choice. It did a pretty good job of catching reasonable exceptions in the communication code which I probably would have just let throw a traceback until I got a bug report. It did a great job of adding logging and debugging options. And I had it take 4 attempts at making a "fancy" --progress display and picked the one I liked the best.
LLMs give a level of exploration to software development that we haven't seen since the good old days when HP set up several groups to each build their next generation workstation environment, and then picked the best one.
IMHO, the experienced software developers stand at a unique position to take advantage of these LLMs that far outstrip what an inexperienced developer can do.
The problem, as with COBOL, would be the loss of the existing pipeline that created new competent high-level developers.
I, for one, welcome this change very much, disastrous as it sounds.
If you choose to let the AI do all the coding and never dig into the actual code, you'll learn more about how to communicate with an LLM than you will about how to write code.
That is risky. The AI might just hallucinate an explanation.
An LLM, for all its flaws, can do these things exceedingly well (provided it has good source material, of course). This gives you the ability to get answers to your questions and correct your mistakes. I'd argue that these benefits far outweigh the errors an LLM makes in breaking down and explaining material, and make an LLM better than a book in many (not all) cases.
Which is why no one reads only one book on a particular subject, especially if one is new to the domain. Often you bring in an alternate, but correct perspective from another book. The risks of hallucination with LLMs' responses outweighs the advantage of multiple perspectives (even if a broken clock is right twice a day, you still want a working one)
> The risks of hallucination with LLMs' responses outweighs the advantage of multiple perspectives
You can ask a model a question multiple times. You can ask different models the same question. You can ask the same question different ways. You can do this automatically with agentic frameworks or reasoning models. "Multiple perspectives" are compatible with language models.
LLMs are for generating text, not producing knowledge. The actual knowledge come from the signal/noise ratio of the training data, not the training method.
It's like recording a violin in a noisy environment, encoding it with a low-bitrate lossy encoder and trying to get the PCM of only the violin back.
We've tried to use a better environment for recording specific notes from the violin (better training data), adjust the parameter of the lossy encoder (tuning the model), mix other violin samples to the respone (RAG, MCP), and redefining what a violin should sound like (LLMs marketing).
But the simpler solution is to record it correctly and play it later (AKA going to the source).
This doesn’t fall under my understanding of the phrase “vibe coding”. In the tweet from Karpathy which many point to for coining the phrase, he says that when vibe coding you essentially “forget the code exists”. I think it’s distinct from regular ol LLM-assisted development
I initially had Gemini 2.5 build some code and then did a code review of it. I tweaked my design document and then had Claude and ChatGPT take a stab at it and at this point I wasn't looking at the code at all. These implementations had some deadlock or early termination problems so I asked for some different async and threaded implementations, but ultimately they went off into the weeds in ways I couldn't really reason about, so I went back to a straight select/procedural implementation and had it start tracking outstanding blocks, which got something working reliably. Then I asked for fancy progress displays, tried several attempts at that. Then at this point I started cleaning up the code some, had it add type annotations, etc...
Big swath of pure or nearly pure vibe coding in the middle there.
I was really curious about how it helped you iron out the problems but couldn't figure it out from your description
Then I started seeing issues with how I had phrased the terminal conditions, the client terminating immediately upon reaching the EOF on input file, and not reading any further server messages. I considered making a 3-way ACK to signal end, but instead changed the design to tell it to keep track of outstanding un-acked packets and only send the EOF/EOFACK when there were no outstanding messages.
The design is that the client generates a bunch of "xxhash64" messages for each block of data, and the server replies to each one with either "ok" (I already have that data) or "send" (send me that data). The "send" is replied to with either "data" (here's the data) or "zero" (the block is entirely NUL).
If I wanted to open a large world class fast food chain, I wouldn't be cooking home meals then. That would be silly.
Copilot can help me cook the equivalent of a McDonalds special in terms of software. It's good, I think McDonalds is delicious.
But it cannot help me cook a home meal software. It will insist that my home made fries must go in a little red box, and my fish sandwich needs only half a slice of cheese.
Deeply into that metaphor, maybe someone who has only worked for fast food chains might forget that a lot of good industrial dishes are variations of what previously were home cooked meals.
I am glad that copilot can help young cooks to become fast food workers. That really looks like something.
Well, I take pleasure in cooking home meal software. Can you make a copilot that helps with that?
You know what, nevermind. I don't need a copilot to cook home meals. I have a bunch of really good old books and I trust my taste.
It's not like some big company is going to be interested in some random amateur dish, is it? It was definitely not cooked for them.
If you prompt an LLM like an architect and feed it code rather than expecting it to write your code, both ChatGPT 4o and Claude 3.7 Sonnet do a great job. Do they mess up? Regularly. But the key is to guide the LLM and not let the LLM guide you, otherwise you'll end up in purgatory.
It takes some time to get used to what types of prompts work. Remember, LLMs are just tools; used in a naive way, they can be a drain, but used effectively they can be great. The typing speed alone is something I could never match.
But just like anything, you have to know what you're doing. Don't go slapping together a bunch of source files that they spit out. Be specific, be firm, tell it what to review first, what's important and what is not. Mention specific algorithms. Detail exactly how you want something to fit together, or describe shortcomings or pitfalls. I'll be damned if they don't get it most of the time.
You mean like in a fast food chain?
I know how to use it. All of that was already implied in my original comment. Sometimes though, I want to cook without a rigid mindset.
But with an LLM, that number goes to about 500 or so, 200 of which are real code and not definitions of some kind. Truthfully, that’s where the LLMs shine. I have this enum with 50 variants, and need to build a dictionary (with further constraints and more complex objects). That shit takes forever even with cut & paste, unless you code with code, and those one-offs aren’t my cup of tea anymore.
McDonald’s is a system for delivering food fast and cheap at scale. It breaks down when making one burger.
If you’re not a trained chef, you will not pump out Michelin star meals, even using the same recipe.
There is nothing wrong with going with the consensus answer. Sure, you can't invent an O(N) sorting algorithm with the consensus, but LLMs can definitely write good, maintainable code. Michelin star analogy might be an exaggeration, but it can help cook a good homemade meal, possibly even better than most wannabe home cooks.
Maybe AI is useful for some things after all.
I guess you are right, though. AI is useful for some things. At least it reads the prompt before responding.
Precisely.
For most things, yes. But you don't hire experts to tell you what everyone else knows. You hire experts to give you the angle.
For example, an LLM trained on all AWS documentation wouldn't be super useful (possibly less useful than traditional full-text search), because what you really want to know is the things that AWS didn't write, the writing "between the lines", all the things that DynamoDB can't do.
That's perfectly fine, because that's what you expect from human work to begin with. From natural text like blog article or tecnical reports to software changes, all output is expected to comply with patterns we are already familiar with. Heck, look at pull requests, where you ask your audience to evaluate your work hoping to reach a consensus.
That's perfectly fine. It doesn't take a genius to write a webpage or a report.
I found one of my license texts for a hobby project[0] in a Microsoft product[1]. I have no idea how they use it.
Copilot is very poor at understanding how the DSL arrangement of the project works or how to assemble validator trees. Before Copilot many other "enterprise" stuff like JetBrains IDEs had problems auto-completing it. But it was always good enough for me, for my own things.
It obviously have value. Maybe they're lazy and using it for something unrelated like the TLD list. I really don't know. What I know is that something of my amateur dish ended up being served in what I consider to be a fast food chain.
No, copilot is not endlessly variable. It works best with typed languages and an enterprise mindset. That's exactly the point against "vibe coding" they push. MS seems to despise anything that doesn't follow that.
Look, I love typescript and all that stuff. As I said, McDonalds meals are delicious.
But I don't want that approach for everything I code. This project is not the only one with that issue. I noticed the same in a portable shell compatibility layer, it can only help if I make it to look like other script languages.
"You're asking too much, how can copilot help with a language it was not trained on?"
I am not asking it to be good at that. Just don't make asshole posts saying anything other than MS pasteurization is "vibes".
If they paid me, then that would be another story. But they didn't, and used my "vibe" project anyways.
I am not complaining about software I released for free, before anyone tries to go for that.
I am making an observation about kinds of software that copilot can't help with. My kind of software development, in which I might take months to decide a name or redo everything several times just to try a new kind of sauce.
[0]: https://github.com/Respect/Validation [1]: https://support.microsoft.com/en-us/office/microsoft-teams-t...
I don't think your assumption is true. The key factor is the corpus used to train it, and the context you fed it. I already experienced Copilot fumbling references to methods of a simple Java class, whereas it pulled off thinks like ARM templates flawlessly. You need to understand that internally LLMs do probability-based text completion. If you feed them enough context to maximize the probability they output things you expect, they do so.
How can I feed to it context about a DSL that does not exist yet?
My example referred to ARM templates. They do exist, don't they?
For the same reason, they tend to handle XML better than JSON - sure, both are trees, but XML is has redundancy, and that redundancy helps keep the model on the rails, so to speak.
I actually wonder if the perfect LLM language would be something a lot more like COBOL - not in a sense of being similarly high level, but rather verbosity. And perhaps also being closer to natural English, which is, after all, still a lot of its training set, especially for reasoning stuff. For query languages they seem to like SQL the most of all the things I've tried, and I strongly suspect that it's the same underlying cause.
Of course, a new language designed like that would have the fundamental problem that there's no existing corpus in such a language. Then again, if it is also designed such that one can reliably convert e.g. from Java to that language, then perhaps we can still pull that off.
The reason they like SQL more than other query languages is primarily that their training data has orders of magnitude more of it than any other query language, and that advantage is so huge that any other possible advantage would probably have comparatively negligible effect.
No. The likes of Copilot help you cook the meal you'd like, how you'd like it. In some cases it forgets to crack open eggs, in other cases it cooks a meal far better than whatever you'd be able to pull together in your home kitchen. And does it in seconds.
The "does it in seconds" is the key factor. Everyone heard of the "million monkeys with a typewriter" theorem. The likes of Copilot is like that, but it takes queues and gathers context and is supervised by you. The likes of Copilot is more in line with a lazy intern that you need to hand-hold to get them to finish their assignment. And with enough feedback and retries, it does finish the assignment. With flying colors.
"Does it in seconds" is kinda true if you ignore the "supervised by you" part. That part - reading through the generated diffs to catch (sometimes very subtle) errors - can take a while.
And sure, you'd have the same problem with the intern. But the intern learns from you. The LLM will make the same mistakes again in the same circumstances. Yes, I know they do have memory now, but it my experience that RAG-based stuff is not particularly reliable. I suspect this might improve once hardware is fast enough that you can use a (smallish) helper LLM for RAG with reasonable perf, but we aren't quite there yet.
My main takeaway from vibe-coding is that nobody cared enough to fill that niche and expectation. And it was really frustrating, yet we're getting there through convolutated, inefficient and borderline barbaric means.
People are still lamenting after HyperCard. Automation on windows or macos didn't go anywhere. Shortcuts were a better step into that direction but I feel it got stuck in the middle as Apple wasn't going to push it further. Android had more freedom yet you won't see a "normal" user do automation either.
If we're going to point the middle finger at vibe-coding, I wish we had something to point to as a better tool for the general population to quickly make stuff they want/need.
(Doing it as a professional dev is to me another debate, still with nuance in it IMHO. I'd also love better prototyping tools to be honest.)
I think what the target demographics want is that level of modularity and block building.
Is it:
- just pressing tab and letting copilot code whatever?
- asking an llm to do everything that it can and falling back on your own brain as needed or when its easier.
I guess probably more like the latter. I was just surprised to hear there was a special term for it. I thought everyone was doing that.
Its so fast to do these things. Seems dumb not to use an LLM to interpret an error message if you don’t immediately know the problem. And if it doesn’t work you can alway abort and do it yourself. The whole code gen and debugging process might take 45 seconds.
Sure its dumb if you’re writing some short simple function. And you need to be in the loop and understand all the changes you are making in a professional setting. But that just sounds like the basic workflow of coding with an LLM.
So if you are having the agent write a script to call an endpoint and do something with the result. Have it write a script that does that and then ask it to run the script itself. It will then be able to close the loop by itself and iterate on the code until it does what you want (usually).
It is also very useful to ask the agent to write a test for something and then run that test to close the loop in the same way.
I like to think I learned this from the person who killed YOLO a decade ago (she teaches English).
Copilot anticipating based off what you’re typing? An explicit prompt about what you want?
The input is undefined char*, and that's the magic of it. It's kinda funny we're reaching the phase where some people throw Syntax Error on a missing colon, but AI be like "I want sfssfsfsffsfs" "ayyy got you covered". How the tables have turned.
It throws away decades of software engineering principles in favour of unchecked AI output by clicking “accept all” on the output.
You would certainly NOT use ‘vibe coded’ slopware that powers key control systems in critical infrastructure such as energy, banking, hospitals and communications systems.
The ones pushing “vibe coding” are the same ones who are all invested in the AI companies that power it (and will never disclose that).
An incident involving using ‘vibe coded’ slopware is just waiting to happen.
You would certainly NOT use ‘OUTSOURCED’ slopware that powers key control systems in critical infrastructure such as energy, banking, hospitals and communications systems.
An incident involving using ‘OUTSOURCED’ slopware is just waiting to happen.
... oh, wait...
Yes, outsourcing is famously difficult to get right, and yet we've been doing it in every industry for decades now. And the famous examples become famous when something fails (boeing et. all) but there are also success stories out there. The trick, as always, is in the implementation.
My other point was that every "new thing" gets the same treatment. It doesn't work. Oh, it works but on toy problems. Well, it works on some complicated problems, but it's expensive. Ok, it works on a variety of problems, it's cheap, but it's unmaintainable. And so on, and so forth.
When someone reviews vibe coded patches and gives feedback, what should the reviewer expect the author to do? Pass the feedback onto their agent and upload the result? That feels silly.
How has code review changed in our brave, vibey new world? Are we just reviewing ideas now, and leaving the code up to the model? Is our software now just an accumulation of deterministic prompts that can be replayed to get the same result?
The next step seems that we will write programs in prompts. Maybe we will get a language that is more structured than the current natural language prompts and of much higher order than than our current high level programming constructs and libraries.
Quit his job or stop contributing. Nobody is helped by people blindly commiting code they don't understand.
As it is easier and cheaper to do anything, the result is low-quality products. This of course serves well for the real professionals, who can now sell premium quality products at premium pricing for those who know and want the quality.
With the cost of developing software dropping to near zero, there are a whole class of product features that may not be relevant for most users.
> You haven’t considered encoding, internationalization, concurrency, authentication, telemetry, billing, branding, mobile devices, deployment.
Software, up until now, was only really viable if you could get millions (or billions) of users.
Development costs are high, and as a product developer you're forced to make software that covers as many edge cases as possible, so that it can meet the needs of the most people.
I like to think of this type of software as "average" -- average in the sense that it's tailored not to one specific user or group of users, but necessarily accommodating a larger more amalgamous "ideal" user.
We've all seen software we love get worse over time, as it tries to expand it's scope to more users.
Salesforce may be the quintessential example of this, with each org only using a small fraction of it's total capabilities.
Vibe coding thus enables users to create programs (rather than products) that do exactly what they want, how they want, when they want. Users are no longer forced into the confines of the average user, and can instead tailor software directly to their needs.
Is that software a product? Maybe one day. For now, it's a solution.
I don't think you can ship a fully baked product made exclusively with AI coding yet. But it's *really* useful for investigative product development - e.g. when you have a few different ideas for how something might work but aren't quite sure what's best. I used to regret that I never mastered Figma or other mockup tools - now I never will have to.
They do tend to complicate things, with all of their moats and such. I never wanted a product that did the thing, I really just wanted to do the thing. "Works on my machine" might be good enough if you're unlikely to want to repeat yourself on a different machine.
Would you vibe code your daily driver car's ECU or a high-frequency trading application to use for your 401k? If you were to do these things (more power to you), I rather suspect you'd still do a whole lot of research and critical thinking beforehand, which sort of obviates the "vibes".
At that point, I don't think we'd be using the AI to code something and then switch to a mode where we're now using the recently-coded thing.
We'd just ask the the car to go, and it would happen. Then we'd suggest a strategy for making trades, and that would happen. If the AI decided to write some code and fork a subprocess along the way... that would be an detail that we'd be unaware of.
'Vibe coding' is merely an instance.
Vibe coding, or VB tricks and hacks 25 years ago, or whatever, sure, do it if that works for you, but that's not a product you can maintain for a customer base. It's a program, not a product.
A lot of profitable or otherwise successful software has been built by people who reasonably couldn't be called software engineers or computer scientists or whatever academic title. With Excel, Access, VB, Delphi, Wordpress. I'm sure there's an astrologer somewhere that made OK money from a hack in Delphi or VB for divining the stars on a computer.
It shouldn't be called "vibe" coding, it seems more like glue coding, which some people have been doing and sometimes made careers out of for a very long time. Wordpress was for a long time the big thing in this area, it allowed (probably) millions of people to call themselves web designers and web developers without actually becoming competent in software design.
Many corporations selling software have developers that aren't much better than that. People that might do 'SELECT * FROM table WHERE foreign_id = ?' to check whether there exists any such row, or create full copies of data to add versioning to some product instead of keeping it as metadata, or generate and store hundreds or thousands of gigabytes of dumb data representing business processes (like bookings and such) for the coming century instead of inferring it from rules on the fly. The corporations where I've seen this are generally considered fine and good software companies by customers and competitors alike.
A new RAD, a new Wordpress, a new MICROS~1 Access or a new LotusNotes isn't particularly revolutionary. I suspect the 'hype' is more about getting people in IT to accept the technology as such and not revolt against the broader applications against other, non-technical, people; for control, disciplining, harassment, surveillance, war or whatever nefarious, tyrannical purpose.