Breaking up with vibe coding
lucasaguiar.xyz
lucasaguiar.xyz
I've personally found LLMs to have recently crossed the "uncanny valley" of programming for me, meaning that I'm much much more productive than without them.
I find that if you're really good at describing a problem and the constraints you want to solve, using a language it knows well (like Go) and following well known patterns in that language, you can describe thousands of lines of code and get accurate results.
Maybe this isn't vibe coding speicfically, but I actually review every line of code that the LLM puts out. It doesn't take long if you know what you're reading, and the LLM make a weird solution to the problem. Usually if I'm specific about how I want it solved, it does it well.
I also found it useful to say "Please tell me your plan to implement the solutions and ask me about any ambiguities that need clarification." In other words, don't make your own assumptions.
The results are incredible. Thousands of lines of code that maybe not stylistically like mine, but are structurally very accurate to what I'm looking for. Giving function/interface signatures and code examples works wonders.
It really was another leap forward for me with AI. Using it for exploratory analysis of solutions is the best rubber duck I’ve ever had
It contains:
- AI assistant instructions (with things like: "for a new task, ALWAYS first present and discuss different approaches before starting implementation", "update this file with relevant information while working", "unless prompted otherwise, work on the first task in the current tasks section of this file")
- basic info about the project: general, technical (general architecture, dev setup, etc.)
- locations of generally important stuff in the repo
- general context of the current task
- task specific locations of important stuff in the repo
- relevant (non-sensitive) task specific data/variables
- list of the current active/relevant tasks.
I tweak the file a bit for a new task and then just add only the taskfile.md as context, add a '.' as text in the Cursor agent chat window and it's off. Works like a charm.
Next step will be to add MCP servers and just point it at an issue in the relevant issue tracker in the task file.
But yeah, the better you know how to describe what you want, the better the results. Not really a difference from humans, eh?
There was a time when I felt my greatest programming ability is googling, I think that transferred well to LLMs since googling and asking question on stackoverflow required a very similar skill set.
Often grepping through a codebase where I know another example resides of something similar with a vague idea of what i'm looking for helps.
I've had basically zero luck with RAG myself so far, especially as I always try to go the local route for better infosec vs relying on cloud solutions like copilot. But trying to tap an LLM to help grok a public github repo wouldn't cover any sensitive data so I wouldn't mind using the cloud for a task like that.
Especially if the behavior is strange enough that trigger the look into the source.
Also great with figuring a regex pattern in 5-10s instead of getting derailed by SO or Google search.
AI handles the boring bits very well, and to me that is the most energetically draining part of coding (the hard parts are fun and invigorating).
Whether it only fits the current codebase’s context or not doesn’t really matter, you just give it important samples from and info about the codebase at the start of your prompt. My baseline prompt length is ~30k tokens due to that.
I do review and polish everything that’s generated though, as needed. Vibe coding (not even reading / understanding the generated code) I believe was coined primarily for having fun in side projects. If you’re using it for production code then you’re likely holding it wrong.
and if I am familiar, how is prose literally ever more terse than the actual code?
to me it just comes down to which one is less typing and that's still Just Writing The Code Myself as far as I can tell
I mean maybe it'd be different if I spent hundreds of dollars on 4o-pro but trying the free models have made me want to pay even less, it's all a massive waste of time compared to just using Kagi to look things up
All the time. Open an android project and tell Claude to create a social feed, or as a recent example I told Claude to add waitlist support to a registration system in Django. It did a better job than most juniors I've worked with, and I think it cost a couple dollars. I tuned it a bit afterwards, but saved me lots of time and energy. With one one sentence prompt it updated the models, created the migrations, found and updated all the views...
Writing thousands of lines? I'm actively thinking about the specific method.
Reading thousands of lines of someone else's code? I might fall into a more passive mode and miss a problem.
I am more likely to do the reverse: write it myself, have LLMs summarize it or suggest how to test it/break it/whatever.
In many situations a sufficiently-described statement of "here is exactly what I want the code to do" is not significantly easier to write than the code itself. Especially when the AI is doing the annoying tedious bits through autocomplete suggestions, vs letting it try to do the whole thing based on a sufficiently-described prompt.
I think I'll experiment with this as well.
Or just bad design. I use AI a fair amount for my personal work (my employer currently bans it), and what I've found is it's a great accelerant, BUT you have to be super vigilant with it to keep the quality decent. It's very good at producing code that will make it past your average code review but has design issues that are going to make things harder down the road and getting it to refactor these itself can be quite difficult at times. I generally find myself in a loop of asking the AI to do something, doing a few rounds of refinement with it, especially around test cases, then a manual refactor/cleanup pass over the tests, followed by a manual refactor of the code.
When I read accounts of other people gushing over AI, allowing them to do some semi-complicated thing in under an hour, it really makes me worry about how this is going to affect the readability of the average codebase in a few years' time.
The Python solution took me most of a day to write, and I've been curious on how an LLM would handle it.
This is why mathematics has stayed with us in spite of the fact that it's eminently possible to express mathematical ideas in natural language. The precision, clarity, and brevity we gain from formalization is useful.
I think code falls into a similar bucket. Reaching for LLMs makes some amount of sense for throwaway code, but not for serious systems. Code is not merely the raw material with which you happen to realize fuzzy ideas, it is also a means of making once fuzzy ideas precise—and for this benefit to be realized, the communicating human has to participate in the formalization. I don't want to have to deal with a world in which all our systems were authored through fuzzy imprecise language and half-baked, not completely formally analyzed ideas. That's a step backwards. Life isn't only about cost cutting and short term "efficiency".
Sometimes I'm unsure about the high level architecture, the trade offs of the various possibilities. I find LLMs do a good job summarising them.
It's also OK for low stakes personal tools where you're willing to risk avoidable bugs since they'll only affect you.
For code you intend to actually to deploy to other people or maintain long term you need to move beyond vibe coding and actually look at, understand and iterate on the code that it produces.
The moment you start doing that it stops being vibe coding, which is defined by the act of not caring about the code that was produced for you.
I know people have different mental workflows for solving problems, but I never understood the concept of building software prototypes to see if something is possible, unless you're doing scientific research. I can understand building prototype for demo purposes (to avoid technical speak and to have a common reference for discussions).
Either something is possible or it is not. If you don't know, that just means you're lacking knowledge (and maybe it's never been done before). If it's possible, the only thing left to figure out is the cost of the solution (and if it's acceptable). The nice thing about software is that you don't have to commit everything at the beginning. You may build a working skeleton of the solution that you flesh out as the project goes, but I don't see that as a prototype.
And I went off to write like a quick version using each and when I got to trying to handle consuming pagnated responses, it was a terrible terrible fit for one of the frameworks, and required only like 20 lines in the other.
I then took that back to the planning phase and used that to decide how the entire application was going to be structured, as we didn't want two competing frameworks.
I've worked directly on consumer software used by hundreds of millions of users and many people I worked with prototyped in various ways before scrapping it to write production code.
Exactly - and that's what I use prototyping to figure out.
It's not just "is something possible?" - it's "is this possible given the time, tools and skills available to me?"
Here's a really simple example from a couple of years ago: I wanted to build a tool that exported my Apple Notes to a SQLite database.
Is that possible to build? What are the export options available?
I ended up using an LLM to write AppleScript (a language I had no prior experience with) to extract the notes data, and it worked great. I wrote up that prototyping experiment here: https://til.simonwillison.net/gpt3/chatgpt-applescript
> Exactly - and that's what I use prototyping to figure out.
Perhaps "prototyping" is an inappropriate descriptor here. A better one might be "one-off disposable code" considering the example described.
Peer developers could easily interpret prototyping an AppleScript ETL as only requiring minimal research into identifying its applicability and subsequent reusable definition.
However, a "one-off disposable AppleScript approach" expresses minimal investment (time) and also conveys a higher risk tolerance to potential incorrect code an LLM might generate.
HTH
In a one-off I don't care about a few output errors that I can fix by hand once. Good enough is enough.
A prototype requires to explore the limitations and estimate the effort required to eventually reach correctness. This estimate juggling time, cost, motivation, knowledge and "human resources" together.
It's not always you as in yourself but you as in your company or your team. You yourself might know it is possible but management might not believe in it. They ask for a prototype to get you the resourcing for the project to proceed.
Because everyone knows how much cooler the "hacker / coder" man scenes in movies would have looked like if they were just prompting ChatGPT.
More seriously, my main issue with vibe coding (outside of throwaway prototypes) is that using the "code runs" as a metric isn’t a good way to tell whether the code is actually correct. It just feels like a great way to introduce sublte edge case bugs later downstream. You do your application a disservice by not at least taking a minimal amount of time to review the LLM output.
$100,000 / (50 weeks * 40 hours) => ~$50 per hour.
Extremely conservatively, if this saves a developer ~30 minutes per day, it pays for itself. If you spent $100 per day, it would still be worth it.
And when it sends the same developer down a rabbit hole of "something weird is going on" for a couple days (using your math - 2 x 8 * $50 / hr = $800), does it still pay for itself?
If you lose 8 hrs to weirdness, but save 10 hrs somewhere else, then yeah it does pay for itself.
My argument is that it doesn't need to improve overall productivity much to have positive ROI. Whether you agree that these tools improve productivity or not is another question.
> My argument is that it doesn't need to improve overall productivity much to have positive ROI.
That makes sense.
Problem is, there exists a paradox when relying upon LLM's to make software solutions:
People use them to "save time."
Saving time by outsourcing understanding to LLM's
circumvents learning.
Without learning, the tool becomes a crutch.
When the crutch breaks, there is nothing behind it to fall
back onto.First I create the trunk (a broad outline of what I wish to accomplish). Then I concentrate on filling out the branches (the individual bullet points of that broad outline). Then I drill down into the twigs and then the leaves.
It's a different workflow than what I used to follow before LLMs arrived, where I tended to build bottom-up, rather than top-down. But it seems to keep both me and the LLM on track.
Even though it's structured, is it still vibe coding? Well, I think of vibe coding mostly in terms of the psychology. If whatever I'm doing induces a feeling of flow, then that's the vibe I'm going for. https://en.wikipedia.org/wiki/Flow_(psychology)
> First I create the trunk (a broad outline of what I wish to accomplish). Then I concentrate on filling out the branches (the individual bullet points of that broad outline). Then I drill down into the twigs and then the leaves.
This is also known as "top-down", Structured Programming[0], and/or Imperative Programming[1].
The only difference here is the language used.
And no, the vibe is largely the same as before, there was google, then stackoverflow, and then now here's an AI that would most likely take care of the specific question and generate boilerplate if you frame the question correctly.
Framing the question correctly isn't a vibe, it requires logical thinking that feels almost like coding, and results need to be checked. Even then, I usually don't go a day without spotting some sort of hallucination (correcting obvious wrongs during prompting doesn't count). And testing is essential because sometimes that's how you can catch this kind of hallucination.
I do find learning in terms of project and even coding itself improves a lot faster with a capable LLM. It introduces a lot of idiomatic ways to code, and provides a lot of practice to fix whatever didn't work or add things that are just easier to type then asking it to fix.
Instead treat the model output as a spike. By this I mean - first ask for the feature, then put the changes into a patch or stash. Reset with `git reset --hard`, then hand-author the change using the LLM output you stashed as a reference.
It is also awesome to one-shot a giant feature with Gemini, thanks to its long context:
> cd $REPO_ROOT && repomix && llm -m gemini-2.5-pro-exp 'attached is my entire codebase. Please add $FEATURE. In your output give all file paths and find/replace blocks needed for the implementation plus any commands and other instructions I will need.' -a repomix-output.xml
(Thanks Simon Willison and the creators of repomix.)
Welcome to the wild west of what we should call “agentic coding”. Serious professionals are doing this stuff, it isn't pure vibes. You’re at the forefront, upending the way software has been built for decades. Theres huge ROI right now hacking together your own workflows. While it’s still DIY, enjoy being part of the shared learning process & please share what you learn :)
I find AI very helpful for some programming problems and GH Copilot can be very useful sometimes (especially, when it is not overeager), but this vibe coding business is beyond me.
(I only used OpenHands and Trea in Docker. I haven't tried more premium products due to budget constraints. In hindsight, maybe I should have.)
Less structure, less optimal, but requires lower mental energy. So, it is likely to win, just like social media feeds won over RSS feeds.
And how many people use that? How many who are not geeks?
You're supposed to use openai whisper, not your keyboard, at least according to the guy who coined the term vibe coding.
If all the subsystems are simple and sparesly connected, any bug will likely be isolated and easy to detect.
You can build complex systems using simple subsystems while being code blind.
Context problem is mitigated with sparse architectures.