The role of developer skills in agentic coding
martinfowler.com
martinfowler.com
1. Anecdotally, AI agents feel stuck somewhere circa ~2021. If I install newer packages, Claude will revert to outdated packages/implementations that were popular four years ago. This is incredibly frustrating to watch and correct for. Providing explicit instructions for which packages to use can mitigate the problem, but it doesn't solve it.
2. The unpredictability of these missteps makes them particularly challenging. A few months ago, I used Claude to "one-shot" a genuinely useful web app. It was fully featured and surprisingly polished. Alone, I think it would've taken a couple weeks or weekends to build. But, when I asked it to update the favicon using a provided file, it spun uselessly for an hour (I eventually did it myself in a couple minutes). A couple days ago, I tried to spin up another similarly scoped web app. After ~4 hours of agent wrangling I'm ready to ditch the code entirely.
3. This approach gives me the brazenness to pursue projects that I wouldn't have the time, expertise, or motivation to attempt otherwise. Lower friction is exciting, but building something meaningful is still hard. Producing a polished MVP still demands significant effort.
4. I keep thinking about The Tortoise and The Hare. Trusting the AI agent is tempting because progress initially feels so much faster. At the end of the day, though, I'm usually left with the feeling I'd have made more solid progress with slower, closer attention. When building by hand, I rarely find myself backtracking or scrapping entire approaches. With an AI-driven approach, I might move 10x faster but throw away ~70% of the work along the way.
> These experiences mean that by no stretch of my personal imagination will we have AI that writes 90% of our code autonomously in a year. Will it assist in writing 90% of the code? Maybe.
Spot on. Current environment feels like the self-driving car hype cycle. There have been a lot of bold promises (and genuine advances), but I don't see a world in the next 5 years where AI writes useful software by itself.
My experience is the same, though the exact dates differ.
I assume LLMs gravitate toward solutions that are most represented in their training material. It's hard to keep them pulled toward newer versions without explicitly mentioning it all the time.
Does you think this feeling reflects the usual underestimation we're all guilty of, or do you think it's accurate?
I'm using Cursor mostly for exploratory/weekend projects. I usually opt for stacks/libraries I'm less familiar with, so I think there's some optimism/uncertainty to account for there.
I think there's another aspect to progress involving learning/becoming fluent in a codebase. When I build something from scratch, I become the expert, so familiar that later features become very easy/obvious to implement.
I haven't had this experience when I take a heavily agent-driven approach. I'm steering, but I'm not learning much. The more I progress, the harder new features feel to implement.
I don't think this is unique to working with AI. I guess the takeaway is that attention and familiarity matter.
But how long before AI is Good Enough that the point is irrelevant?
It says here that it'll only be another 6 months!
That resonates with me. It actually brings back some nostalgic memories about setting up and constantly tweaking bots in World of Warcraft (sorry, I was young).
There's something incredibly engaging and satisfying about configuring them just right and then sitting back, watching your minions run around and do your bidding.
I get a similar vibe when I'm working with AI coding assistants. It's less about the (currently unrealistic) hype of full replacement and more about trying to guide these powerful, and often erratic, tools to do what they're supposed to.
For me, it taps into that same core enjoyment of automating and observing a process unfold that I got with those WoW bots. Perhaps unsurprisingly, automation is now a fairly big part of my job.
Clone the dependency you want to use in the directory of your code.
Instruct it to go into the directory and look at that code in order to complete task X: "I've got a new directory xyz, in it contains a library to do feature abc. I'll need to include it here to do A to function B and so on"
The weird version mixing bug will disappear. If it's closed source, then just do the documentation.
You need to line up the breadcrumbs right.
#2 is "create a patch file that does X. Do not apply it". Followed by "apply the previous patch file". Manually splitting the task fixes the attention.
Another method is to modify the code. Don't use "to-do" it will get confused. Instead use something meaningless like 1gwvDn, then at the appropriate place
[1gwvDn: insert blah here]
Then go to the agent and say
"I've changed the file and given you instructions in the form [1gwvDn:<instructions>]. Go through the code and do each individually.
Then the breadcrumbs are right and it doesn't start deleting giant blocks of code and breaking things
#3 You will never start anything unless you convince yourself it's going to be easy. I know some people will disagree with this. They're wrong. You need to tell yourself it's doable before you attempt it.
#4 is because we lose ownership of the code and end up playing manager. So we do the human thing of asking the computer to do increasingly trivial things because it's "the computer's" code. Realize you're doing that and don't be dumb about it.
Yes, it can write everything if I provide enough context, but it ain't 'Intelligence' if context ~= output.
The point here is providing enough context itself is challenging and requires expertise, this makes AI ides unusable for many scenarios.
I mostly jest but your comment comes off quite unhelpful and negative. The person you replied to wasn’t blaming the parent comment, just offering helpful tips. I agree that today’s AI tools aren’t perfect, but I also think it’s important for developers to invest time refining their toolset. It’s no different from customizing IDE shortcuts. These tools will improve, but if devs aren’t willing to tinker with what they use, what’s the point?
I don't know if it's explicitely said but if you call it agentic, it sounds like it can do stuff independently (like an agent). If I still need to hand feed it everything, I wouldn't really call it agentic.
Where is the consensus position that is demonstrably more effective than traditional development?
My observation is that it's somewhere close to "start this project and set an overall structure", after that nobody seems to agree.
Much of the marketing around agents is about not needing them. Zuck said Meta's internal agent can produce the code of an average junior engineer in the company. An average junior engineer doesn't need this level of steering to know to not include a 4 year old outdated library in a web project.
We all know that this technology isn't magic. In tech spaces there are more people telling you it isn't magic than it is. The reminder does nothing. The contextual advice on how to tackle those issues does. Why even bother with that conversation, you can just take the advice or ignore it until the technology improves since you've already made up your mind about the limit you or others should be willing to go.
If it doesn't meet the standard of what you believe is advertised than say that. Not, "workarounds" are problematic because they obfuscate how someone should feel about how the product is advertised. Maybe you are an advertising purist and it bothers you, but why invalidate the person providing the context into how to utilize those tools in their current state better?
I didn't say it's magic. I said what it is advertised as.
> The reminder does nothing. The contextual advice on how to tackle those issues does.
No, the contextual advice doesn't help because it doesn't tackle the issue because the issue is "It doesn't work as advertised". We are in a thread of an article whose main thesis is "We’re still far away from AI writing code autonomously for non-trivial tasks." Giving advice that doesn't achieve autonomous writing code for non-trivial tasks doesn't help achieve that goal.
And if you want to talk about replies that do nothing. Calling the guy a Luddite for saying that the tip doesn't help him use the agent as an autonomous coder, is a huge nothing.
> since you've already made up your mind about the limit you or others should be willing to go.
Please read the article and understand what the conversation is about. We are talking about the limits that the article outlined, and the poster is saying how he also hit those limits.
> If it doesn't meet the standard of what you believe is advertised than say that.
The article says this. The commenter must have assumed people here read the articles.
> why invalidate the person providing the context into how to utilize those tools in their current state better?
Because that context is a deflection from the main point of the comment and conversation. It's like in a thread of mechanics talking about how an automatic tire balancer doesn't work well, and someone comes in saying "Well you could balance the tires manually!" How helpful is that?
They just roll eyes and continue to do copies of the files they work on.
I just roll eyes on that explanation because it exactly feels like additional work I don’t want to do. Doing my stuff the old way works right away without having to do some explanation tricks and setting up context for the tool I expect to do correct thing on the first go.
I was personally clocking in about 1% of the openrouter token count last year every day. openrouter has grown quite a bit, but I realize I'm certainly in the minority/on the edge here.
And in the time I line up the breadcrumbs to help this thing to emulate an actual thought process, I would probably alrady have finshed doing it myself, especially since I speed up the typing-it-all-out part using a much less "clever" and "agentic" AI system.
We all know that 2 years is a lifetime in tech (for better or for worse), and we've all trained ourselves to keep up with a rapidly changing industry in a way that's more efficient than fully retraining a model with considerably more novel data.
For instance, enough people have started to move away from React for more innovative or standards-based approaches. HTML and CSS alone have come a long way since 2013 when React was a huge leap forward. But while those of us doing the development might have that realization, the training data won't reflect that for a good amount of time. So until then, trying to build a non-React approach will involve wrestling with the LLM until the point when the model has caught up.
At which point, we will likely still be ahead of the curve in terms of the solutions it provides.
I'm not fuding LLM, I use it everyday, all the time. But it won't make me a better engineer. And I deeply believe that becoming a good engineer helped me becoming a better human, because how the job make you face your own limits, train you to be humble and constantly learning.
Vibe coding won't lead to that same and sane mindset.
But in order to be able to find the right balance, one does need to learn fully what the agent can do, and have a lot of experience with that way of coding. Otherwise the mental model is wrong.
What I do is to write pure functions with LLM. Once I designed the software and I have the API, I can tell the model to write a function which does a specific job where I know the inputs and outputs but I'm lazy to write the code itself.
A lot of the other problems mentioned still hold though.
> There have been a lot of bold promises (and genuine advances), but I don't see a world in the next 5 years where AI writes useful software by itself.
I actually think the opposite: that within five years, we will be seeing AI one-shot software, not because LLMs will experience some kind of revolution in auditing output, but because we will move the goalposts to ensure the rough spots of AI are massaged out. Is this cheating? Kind of, but any effort to do this will also ease humans accomplishing the same thing.
It's entirely possible, in other words, that LLMs will force engineers to be honest about the ease of tasks they ask developers to tackle, resulting in more easily composable software stacks.
I also believe that use of LLMs will force better naming of things. Much of the difficulty of complex projects comes from simply tracking the existence and status of all the moving parts and the wires that connect them. It wouldn't surprise me at all if LLMs struggle to manage without a clear shared ontology (that we naturally create and internalize ourselves).
Totally agree with this point. Software engineering will adapt to work better with LLMs. It will influence how we think about programming language design, as an interface to human readers/writers as well as for machines to "understand", generate, and refine.
There was a recent article about how LLMs will stifle innovation due to its cutoff point, where it's more productive using older or more mature frameworks and libraries whose documentation is part of the training data. I'm already seeing how this is affecting technical decisions at companies. But then again, it's similar to how such decisions are often made to maximize the labor pool, for example, choosing a more popular language due to availability of experts.
One thing I hope for is that we'll see more emphasis on thorough and precisely worded documentation. Similarly with API design and user interfaces in general, possibly leading to improvements in accessibility also.
Another aspect I think about is the recursive cycle of LLM-generated code and documentation being consumed by future LLMs, influencing what kinds of new frameworks and libraries may emerge that are particularly suited for this new kind of programming, purely AI or human/AI symbiosis.
Being on this planet long enough, I've learned this won't happen and in fact the quality of such will degrade making the AI using them degrade and we'll all have to just accept these flaws and work around them like so many myriad technical flaws in our current systems today
I’ll take the other side of that bet. The software industry won’t make things easier for LLMs. A few will try, but will get burned by the tech changing too fast to target. Seeing this, people will by and large stay focused on designing their ecosystems for humans.
My worry is that we get an overnight sea-change like 4o image generation. The current tools aren't good enough for anything other than one-shotting, and then suddenly overnight, they're good enough to put a lot of people out of work.
Is that really hype? I mean there companies or person(s) hyping it up, but there is also Waymo and Pingshan (Baidu) for example actually rolling it out. It's a lot less hype than AI coding.
> Anecdotally, AI agents feel stuck somewhere circa ~2021.
That's only part of the problem. It's also stuck or dead set on certain frameworks / libraries where it has training data.
> I keep thinking about The Tortoise and The Hare.
This implies that AI is currently "smart". The hare is "smarter" and just takes breaks i.e. it can actually get the job done. With the current "AI" there are still quirks where it can get stuck.
I'm thinking back to various promises self-driving would be widespread by 2016. These set a certain expectation for how our roads would look that I don't think has been realized a decade later (even as I've ridden in Waymos/FSD Teslas.)
But no, we have Waymo, I guess, so that means Musk was right?
I learned to only used ORM’s for basic stuff, which they are very much useful, but when things got a little bit complicated to drop back to hand coding SQL.
More broadly, wherever you might have relied on a heavyweight dependency, you can often replace it with AI-generated code tailored to the task at hand. You might think this would increase the maintenance burden, but au contraire: reviewing and taking ownership of AI-generated code is often simpler and more sustainable in the long term than dealing with complex libraries you don’t fully control.
This reminds me of when cross-platform was becoming big for mobile apps and all of us new app developers would put up templates on GitHub which gave a great base to start on, but you really quickly realized you'd have to change a lot of it for your use case anyways.
Something I do quite a lot is throwing back and forth a discussion over a particular piece of code, usually provided with little to no context (because that's my task to worry about), hammering it until we get that functionality correct, then presenting it with broader context to fit it in (or I simply do that part by hand).
Here is how I don't use it: As an agent that gets broad goals that he is supposed to fulfill on its own.
Why? Because the time and effort I have to invest to ensure that the output of an agentic system is in line with what I actually try to accomplish, is simply too much, for all the reasons outlined in this excellent article.
Ironically, this is even more true, since using AI as an incredibly capable writing assistant, already speeds up my workflow considerably. So in a way, less agentic AI empowers me in a way that makes me more critical of the additional time I'd have to invest to play around the quirks of agentic AI.
> a discussion over a particular piece of code [...] hammering it until we get that functionality correct
Care to provide an example of sorts?
> then presenting it with broader context to fit it in
So after you have a function you might convert it to a method of a class. Stuff like that?
For example, recently I needed to revise some code I wrote a few years back, re-implementing a caching mechanism to make it work across networked instances of the same software. I had a rough idea how I wanted to do that, and used an LLM to flesh out the idea. The conversation starts with an instruction that I don't want any code written until I ask for it, then I describe the problem itself, let it list the key points, and then present my solution (all still as prose, no code so far).
Next step, I ask for its comments, and patterns/implementation details how to do that, as well as alternatives to those. This is the "design phase" of the conversation.
Once we zoom in on a concrete solution, I ask it to produce a minimal example code of what we discussed. Then we repeat the same process, this time discussing details about the code itself. During that phase I tell it what parts of the implementation it doesn't t need to worry about and what to focus on, keeping it from going off to mock-up-wonderland.
At the end it usually gets an instruction like "alright, please write out the code implementing what we discussed so far, in the same context as before".
This gives me a starting point to work from. If the solution is fairly small, I might then give it some of the context this code will live in, and ask it to "fill in the blanks" as it were...often though I do that part myself, as its mostly small refactoring and renaming.
What I find so useful about this workflow, as opposed to just throwing the thing at my project directory; it prevents the AI from getting side tracked, lost, as it were, in some detail, endlessly chasing its own tail trying to make sense of some compiler error. The human in the loop (yours truly), sets the stage, presents the focus, and the starting point is no code at all, just an ephemeral description of a problem and a discussion about it, grounding all the latter steps of the interaction.
Hope that makes sense.
You still have to pedal, steer, and balance, but you're much faster overall.
You're right though, you need to be able to steer, but you don't necessarily need to be able to map read.
Case in point, I recently stood up my first project in supabase, cursor happily created the tables, secure RLS rules etc in a fraction of the time it would take me.
To stop it getting spaghetti I had to add a rule "I'm developing a first version - add everything to an SQL file that tears down and recreates everything cleanly".
This prevented hundreds of migration files being created, allowed me to retain and context, and every now and then ask "have you just made my database insecure", which 50:50 resulted in me learning something, or a "whoopsie, let me sort that".
If I wasn't aware of this then it's highly likely my project would be full of holes.
Maybe it still is, but ignorance is bliss (and 3 different LLMs can't be wrong can they?!)
Why are experienced developers so enthusiastic about chaining themselves to such an obviously crappy and unfulfilling experience?
I like writing code and figuring stuff out, that's why I chose a career in software development in the first place.
There is a lot of tedium in software development and these tools help alleviate it.
There are tools and then there are the people holding the tools. The problem is no-one really knows which one AI is going to be.
If I need to call the VideoService to fetch some data, I don't want to spend time writing that and the tests that come with it. I'd rather outsource that part.
But this method of getting there makes me feel like I'm degraded to being the assistant and the machine is pulling my strings; and as a result I become dumber the more I do it, more dependent on crap tech.
Luddism is a strange philosophy for a software engineer.
Luddism and critically evaluating the net benefit and cost of a piece of tech, are 2 very different things.
And the latter is not strange for a SWE at all, in fact I'd say it's an essential skill.
>I like writing code and figuring stuff out
This is an alien mentality to me.
I find pleasure in crafting the solution, sculpting it by hand; putting everything I've got into making it fit the problem like a glove. Coding to me is an artistic way of expressing myself, exploring, always improving; it's part of the fun to me.
And I don't like following instructions in general, don't like being programmed.
So what's in it for you then, if not the problem solving?
And yes, there already are solutions to many common problems. The task then becomes finding a best-fit, and adapting these solutions to the specific needs of the usecase...which is another instance of the same task.
It's like failing to adopt compiled code and sticking to punch cards. Or like refusing to use open source libraries and writing everything yourself. Or deciding that using the internet isn't useful.
Yes, developing as a craft is probably more fulfilling. But if you want it to be a career you have to adapt. Do the crafting on your own time. Employers won't pay you for it.
And when they have forgotten all about how to actually write software, the market is mine.
The excitement around AI coding tools isn't about chaining yourself to a crappy experience — it's about having support to offload cognitive overhead, reduce boilerplate, and help spot potential missteps early.
Sure, the current gen AI isn't quite there yet, but it can lighten the load, leaving more space to solving interesting problems, architecting elegant solutions and "figuring stuff out".
That's why minimizing the generated code is important as well as working on smaller parts at once to avoid what the author refers as "too much up-front work" -- It is also easier mentally when you can iterate on this whole process in seconds rather than days in a pull request review.
That's not why I got into software development. I got into it to make money. I think most people in Silicon Valley these days are the same mentality. How else could you tolerate the level of abuse you experience in the workplace and how little time you get to really dig on that particular aspect of the job?
This is a website that is catered to YC/Silicon Valley. My perspective is going to be common here.
I'm firmly in the problem solver/hacker/artist camp.
Which I guess is why we're more concerned about the current direction. Because we value those aspects more than anything; consider them essential to creating great software/technology, to staying human; and that's exactly what GenAI takes away.
I see how not giving a crap about anything but money means you don't see many problems with GenAI.
AI really is just like us!
There's very often a heap of dev tools, introspection, logging, conversion, etc tools that need to be build and maintained. I've had a lot of luck using agents to make and fix these. For example a tool that collates data and logs in a bespoke planning system.
It is a lot of generated boilerplate off the critical path to build these tools and I just don't want to do it most days.
I am thinking about build systems and shell scripts. I see people everyday going to AI before even looking at the docs and invariably failing with non-existent command line options, or worst options that break things in very subtle ways.
Same people that when you tell them why don't you read the f-ing man page they go to google to look it up instead of opening a terminal.
Same people that push through an unknown problem by trial and error instead of reading the docs first. But now they have this dumb counselor that steers them in the wrong direction most of the time and the whole process is even more error prone.
Time to learn some Emacs/Vim and Awk/Perl
To take an example from the article: code re-use. When I'm writing code, I subconsciously have a mental inventory of what code is already there, and I'm subconsciously asking myself "hey, is this new task super similar to something that we already have working (and tested!) code for?". I haven't looked into the details of the initial prompt that a coding agent gets, but my intuition is that an addition to the prompt instructing the agent to keep an inventory of what's in the codebase, and when planning out a new batch of code, check the requirements of the new tasks against what's already there.
Yes, this adds a bunch of compute cycles to the planning process, but we should be honest and say "that's just the price of an agent writing code". Better planning > ability to fix things.
The hard part is that finding a local optimum for prompting style for one LLM may or may not transfer to another depending on personality post-training.
And whatever style works best with all LLMs must be approaching some kind of optimum for using English to design and specify computer programs. We cannot have better programs without better program specifications.
GP was pondering about code re-use. My typical use involves giving an entire file to the LLM and asking the LLM to give the entire file back implementing requested changes, so that it's forced to keep the full text in context and can't get too off-track by focusing on small sections of code when related changes might be needed in other parts of the file.
I think all of this is getting at the fact that an LLM won't spit out perfect code in response to a lazy prompt unless it's been highly post-trained to "reinterpret" sloppy prompts just as academically as academic prompts. Just like a human programmer, you can just give the programmer project descriptions and wait for the deliverable and accept it at face value, or you can join the programmer along their journey and verify their work is according to the standards you want. And sometimes there is no other way to get a hard project done.
Conversely, sometimes you can give very detailed specifications and the LLM will just ignore part of them over and over. Hopefully the training experts can continue to improve that.
This is a solved problem!
What i completely miss in these LLM parrots-agents-generators, is the learning. You can't teach them anything. They would not remember. Tabula rasa / Clean slate, every time. They may cite Shakespeare - or whatever code scrubbed from github - and concoct it to unrecognizability - but that's it. Hard rules or guardrails for every-little-thing are unsustainable to keep (and/or create) - expert-systems, rule-based no-code/low-code.. has been unsuccessful for decades).
Maybe, next AI wave.
And, there's no understanding. But that also applies to quite some people :/
For example, I have had good success in test first development as a rule. That means that I can make sure it has the specifications correct first.
Agents are a mix of models, prompts, RAG and an event loop.
Lack of reuse
AI-generated code sometimes lacks modularity, making it difficult to apply the same approach elsewhere in the application.
Example: Not realising that a UI component is already implemented elsewhere, and therefore creating duplicate code.
Example: Use of inline CSS styles instead of CSS classes and variables
This is the big one I hit for sure. I think it's a problem with agentic RAG, where it only knows the files it's looked in and not the overall structure or where to look for things, so it just recreates them.That said, I think there are 3 items that are important:
- Quickly grasp a new framework or a new language. People might expect you to do so because of AI's help. 2 weeks might be the maximum, instead of the minimum. The same for juniors.
- Focus on the real important things. So instead of trying to memorize a shell script you are going to use a couple of times per year, maybe use the time to learn something more fundamental. You can also use AI to help you to bootstrap the learning. If you need something for interviews, spend a week to memorize them.
- Be willing to exclude AI from your thought process. If you rely AI on everything, including algorithms and designs, this might impact your understanding.
- Be willing to exclude AI from your thought process. If you rely AI on everything, including algorithms and designs, this might impact your understanding.
Most of the time I'm using AI for problem space mapping (I'm doing dirt simple CRUD dev right now) and decomposition. It's ok at that, but even the deep research mode of Claude leaves some things to be desired.I feel like an editor now, more than an engineer. I know the kinds of things I'm looking for, and I use AI to walk a solution in. Either I use the output of the LLM as-is (for throwaway stuff) or I use it as a jumping off point for my own work _without_ the AI.
I agree. I think it's fine to do so. I usually prefer to write my code without AI (except for bootstrapping it).
In my work as a DE, I mostly use AI to write scripts for me. For example, how to do this in PySpark? I kinda refused to memorize any of these because I'm simply not very interested, and I can always spend a week to memorize the fundamentals if I need.
In my side projects, I use AI extensively. Same as you, I use AI for problem space mapping, or sort of. For example, I have some source code, how do I structure them better? I have read the MIDI standard and thought this piece of binary code means blah, can you please confirm for me? Well AI is OK for these kinds of work.
I have been using LLMs more and more and I think they are great for brainstorming and editing and getting to an 80% solution. But just like an editor needs to know the fundamentals of writing, you still need to know the fundamentals of software to be effective as an engineer.
Side note: I remember when I first came upon a really sophisticated system that made me feel more like an assembler than a developer. It was drupal in the late 2000s. I ran away from that as fast as I could, even though the money was good.
In my side projects I mostly use C/C++ so auto-completion helps me to find a struct member or something similar.
I guess it can become quite complicated when the projects becomes very large.
1. Complete Vibe coding -- greenfield and just playing or doing a small quick prototype
2. Get out of my IDE, but know my code base -- small snippets, methods, classes that add desired functionality but off to the side completely -- and don't run things in my shell that's just wrong
What I don't like it current codebases getting whacked because it decided to downgrade a dependency or lie about a function signature.
I've dabbled with plugins for intellij but wasn't really happy with those. But ever since chat gpt for desktop started interfacing directly with jetbrains products (and vs code as well), that's my goto tool. I realized that I like being able to pull that up with a simple keybinding and it auto connects to the IDE when I do. I don't need to replace my tools and I get to have AI support ready to go. Most of the existing plugins seem to insist on some crappy auto complete, which in a tool that offers a lot of auto complete features already is a bit of an anti feature. I don't need clippy style autocomplete.
What matters here is the tool integration, not the model quality. Better tool integration means better prompts with less work and getting better answers that way.
Example: I run a test, it fails with some output. I had this yesterday. So I asked, "why is this failing" and had a short discussion about what could be wrong. No need for me to specify any detail; all extracted from the IDE. We ticked off a few possible causes, I excluded them. And then it noticed a subtle change in the log messages that I had not noticed (a co-routine context switch) that turned out to be the root cause.
That kind of open ended debugging is a bit of a mixed bag. Sometimes it finds stuff. Mostly it just starts proposing solutions based on a poor analysis of the problem.
What works pretty reliably is:
- address the TODOs / FIXMEs, especially if you give it some examples of what you expect
- write documentation (very good for this)
- evaluate if I covered all the edge cases (often finds stuff I want to fix)
- simple code transformations (rewrite this using framework X instead of Y)
I don't trust it blindly. But it's generally giving me good code and feedback. And I get to outsource a lot of the boring crap.
It's like Rubymine is "home" for me - and chatGPT's macOS client has become another "home" for me so it's quite convenient that they talk to each other now.
I have a little FOMO about Cursor though. ChatGPT will automatically apply its suggested changes to my open editor - but I have the sense Cursor will do a bit more? Apply changes to multiple files? And have knowledge of your whole project, not just open files? Can someone fill me in
If Claude could write the code directly unsupervised, it would go wild and produce a ton of garbage. At least if the code it writes in the browser is any indication. It's not that it's all bad, but it's like a very eager junior dev -- potentially dangerous!
Imagining a codebase that is one or two orders of magnitude larger, I think Claude would be useless. Imagining a non-expert driving the process, I think Claude would generate a very rickety proof of concept then fall over. All that said, I wish I had this tool when developing my previous game. Especially for a green field project, it feels like having access to the internet versus pulling reference manuals -- a big force multiplier.
It outweighs the supposed productivity boost of LLMs by at least one order of magnitude if not more.
I've read comments like this many times and I'm genuinely surprised at the coexistence of "productivity boost" and "15k lines".
Am I the only one that feels like 15k is a tiny project even in non-boilerplatey languages? That's not even past the prototyping stage of a small project.
Am I completely out of touch with a modern project's scale?
"Misunderstood requirements" and "overly complex implementations" are practically our mascots at this point. We're slowly untangling this chaos through better upfront convos and iterative reviews, but man, habits die hard. Anyone else navigating these pitfalls totally unaided by AI?
I think there's a not missing there. Why would you preface that they are categorically and always bad? Makes more sense the other way round.
Also grammar error "effected" instead of "affected" in the footer.
Also, right now engineers are hyper optimized in the code aspects but not thinking about the context into cursor and context out of cursor.
Like the amount of copy paste from Notion / JIRA / Sentry and the amount of output like summarizing the git commits and PRs, slack and other “over communication” you have to do these days. This is the area I think we can more easily automate away.
How do you picture the human in the loop?
It is "Slopware Engineering".
It’s frustrating
I am fiddling with tools like Cursor, Aider, Augment Code, Roo Code and LLMs like GPT, Sonnet, Grok, Deepseek to try to decide whether I can use AI for what I need, and if yes, identify some good workflows. I've read experiences of other people and tried my own ideas. I've burnt countless tokens, fast searches and US dollars. Working with AI for writing code is painful. It can break the code in ways you've never imagined and introduce bugs you never thought are possible. Unit testing and integration testing doesn't help much, because AI can break those, too.
You can ask AI to run in loop, fixing compile errors, fixing tests, do builds, run the app and do API calls, to have the project building and tests passing. AI will be happy to do that, burning lots of dollars while at it.
And after AI "fixes" the problem it introduced, you will still have to read every goddam line of the code to make sure it does what is supposed to.
For greenfield projects, some people recommended crafting a very detailed plan with very detailed description and very detailed specs and feed that into the AI tool.
AI can help with that, it asks questions I would never ask for an MVP and suggests stuff I would never implement for an MVP. Hurray, we have a very, very detailed plan, ready to feed into Cursor & Friends.
Based on the very detailed plan, implementation takes few hours. Than, fixing compile errors and fixing failing tests takes a few more days. Then I manually test the app, see it has issues, look in the code to see where the issues can be. Make a list. Ask Cursor & Friends to fix issues one by one. They happily do it and they happily introduce compilation errors again and break tests again. So the fixing phase that last days begins again.
Rinse and repeat until hopefully we spend a few weeks together (AI and I) instead on me building the MVP myself in half time.
One tactic which seems a bit faster, is to just make a hierarchical tree of features, ask Cursor & Friends to implement a simple skeleton, then ask them to implement each feature, verifying myself the implementation after each step. For example, if I need to log in users, just ask to add logging in code, the ask to add an email sender service, then ask to add email verification code.
Structuring the project using Vertical Slice Architecture and opening each feature folder in Cursor & Friends seems to improve the situation as the AI will have just enough context to modify or add something but can't break other parts of the code.
I dislike that AI can introduce inconsistencies in code. I had some endpoint which used timestamps and AI used three different types for that DateTime, DateTimeOffset and long (UNIX time). It also introduced code to convert between the types and lots of bugs. The AI uses some folder structure for a part of the solution and other structure for other parts. It uses some naming conventions in some parts and other naming conventions in other parts. It uses multiple libraries for the same thing, like multiple JSON serializing libraries. It does things in a particular way in some parts of the application and in another way in other parts. It seems like tens of people are working in the same solution without anyone reading the code of the others.
While asking AI to modify something, it will be very happy to modify things that you didn't ask to.
I still need to figure out a good workflow, to reduce time and money spent, to reduce or eliminate inconsistency, to reduce bugs and compile errors.
As an upside using AI to help with planning seems to be good, if I want to write the code myself, because the plan can be very thorough and I usually lack time and patience to make a very detailed plan.
I wish articles about AI assistance would caveat this at the start. 15k LOC is a weekend hackathon project, which is all well and good, but not reflective of the work that 99% of developers are doing in their day jobs.
There's some heavy assumptions about boilerplate or autogenerated code going on in that estimate, as I don't think very many average 5 characters a second over 16 hours.
honestly, i don't see it. don't you stop to pee?
The issue is that Cursor tends to be demoed for incredibly small, green, and simple projects.
Most of us are working on codebases with at least over 10 million lines. I would love an AI agent that can massive infrastructure migrations with only a bit of oversight. Didn’t Shopify do something like that recently?
I think this is still an area that needs a lot of work.
The industry average seems to be around 100 LOC per day per developer. So if you have a team of 10 that’s only 15 days of work. Once you’re involved in some existing legacy code base it’s likely in the millions.
This is the trick. Human in the loop, not human hiding in an ivory tower after uttering a single command. This is ~effectively what I see a lot of shops doing right now:
"Clean up the codebase please. Apply best practices :D. OH. By the way, heres a laundry list of 100 things to NOT do: <list begins>".
I get a lot more uplift out of use cases like:
"Please generate a custom Stream implementation that is read-only and sources bytes from an underlying chunked representation. Mock the chunk loading part. Primarily demonstrate the ReadAsync method and transition logic between chunks."
The internet is full of javascript/html/css info. Some wrong some obsolete some right and current, but there is data.
How about the more peasant languages?
The deeper into a nerdy domain I go, the less likely it is to understand what's going on. And more broadly, it seems that usefulness steeply declines for larger files and projects. It just can't fit enough context into its window to make sense of complicated things.
As an extreme example, I told it to average two angles, and it just took the mean of them. Even after prompting, it thought nothing wrong of that. (The average of 1° and 359° is 0°, not 180°.) So it goes for many domains outside of webdev, UI, data science, scripting.
An example is asking for simple Kalman filter, limiting to 2x2 matrix to avoid the need for LU decomposition. If you ad to the prompt a constraint to not use Numpy, which almost everything in the corpus does.
Even with LRM's having a high enough top-k accuracy, so that at least one correct solution given in k guesses seems to be the trick.
Perhaps Pyhon+Numpy is a language barrier but the errors without Numpy seem really trivial, similar to what one would see on an obscure language. It is different across different models, but getting stuck generating verification code with divide by zero to giving up and producing code that uses numpy are failure modes I have seen.
Professor Subbarao Kambhampati's explanation really helps here IMHO.
https://bsky.app/profile/rao2z.bsky.social/post/3lkjnrrv2qk2...
"Compiling the signal verifier" in, at least superficially to me, is a good intuition on where these fail.
The limits of Top-K and heavy tail dependance in many tasks will be something painful, I think we will need more expertise and not less among programers just due to the failings of us humans and our over trust of automation etc...
How we change the career path to develop tacit and technical abilities is a big question personally.
Are there any? Honest question.
Why did they choose a circle diagram over a pyramid?
I noticed that I am capable of producing software beyond my own understanding. It wouldn't surprise me if the same is true of AI!
I'm typically pretty gentle in real code reviews but that one is a serious "what the fuck are you even doing" if it were a human.
Adding a top-level context-rule in claude.md doesn't fix it reliably.
Agent generated broken code? An agent can discover that, provide feedback on the pull request, and close it, forcing the coding agent to respond to the feedback.
As long as you have 10 agents doing software engineering analysis for every 1 agent you have writing code, my suspicion is that a lot of this babysitting can be avoided.
At least theoretically.. I haven't got all of this infrastructure linked myself to try.
Not sure if this is an argument against atheism or foxholes.
I suppose it is Thoughtworks after all working to expand mindshare by defining buzzwords.
I recently built out a project where I was able to design 30+ modules and only had 4 generation errors. These were decent size modules of 700-5000 lines each. I would classify the generation errors as related to missing specification -- i.e., no you may not take an approach where you import another language runtime into memory to hack a solution.
Sure, in the past, AI would lead me on goose chases, produce bad code, or otherwise fail. AI in 2025 though? No. AI has solved many quirky or complex headscratchers, async and distributed runtime bugs, etc.
My error rate with Claude-3.7-sonnet and OpenAI's O3-mini has dropped to nearly zero.
I think part of this is how you transfer your expert knowledge into the AI's "mindspace".
I tend to prompt a paragraph which represents my requirements and constraints. Use this programming language. Cache in this way. Encrypt in this way. Prefer standard library. Use this or that algorithm. Search for the latest way to use this API and use it. Have this API surface. Etc. I'm not particularly verbose either.
The thinking models tend to unravel that into a checklist, which they then run through and write a module for. "Ok, the user wants me to create a module that has these 10 features with these constraints and using these libraries."
Maybe that's a matter of 25yrs of coding and being able to understand and describe the problem and all of its limits and constraints quickly but I find that I get one-shot success nearly every time.
I'm not only laying out the specification, but I also have the overall spec in my mind and limit the AI to building modules to my specifications (apis/etc) rather than trying to shove all of this into context. Maybe that is the issue that some people have. Trying to shove everything (prior versions of the same code, etc) into one session.
I always start brand new sessions for every core task or refactoring. "Let's add caching to this class that expires at X interval and is configurable from Y file and dependency injected to the constructor". So perhaps I'm unintentionally optimizing for AI but this fairly easy to do and has probably led to a 5-10x increase in code I'm pushing.
Huge caveat here though, I mostly operate on service/backend/core lib/api code which is far less convoluted than web front-ends.
It's kind of sad that front-end dev will require 100x context tokens due to intermingling of responsibilities, complex frameworks, etc. I don't envy people doing front-end dev work with AI.
Like others I agree humans are not getting replaced anytime soon though. For all the current hype current AI technology is pretty dumb. Give it a decade or so though and everything we are currently doing will seem like Stone Age technology.
Aside: sometimes I really wonder if humanity is trying to automate itself out of existence.