No, it's not going to write all your code for you. Yes your skills are still needed to design, debug, perform teamwork(selling your designs, building consensus, etc), etc.. But it's time to get on the train.
Claude 3.5 was actually where it could generate simple stuff. Progress kind of tapered off since tho, Claude is still best but Sonnet 4.5 is disappointing in that it does't fundamentally bring me more than 3.5 did it's just a bit better at execution - but I still can't delegate higher level problems to it.
Top tier models are sometimes surprisingly good but they take forever.
And from reading through the forums and talking to co-workers this was a common experience.
Especially Claude, where if you check the forums everyone is complaining that it's gone stupid the last few months.
Claude's code is all over the place, and if you can't see that and are putting it's code into production I pity your colleagues.
Try stopping. Honestly, just try. Just use claude as a super search engine. Though right now ChatGPT is better.
You won't see any drop in productivity.
Its like terminal autocomplete on steroids. Everything around the code is blazing fast.
Secondly it depends what you're using it for within web dev. One shot an entire app? I did that recently for a Chrome extension and while it got many things wrong that I had to learn and fix, it was still waaaaaay faster than doing it myself. Especially for solving stupid JS ecosystem bugs.
Nobody sane is suggesting you just generate code and put it straight into production. It isn't ready for that. It is ready for saving you a ton of time if you use it wisely.
And I do web dev, the code is rubbish. It's actually got subtle problems, even though it fails less. It often munges together loads of old APIs or deprecated ways of doing things. God forbid you need to deal with something like react router or MUI as it will combine code from several different versions.
And yes, people are using these tools to directly put code in. I see devs DOING it. The code sucks.
Vibe coded PRs are a huge timesink that OTHER people end up fixing.
One guy let it run and it changed code in an entirely unrelated part of the system and he didn't even notice. Worse, when scanning the PR it looked reasonable, until I went to fix a 350 line service Claude or codex had puked out that could be rewritten in 20 lines, and realized the code files were in an entirely different search system.
They're also generally both terrible at abstracting code. So you end up with tons of code that does sweet FA over and over. And the constant over engineering and exception handling theatre it does makes it look like it's written a lot of code when it's basically turned what should be a 5 liner into an essay.
Ugh. This is like coding in the FactoryFactoryFactory days all over again.
IMO LLMs are still at the point where they require significant handholding, showing what exactly to do, exactly where. Otherwise, it's constant review of random application of different random patterns, which may or may not satisfy requirements, goals and invariants.
Also - for most seasoned developers, actual dev activity is miniscule part of overall efforts. If you are churning code like some sweatshop every single day at say 45 its by your own choice, you don't want to progress in career or career didn't push you up on its own.
What I want to say - that miniscule part of the day when I actually get my hands on the code are the best. Pure creativity, puzzle solving, learning new stuff (or relearning when looking at old code). Why the heck would I want to lose or dilute this and even run towards it? It makes sense if my performance is rated only based on code output, but its not... that would be a pretty toxic place to be polite.
Seniority doesn't come from churning out code quicker. Its more long the lines of communication, leading others, empathy, toughness when needed, not avoiding uncomfortable situations or discussions and so on. No room for llms there.
So now with AI, that's even quicker. And I can do it more easily during the half relevant part of meetings, which I have a lot more of nowadays. When I have real time to sit and code, I focus on the hardest and most interesting parts, which the AI can't do.
It is always the talking that transitions "here's quick proof of concept" to "someone else will implement this fully and then maintain". One cannot be catapulted if they cannot offload the implementation and maintenance. Two quick proof of concept ideas you are stuck with and it's already your full capacity. one either talks their way out to having a team supporting them or they find themselves on a PIP with a regular backlog piling up.
And if you are not "churning code like some sweatshop every single day" those hours are not "hey, let's bang out something cool!", it's more like "here are 5 reasons we can't do the cool thing, young padawan".
I don't think that's a relevant metric. "learning" rate of humans versus LLMs. If you expect typical LLMs to grow from juniors to competent mids and maybe even seniors faster than typical human, then there is little point to learn to write code, but rather learn "software engineering with artificial code monkey". However, if that turns out to not be true, we have just broken the pipeline producing actual mids and seniors, who can actually oversee the LLMs.
they might be poor at it, but if you do everything you specified online and through a computer, then its in an LLMs domain. If we hadnt pushed so hard for work from home it might be a different story. LLMs are poor on soft skills but is that inherent or just a problem that can be refined away? i dont know
Most of my day isn't coding. But sometimes it is. On those days, AI helps me get back to doing the important stuff. Sure, I like solving problems and writing code, but where I add value to my company is in bringing solutions to users and getting them into production.
I'm a systems/embedded engineer. In my 30 years of being employed I've written very little code, relatively speaking. I am not a code monkey cranking out thousands of lines per day or even week. AI is like having an on-demand intern who can do that if I need to, however. I basically gained an employee for free. AI can also saving me time debugging because look, I'm old and I really don't write all that much code. I mess up syntax sometimes. I can't remember some stupid C++ rule or best-practice sometimes. Now I don't have to read a book or google it.
AI is letting me put my experience and intuition to work much more efficiently than ever before and it's pretty cool.
It will give the developer a leg up in the future when the mature tools are ready. Just like the people who surfed the 90s internet seem to do better with advanced technology than the youngsters who've only seen the latest sleek modern GUI tools and apps of today.
Using AI, I constantly realize that a-typical patterns are much rarer than I thought.
Humbling.
I'm happy, it's happy, I've never been more productive.
The longer I do this, the more likely it is to one-shot things across 5-10 files with testing passing on the first try.
I think there's an obsession, especially in more veteran SWEs to think they are creating something one of a kind and special, when in reality, we're just iterating over the same patterns.
The teams that have embraced AI in their worlflow have not increased their output compared with they ones that don't use it.
AI Companies have invested a crazy amount of money into a small productivity gain for their customers.
If AI was replacing developers it wouldn’t cost me $20-100/month to get a subscription.
I will get all the goverment IT contracts and make billions in a few months.
Nobody does it because LLMS are a fucking scam, like crypto, and I am tired of pretending is not.
Good for you, though.
I haven't found that to be true
I'm of the opinion that anyone who is impressed by the code these things produce is a hack
Whoever says is time to move to LLMS is clueless.
Cloudfare gets blocked in some parts of Europe on the weekend... Only the DNS, not really blocked.
Football is more important.
They block random Cloudfare IPs related with sport streaming sites.
I have not even seen a CRUD app with real users wrote using AI tools.
Know this: someone is coming after this already.
One day someone from management will hear about a cost-saving story at a dinner table, the words GPT, Cursor, Antigravity, reasoning, AGI will cause a buzzing in her ear. Waking up with tinnitus the next morning, they'll instantly schedule a 1:1 to discuss "the degree of AI use and automation"
Yesterday, GitHub Copilot declared that my less-AI-weary friend’s new Laravel project was following all industry best-practices for database design as it storing entities as denormalized JSON blobs in a MySQL 8.x database with no FKs, indexes, constraints, all NULL columns (and using root@mysql as the login, of course); while all Laravel controller actions’ DB queries were RBAR loops that did loaded all rows into memory before doing JSON deserialisation in order to filter rows.
I can’t reconcile your attitude with my own personal lived experience of LLMs being utterly wrong 40% of the time; while 50% of the time being no better or faster than if I did things myself; another 5% of the time it gets stuck in a loop debating the existence of the seahorse emoji; and the last 5% of the time genuinely utterly scaring me with a profoundly accurate answer or solution that it produced instantly.
Also, LLMs have yet to demonstrate an ability to tackle other real-world DBA problems… like physically installing a new SSD into the SAN unit in the rack.
You can trow all AI you want, but at the end of the day you get what you pay for.
Feed an LLM stack traces or ask it to ask you questions about a topic you're unfamiliar about. Give it a rough hypothesis and demand it poke holes in it. These things it does well. I use Kagi's auto summariser to distil search results in to a hand full of paragraphs and then read through the citations it gives me.
Know that LLMs will suck up to you and confirm your ideas and make up bonkers things a third of the time.
You are doing yourself a huge disservice.
I'm a long time SWE and in the last week, I've made and shipped production changes across around 6 different repos/monorepos, ranging from Python to Golang, to Kotlin to TS to Java. I'd consider myself "expert" in maybe one or two of those codebases and only having a passing knowledge of the others.
I'm using AI, not to fire-and-forget changes, but to explain and document where I can find certain functionality, generate snippets and boilerplate, and produce test cases for the changes I need. I read, review and consider that every line of code I commit has my name against it, and treat it as such.
Without these tools I'd estimate being around 25% as effective when it comes to getting up to speed on unfamiliar code and service. For that alone, AI tooling is utterly invaluable.
AI and LLMs are tools. The best tools tend to be highly focused in their application. I expect AI to eventually find its way to various specific tool uses, but I have no crystal ball to predict what those tools might be or where they will surface. Although I have to say that I have seen, earlier this week, the first genuinely interesting use-case for AI-powered code generation.
A very senior engineer (think: ~40 years of hands-on experience) had joined a company and was frustrated by lack of integration tests. Unit tests, yes. E2E test suite, yes. Nothing in between. So he wrote a piece of machinery to automatically test integration between a selected number of interacting components, and eventually was happy with the result. But since that was only a small portion of the stack, he would have had to then replicate that body of work for a whole lot of other pieces - and thought "I could make AI repeat this chore".
The end result is a set of very specific prompts, constraints, requirements, instructions, and sequences of logical steps that tell one advanced model what to do. One of the instructions is along the lines of "use this body of work I wrote $OVER_THERE as a reference". That the model is building iteratively a set of tests that self-validate the progress certainly helps. The curious insight in the workflow is that once the model has finished, he then points the generated body of work to another advanced model from a different vendor, and has that do an automated code review, again using his original work as a reference material. And then feeds that back to the first model to fix things.
That means that he still has to do the final review of the results, and tweak/polish parts where the two-headed AI went off the rails. But overall the approach saves quite a lot of time and actually scales pretty much linearly to the size of the codebase and stack depth. To quote his presentation note, "this way AI works as a highly productive junior that can follow good instructions, not as a misguided oracle that comes up with inventive reinterpretations."
He made modern AI repeat his effort, but crucially he had had to do the work at least once to know precisely what constraints would apply. I suspect that eventually we'll be seeing more of these increasingly sophisticated but very narrowly tailored tooling use cases to pop up. The best tools are after all focused, even surgical.
ß: Who could have predicted in 1900 that radioactive compounds would change fields ranging from medicine to food storage?
Or, more correctly, they don't work well for my problems or usage. They can at best answer basic questions, stuff you could lookup using a search engine, if you knew what to look for. They can also generate code for inspiration, but you'll end up rewriting all of it before you're done. What they can't do it solve your problem start to end. They really do need a RTFM mode, where they will just insult you if you're approach or design is plain wrong or at least just let you know that it will now stop helping as you're clearly of the rails.
We need to bubble to pop, it'll be a year or two, the finance bros aren't done extracting value from the stonks. Once it does, we can focus on what's working and what isn't and refine the good stuff.
Right now the LLMs are the product, and they can't be, it makes no sense. They need to be embedded within product, either as a built in feature, e.g. CoPilot in Visual Studio, or as plugins, like LSPs.
Clearly others are having more luck with LLMs than I do, and do amazing projects, but that sort of illustrates the point, their aren't ready and we don't have a solution for them to be universally useful (and here I'm even restricting myself to thinking about coding).
Can you give some examples, please?
You’re deliberately disadvantaging yourself by a mile. Give it a go
… the first one’s free ;)