I'm personally much more towards the "I struggle to get it to work well" side, but there have been so many anecdotes from people who are earnestly saying it does all their work for them and that they hit enter, go to bed, and wake up to working software. I'm trying to imagine a hammer or band saw that produces such wildly different results in the hands of different users.
A lot of pithy responses to your comments that contain some pretty bad-faith assumptions about AI "power users", but to answer your question I think part of it genuinely does come down to a "skill issue" - or at least familiarity.
AIs really do have huge blind spots, and there are things it can't do well. But on the flip side, there are modes of operation that do get good results and if you keep it in this "zone" you can get amazing results.
For example, here is my CLAUDE.md that's global to all projects: https://github.com/EspoTek/.claude/blob/master/CLAUDE.md
Note the "Working with unfamiliar data or systems" section - it doesn't stop the models from making wrong assumptions but it does get them to test these assumptions with simple experiments and course-correct before human eyes ever see the result.
Setting up things like CLAUDE.md files and skills so that it doesn't make the same mistakes over and over is a big help too. The model doesn't literally get smarter, but it knows how to avoid pitfalls its fallen into in the past.
If you want to have a chat about it earnestly, my email's in my profile.
For the last several years (since GPT 4o) I have written several apps in use by people, making money, entirely with AI/LLMs (me neither writing nor reviewing the code in any meaningful fashion - other than high level architecture, schemas, etc.) and - yes, in a few hours it can do things that would take a normal human weeks (if ever!). But left unconstrained, it will just pile more and more garbage on the pile.
Fable MAX is the thing that gets me much closer to "just send emails" but even then it doesn't look in my specs directory for the spec, it just goes off and does a Jurassic Park-style "It's a Ruby-on-Rails system! I know this!" and disregards all the other ways we send email and writes its own thing. And often I'll go "Where's the button to do X" and it will say "You're right! Nobody asked for it, so this page is an orphan!"
I happen to use Superpower's (https://github.com/obra/superpowers) "brainstorm -> spec -> plan" workflow in e.g. Fable MAX (for anything non-trivial) then I have Fable send to GPT 5.6 Sol XHIGH for execution, with Fable (either the original, or a separate one, depending on the criticality of the task and the blast radius) being the critical reviewer.
Even then, though, I need to continually guide it against the norms and conventions of the codebase/app, because it makes a ton of assumptions.
It doesn't surprise me that if people don't take a fairly rigorous approach to AI-software development then they'll end up with a mess. Even if you buy into "just re-write it" (I happen to think that's where we're ending up) - we aren't there yet and without e.g. a strong test suite, re-writing it is just as likely to create more bugs than it is to fix the existing ones.
So many businesses are pushing AI hard. People with varying levels of job (in)security want to be seen as "leading the charge" into a post-AI era.
LLMs can do some really impressive things, don't get me wrong. But I regularly find massive blunders in the code it produces.
And that actually makes sense if you understand how LLMs operate and the quality of software on GitHub.
r/DiWHY
Quite simply, difference in skill.
The sysadmin that has configured a few applications in their career would likely get annoyed at how many mistakes an LLM makes.
The manager who barely knows their way around a shell, uses an LLM to configure an application in less than the three business days it will take them, and will claim they are 10x more productive now.
Once again, LLMs are effectively great at triggering the Dunning-Kruger effect. The least you know, the more productive you feel.
Try to be aware that mentioning Dunning-Kruger is recursive.
I am tech lead on a front end project that is typescript with react and everything else as bog standard as possible. In the last year the backend devs have been pressured to be full stack, so have been submitting AI PRs to my repo with Claude. The PRs mostly 'work', but invariably have major problems that wouldn't make it through a code review in the before times.
Front end JS on a CRUD app is the literal best scenario for AI. Professional senior devs are using it, and it can kind of work, but still not really.
I've given it the old college try myself recently, and it is at the same time astonishing and disappointing. Yeah it can do a lot, but there's always something wrong with it. The 'something' is not always a bug; it can be a design issue, performance issue, unused code, duplicate code, etc, thus many people don't notice the problems. But I think for those of us who prided themselves on quality, the stuff doesn't work well enough to use. For those who leaned more towards speed, they're overawed. I am concerned for my career, yes.
but also with my instructions it will churn on prompts for hours. It makes use of my local rags by domain (which it also helped me build and maintain). Most recent example: it’s been on my list to set up headscale for like 5 years now. I know how to do it, I’ve read the docs and kept putting it off, partially cuz Tailscale works and has a generous free tier, and partially because I didn’t want to strand assets that I’d somewhat shortsightedly hardened to only accept ssh from my Tailnet.
claude did it all in 3 hours, while I worked on other stuff, prompting for `op` access when it needed keys, and logging each usage, with minimal involvement from me. All with more documentation, roll back points, and “learnings” propagated to the appropriate rags than I could have created in a weekend session. It made mistakes (some that I likely would have made, like region settings) then reasoned about it, corrected and tested, all transparently, all git backed, all posted about on my Mattermost server as it went, so I can easily search it the way I would of an extremely detailed teammate did it.
I don’t know what to make of it all, but letting it churn through my backlog has been eye-opening.
god, that analogy was awful. sorry. xD
Definitely incomprehensible that someone might have a different take than you about the 'goodness' of your prompt outcome...
Claude though is still slow enough that I think it works out to maybe a 50% improvement at my time to do a task, but sacrificing deep knowledge and learning. I worry that I'll never develop the same skills and intuition that senior engineers developed a decade ago because I'm outsourcing a lot of the planning and thinking.