It never does anything interesting or inspired. It never gets halfway into building something and then questions the entire premise of your app. It's automated milquetoast.
If only middle management and c-suite knew that
It's only the craftsman that look up to the superstars / thought leaders, everyone else just wants inoffensive competency (and that's OK I guess).
I asked Sonnet to port an old game I wrote in college to HTML5.
It threw out the audio and graphics system, made it a single HTML file with no dependencies, no assets, everything dynamically generated, added new backgrounds, and new music.
I just asked it for a port, so I was like wtf, though some of the choices it made were better than the original ones.
Would be nice if that "initiative" could be turned on or off on demand though. I haven't tested that. (Some guy told me "if you want creativity, use Claude. To turn the behavior off, use a different model!")
It's nowhere near the quality of hand-crafted expert human code, but that requires a hand-crafting expert human engineer and takes a long time.
I predict that the latter will be reserved for the highest value, longest lived, or special (high performance, high security) parts of systems.
As of June 25th, 2026. Read your comment in four years, and tell me you believe what you've posted.
Maybe one year? Four months?
The spec is the new code. I expect that in 5 years I will rarely write actual code. I will write specs and herd bots.
That doesn't give you good taste, but .. for yer basic line of business or enterprise app, expectations were already low, and most websites have user-hostile design written into the requirements, so the damage isn't too bad.
I do think we've yet to see what the worst case for government contractor software project + vibe coding is. The benchmark is the Canadian gun registry https://calleam.com/WTPF/?p=1949 "In what may be the worst budget overrun in history, the costs to implement a registry of firearms balloons from $2M to $860M". Now add token spend to that.
The science part is that the outcomes are scientifically verifiable, hence tests.
The art part is that unlike the physical world with the laws of physics, there are nearly limitless ways in code to accomplish those outcomes.
A lot of companies are creating AI assistants which take otherwise deterministic processes and makes them nondeterministic.
For example, I work in financial services and deal with a lot of data vendors. One of the big ones very recently added a chatbot to their UI, which on the second question I asked, provided entirely incorrect numerical but with confidence and 3-decimal place precision.
So the chatbot makes it "easier" to ask things, because you don't need to know which tab in the UI to use / or code to write in their scripting interface / or function to call in their Excel interface / or parameters to pass to get the correct answer.
Unfortunately the chatbot also may completely mistranslate your question, call the wrong function/pass wrong parameters and feed you back nonsense confidently.
Who is this helping?
You can't put an AI in a flow which requires reliable results. I don't think people who are used to determinism have coped with that.
The problem is .. what flows don't need determinism? Search results / recommendation engine / ad targeting ?
Arguably the majority of companies & corporate users are using these tools in cases they expect determinism. Email inbox summaries. Search summaries. What is the value of x queries. You'd be shocked.
My favorite Google AI summary bug/quirk that seems to persist (I just tested it again) - "are Lillies OK for cats". The summary starts with "Yes, lilies are extremely toxic to cats. "
I first hit this with another plant (lavender) where the response was much longer and on an iPhone looked like this:
Yes, lavender (both the plant and its essential oils) [line break]
is considered toxic to cats.
That's not the relevant question, because the actual answer to what you asked is "all flows where human judgement is used".
The thing we need to not blindly use current generation AI for is "things where we accept the combination of an untrained (or barely trained) human with no QA".
IMO, the drive to use AI is not only fully automating a lot of things that shouldn't be so, but also revealing how some of them never should have been in the first place. If your human customer support agent makes stuff up and that got your business a penalty fine, you might discipline or fire them; not so easy when it's an AI that replaced a whole call centre in one go, even when the incident frequency is the same.
Ever since Claude Opus 4.7 it's been helping me greatly ship around 3x to 5x faster, and it's good code if I tell it to follow my existing code's structures. Otherwise you have to create .md guideline files and it works.
It's not perfect, but again, 3x to 5x (!!!)
Web apps and CRUDs, if may I ask? Or is AI helping you with something that you couldn't ever do by yourself? I have mixed results across different technologies like frontend, backend, infra and hardware.
I know from experience that to make it a non-toy project I am probably going to need to spend some, if not most of that saved time cleaning up in the future due to technical debt.
For backend, it is mildly useful at best, helpful with boilerplate and when I know exactly what needs to be built. I have not yet firsthand experienced these 10x, 100x productivity gains.
Frontend, forms,tables and generic dashboards is pretty good, I am sure one could get it done faster and better over the long term with proper technique and methods, but I just hate css
It's much easier to take "well now I have an app, which after using for a bit I find it sucks in <quantifiable list of ways>" and then fix them, than it is to start from absolutely nothing.
This reeks of self flattery. Using AI isn't "hard". Any experienced developer can simply pull in a skill into your harness and have it writing code at a similar level to others with the same harness.