I don't particularly like these tools, but I also can't deny that for some things I'm genuinely able to get to the results I want faster with them than without. Yes, some people will use them to produce low quality software, but in my experience it's mostly people who already would have been producing low quality software as well, just at a much slower pace. Having a higher volume of low quality software is a problem, but it's distinct from the claim that the tools not ever being useful in producing high quality software.
It's possible that from reading my rebuttals you'll think I'm one of the people using these tools to produce low quality software, but it's not clear how to falsify that claim. If that's the conclusion you'd draw, then you probably weren't really open to considering rebuttals in the first place though.
There is literally nothing to see in this subthread.
The modern problem with AI is more the opposite, you give the AI a task that is impossible with the tools at hand and instead of saying "That doesn't work", it starts elaborate workarounds to make it happen anyway.
Might?!
Every agentic workflow should be doing this automatically for nearly a year now.
And yes, I agree with the latter. The lengths AI will go to, to get you a working solution is sometimes scary.
My agents have to pass tests meaning that if their LLM hallucinates, the agent tools capture it and not me.
Reading tests and actually catching issues requires a heck of a lot of effort. IME code reviews of tests are often more laborious than reviewing the code itself since so much if it is reasoning about corner cases.
I'm not saying don't do it, but the LOE is high. Done well I'd argue it runs close to the effort involved in just authoring those tests by hand.
And that's ignoring that the "vibe all the things" crowd is explicitly telling people not to read the code. At all.
(And if you think I'm exaggerating: an exec where I work decided to own rebuilding one of our services. They told me how recently the LLM generated multiple thousands of lines of code and they pushed it with minimal review, figuring we'll just fix the bugs as they happen...).
While technically true the hallucination rates on modern models is low and other checks can ensure that by the time a human sees it it is most likely solid.
For research there is more danger as there is less feedback loop other than other LLM scrutinising the first. For research I get it to come to a conclusion but provide me with links so I can judge. More like advanced search.
Isn’t this entirely context dependent? Where did you get the information that modern models have low hallucination rates? I’d love to see the benchmark if there is one, it seems like it would be useful to track.
Most likely? That’s not reassuring at all. So you’re saying the other checks can result in hallucinations?
But making a decision the human reviewer disagrees with or misunderstanding a spec but maybe not asking a follow up isn't hallucination in my book.
Hallucination is just making stuff up without checking.
Points 1-2 were clearly cited not because of their factual content, but because they boil down to the author mocking the existence of industry trends, and writing them off whole (incl. this one). This is unsound, both because it doesn't actually follow by default, and because it violates the principle of suspending (prior) judgement: https://en.wikipedia.org/wiki/Suspension_of_judgment - it sets both the author and their readers up for a specific conclusion, rather than inspiring nonbias. Hopefully one does not need to explain why this is problematic?
It is further incredibly trite to bring up how AI is an ongoing trend (self-evident), and the fact that trends distort perception (also self-evident). Them fighting fire with fire and bringing their own gut instinct is not any more intellectually respectable.
Point 3 is completely unsupported and misleading. The author probably means that hallucinations are formally unavoidable, but that on its own doesn't carry much weight, and is not the same thing.
Point 4 is factually wrong. AI is an entire academic field with a half a century of history to its name, and this is trivial to learn. Clearly not why it was cited once again either however, but because it's blatantly ill faith too.