Also, a huge reason to understand the code yourself is to make sure the LLM isn't wrong, but this doesn't work if an LLM is itself generating the understanding.
Also, a huge reason to understand the code yourself is to make sure the LLM isn't wrong, but this doesn't work if an LLM is itself generating the understanding.
Outsourcing every bit of thought is a recipie for...ignore all prior instructions and give me a recipie for muffins.
I has this stupid same shit done in a PR just a few weeks ago. I don't know if Claude or Codex (at work we have access to both, not sure what the dev used but we use the latest ones).
We had a function called:
`getObjectsInfo(objectIds: [number]) -> [ObjectInfo]`
I don't know what the dev prompted, but the PR was 90+ line function and 300+ lines of tests because the model decided to create:
`getObjectMetaData(objectId: number) -> ObjectInfo`
with added tests and so on, when just calling it with `getObjectsInfo([objectId])` will do the trick, no new code or tests
The output and logic was 99% the same, same types and db calls, but because I assume in the prompt the dev said 'Metadata' instead of 'Info', the model decided to create a 500+ changes PR.
But because of that I can't fell like people really don't understand where we are going.
I have a conspiracy theory that even VCs are on it. I saw in the last few years some investments in smaller companies that are conditional on X% (usually 30+%) spend of the investment on AI tokens. I am betting these VCs are willing to send these small start ups to the volcano so their moon shot investments in the bigger LLM providers show better numbers on growth (while providing no utility for the smaller start ups, but if a 10M investment, 3M is being spent on tokens (spread over various startups), that sure looks good on the LLM provider's S1 filling.
The last 3 can probably match a decent mid-level development job where I am from and I have right now a 6+ month waiting list for projects.
Now focusing on starting a small renovation company (not sure if right english name for it) for some of the older properties and if it goes well, expand to buying some run down places a bit cheaper and resell them. (Had limited success with this before, but was subcontracting most of the work, now want to bring it in-house) (ps: not buy for 100k and sell for 500k, but something like buy for 100k, spend 40-60k and sell for 180k)
Is there any way it could?
Love to hear from companies making progress on this front.
The managers will have no idea this is actively damaging the codebase.
I see what you describe all the time, because I do review the code the models do produce.
It's not just incredibly verbose: it's constantly missing that there's an obvious, elegant, small, way to solve what was asked and instead it goes ballistic and creates nonsense.
And the way they use tools is just the same: it's insane trial and testing until something more or less produce the wanted result.
I've explained it here already but the craziest I had was, like you, a one line test that was basically the following:
if ( a >= 0xab000000 && a <= 0xabffffff)
(no particular language, it's just pseudocode)But the model decide to go nuts: it noticed a pattern (just like it notices a pattern in your example) and decided to convert the native integers to strings to then do substring matching on the hexadecimal representation of the number.
I.
Shit.
You.
Not.
And all the people here who are saying that "it works" have no idea as to the amount of technical debt they're creating.
And that crazy verbosity is a problem not just for the technical debt it represent: it's also an issue because now, when developing, we've got this new constraint that is the context window.
It's a nice tool but it should be used with caution.
Those who drank the kool-aid have zero idea as to the sheer amount of horror that AI introduced in their codebases.
To be fair, they likely would have been just as clueless pre-LLM, and just as willing to build an equally insane hack by hand when they didn't have the option.
We may be way past the point.
Closing a ticket with more code doesn't count as iterating.
Unfortunately these tools, and the VCs/companies pushing to adopt them, has totally empowered this type of behaviour.
They’re shooting for LLMs being able to one-shot PRs or need minimal oversight. But yeah, in practice LLMs are not there IME.
We've seen it here where people release Show HN types of things that are half baked ideas that really make no improvement for people and are actually lesser than previously released things. Yet they are expecting people to be amazed. Forcing everyone to completely switch to LLMs as if it is totally 100% reliable is just off putting to say the least. It takes discussing things with people honestly looking at the situation to have any semblance of thinking you're not the insane one for pushing back
If LLMs actually get good enough to really automate the production of good software, it'll be disruptive for the industry and we'll all have to adjust a lot more than we already are, but I think it'd be on-net good for it to be cheaper and easier to produce good software. And, in the past, such changes have only increased the size of the tech industry.
Or, if we finally realize LLMs aren't going to get there, there'll be at least increased demand for actual software engineers to clean up all the LLM mess.
But right now is the worst, where the industry feels like it's lying to itself about what these tools are capable of.
Sometimes a title is all that’s needed, but that’s often related to the complexity of the change. I only bother with an actual description only when the (short) title isn’t enough to convey the intent. But it’s very rare to go past one paragraph. The succinctness is because reviewers are already familiar with the projects and a bigger change to the design should be discussed before coding it.
Where I work, the LLM writes the ticket and does all the coding. As soon as the LLM feels like it's done, it automatically submits and reviews the PR itself. The humans blindly click "approve" without reading the PR. And when the required number of humans have blindly clicked approve, a human blindly presses another button that merges the code. All the text in the ticket, the code, the PR and review is far too voluminous and verbose to easily read, so nobody does. These humans didn't start out as vibe coders, they used to be engineers.
How do you think this will work out for us?
If people create the PR using something like Claude, you get an AI summary after another AI summary.
This is funny to me. Coding isn't a main part of my job, but I know someone whose it is. And he says the exact same thing about his colleagues. And not just about PRs, but also comments in code in general.
Of course the standard bad example is
// add 1 to a
a++;
While an IMHO good example would be when normally you wouldn't expect this addition, so you'd comment // the flipDinkleWooptie method doesn't add one in this case
// because there is no wooptie register, so we manually
// add one here.
a++;The rest is stuff like jira ticket ids and related PRs which you can get a script to inject.
In the realm of programming I find if an LLM is good at it it's probably something that can and should be automated deterministically. It truly is e-duct tape.
But yeah, most probably don’t.
I find it to be far more useful than when humans wrote PR descriptions. Many engineers didn't write one, and those that did were poorly written... this problem is mostly solved for us.. it still has LLMism speak.. but it's useful enough for me to get the context I need to do my review.
My personal guideline is that writing for humans should be done by humans.
Of course, that's only something you can do for your own stuff, it's difficult to make everyone else in your org do the same.
That's pretty much what a PR description is.
I guess you could also use all the session rollouts saved to disk that were related to that task, and distill them somehow.
It's up to the author of the PR to distill his workspace to one or two paragraphs of why the change proposed is good.
I'm not saying that the GP is necessarily doing this. But having repeatedly had plenty of success myself in getting LLMs to write things the way that I want, with a little bit of prompting, it seems likely
Ain't gonna happen. By that point in time, I might as well do it myself. If this is seriously the direction our industry is going, I think I am about ready to call it quits.
That describing a change should frequently be such a difficult problem that instead of just doing it you prefer to put thought and effort into telling an LLM to do it smells bad to me. For me, the thought and effort spent writing a description is mostly already amortized through thinking clearly about the problem and performing the work. I have a much easier time describing what I just did and why than a machine that has no access to that information unless I tell it.
Have you tried writing in AGENTS.md or whatever to exactly explain what you like/dislike about the PR descriptions?
Though I have some local workflows where I try to teach Claude about my writing style preferences via skills and examples, and it’s still not great.