Along with the total absence of long-term data, I think the benefit can be (weakly) denied. Maybe not in the employmemt marketplace, but certainly for myself.
Along with the total absence of long-term data, I think the benefit can be (weakly) denied. Maybe not in the employmemt marketplace, but certainly for myself.
I think there are two different claims here:
- developers overestimate productivity gains, which is a solid finding in many of these studies. Skepticism of extremely large productivity gains is warranted and I flatly disbelieve "10x uplift" claims.
- LLMs give no productivity uplift at all, which is much harder to defend. A repeat of the famous METR RCT study did find evidence of improved productivity, and this seems to align with the experience of many experts I trust.
IMO the bigger problem is that ~1.5x individual dev productivity uplift seems to translate into 1.05x uplift across the team. People have been waaaaayyyyy too overconfident about this stuff.
There was an old Onion article about the monetary savings from buying clothes on sale as computed by (a) the women buying the clothes; and (b) their male relatives.
I can context switch between two or three chats, but doing so speeds up my agent use at the cost of making reviews and discovery harder, so it might come out in the wash.
I'm finding this is a skill which I'm slowwwwwwly improving but, for now, my level means I tend to miss mistakes I'd notice if just working on one task. I also had to make accommodations to the way I work to make it achievable, including telling my employer that a 500 GiB disk just isn't enough any more when I need multiple builds in parallel.
AI definitely supercharges the side project myth, but I think we’re going to see waves of developers quitting or losing their jobs because they’re chasing dreams of running a side code project that isn’t going to turn into a business.
The incentive to be more productive at work is that they get to keep their job and not be replaced by a cheaper junior. I think we’re in for a reckoning across the industry as companies realize that there’s little difference between a lazy senior armed with Claude and a halfway motivated junior who has aspirations of growing into something more. The latter costs less and might grow into a better dev.
The productivity uplift measured in some of these reports is from more tasks being done, not from developers giving themselves more idle time during the day.
If we do see job losses from AI, my feeling is that it’s going to be concentrated in people with an attitude like this one where workers try to remain anchored to their old delivery pace. It’s becoming easier than ever to replace a worker in that position with a lower paid junior equipped with Claude who has an interest in learning. There’s no reason to keep a higher paid worker around to copy and paste between Jira and Claude.
I just compressed what would take a team a few months to deliver in one or two months, the amount of gotcha I spot from the existing system we work on everyday is staggering.
How to deliver and keep these feedback is difficult, the other teams are overwhelmed by other tasks, so any tickets or communications will be forgotten in a few months.
And as an external ressources allowed to use frontier model, I have a hard time to keep up with my own pace, and all that institutional knowledge will be lost both by my employer and client.
- chat box interaction
- vendor harness user (the cursor / antigravity crowd)
- SDD Claude users
- bespoke harness creators who let it goooooooo
I mean we already know who’s using all the tokens and driving enough cost to make penny pinchers consider removing other “costs”
It’s not really hard to defend. Because when people says that productivity is uplifted, they are talking about amount of work, not the ROI. That’s why you keep hearing about LOC, amount of PR and prototypes, and the time taken is actually “time to PR” and not “time to production + time spent on bugs”.
This is how LLMs can both result in greater productivity yet still not appear to yield much more benefit than the pre-LLM era.
And there is no benefit to the developer for doing the work first then just sitting idle. That’s how you get people putting more things on your plate.
My gain is that I produce good work in a normal workday and I am absolutely not writing code to 4 am anymore...
Overall for me it is a win for me and my company.
Managers will abuse you if you expand your workday to make sure deadlines don't slip. They will see that the work is getting done and see no urgency to hire more help.
You have to surface capacity issues to management, learn to say "no" diplomatically, and/or "I have capacity for 2 things out of these 5, prioritize which ones you want first."
If I run out of tokens I’m encouraged to spend money to buy more tokens, Claude has an overage setup, and Anthropic runs occasional token sales.
So, to that example: the more often customers are hitting token limits under duress the more likely to open up the wallet and pay to finish.
Run out of your fix? (Tap tap) Buy some more!
Also importantly, can it be a neovim? Neovim hasn't had literal trillions invested over it in the space of just 5 years, local llm's aside can those smaller productivity gains justify the huge investment put into them, and which will continue to be needed for further model development?
A plane literally goes 10x faster than a car, so a 5 hour drive becomes a 30 minutes flight. But if you have to drive to the airport, arrive early, pass check-in, security, boarding, then pick up your luggage on arrival, rent a car and drive to your destination, you may realize that your 30 minute flight took you more than 5 hours in total.
You also have to consider the Amdahl's law: a 10x speedup on 10% of the project is just a 9% speedup overall.
This is totally true for me. I'm working on more features/bugs at work and getting the coding done faster with claude but the coding part is at maximum half of the process and usually less. The rest of it is clarifying the requirements from our product team or given that we're a very legacy place, what the existing state of our infrastructure and other systems are for the new changes to fit in to, etc.
It's a 10% speedup, unless you think that a 100% speedup would reduce completion time to zero rather than cutting it in half.
Assuming that "10% of the project" is measured by time requirements, when you get a 10x speedup on that part of the project, you'll end up completing 100 units of work in 91 (formerly 100) units of time. Your work rate is then 100 / 91 = 1.099 times what it was before; that's an improvement of 10%, not 9%.
(Sanity check: suppose you have a project that will take 100 days. That project is reassigned to someone 10% faster than you. They will take 100 / 1.1 = 90.91 days to finish.
If they were 9% faster, they'd take 100 / 1.09 = 91.74 days to finish.
Does a savings of 9 out of 100 days of work look more like the project that saves 9.09 days, or the one that saves 8.26 days?)
But instead of productivity, I'm much more interested in using it to improve quality. You've got a tireless reviewer who is always ready to review your code and catch any gaps.
The latest was that we need to write in object pooling in their client code, an option that doesn't exist yet in _their_ client code. Like, yo, that is a you-problem on performance bottlenecks; we expect you to provide solutions.
When a person becomes a manager, they do or do not have enough time and expertise to review all of the code that they trust the team to produce.
Managers usually get into automated testing; unit tests, integration tests, acceptance tests, and maybe also BDD syntax
Managers and developers are responsible for setting a test coverage threshold for merge approval.
If there is 100% branch coverage test coverage for a codebase, what would coverage-guided fuzzing or property testing find? If there is 100% branch coverage test coverage for a codebase, what is the value of spending resources on formal verification?
How does the value of LLM-produced 100% branch coverage compare to no-LLM 100% branch coverage?
This is such a salient question. Sometimes (definitely not always) the test suites produced by LLMs are so trivial it's scary. Coverage can be an illusion for sure.
I wrote a tool called tert - I guess it's called an agent harness now - to run various test runners and log test output and coverage output to disk. FWIU stripping spaces from JSON does save tokens. It seems like feeding coverage lines-missing maps into the prompt results in better output, better LLM-authored tests.
"Refactor these tests for maintainability and coverage. Use fixtures, mocks, and parametrization"
Substance coverage - testing the actual logic, edge cases, etc. Not mere lines.
If there is prompt insufficiency, there is probably acceptance test insufficiency.
A more assuming agent could automatically develop a plan that includes presumptive acceptance tests and request feedback before spending tokens
And, if/where we need tests, we write the source so they are few, high value, and complementary. Like actual unit tests, not complex with stuff like mocks just to generate trivial coverage.
Working with an LLM has given me a real eye opener on unwritten requirements. It's like outsourcing. "Yes, you've given me what I wrote down, but I never expected you do to it in that way"
I haven't yet made myself learn the new swarm of concurrent agents with different specializations/agent_instructions methods yet.
Are multiple worktrees worth the cognitive burden and merge overhead?
A merge maintainer is always in code review mode
> Managers usually get into automated testing; unit tests, integration tests, acceptance tests, and maybe also BDD syntax
I can see managers getting involved into acceptance tests, but never in the other type of tests. And the verification mostly is involved into a quick manual testing/watching a demo. Code is not their concern. When there's a bug, they expect you to investigate and fix it.
And then sometimes you report a memory leak and it gets fixed by a VP and you wonder if he doesn't have something better to do.
If devops has done their job, it should be trivial for a manager to contribute to the tests and run the build on git push (or manually re-run the build with the web UI).
If a manager has deploy rights, they should be able to run the tests.
If you have a "I trust my competent team to write good enough tests and test coverage isn't my responsibility" attitude, that's what quality software you'll get back.
There are people producing good and excellent quality software with LLMs. Presumably you must discard low-quality code in order to maintain quality.
There's certainly a limit to code quality with current models. On number of lines of code per unit of time, LLM tools certainly already win.
Can costly automated code review for PRs catch most of the problems before they're under consideration for merge?
For example, the vscode repo has extensive copilot integration. Every PR gets auto code reviewed. But with their tokens or the contributors'?
If I take poor quality code (AI-assisted or not) and spend a few hundred dollars on tokens for a next gen model and agent to get to 100% coverage and review for security bugs and CWE common weaknesses, what quality code will I have without refactoring with proven patterns and type annotations and polishing docstrings?
If you're a manager now and your "team" is a bunch of coding agents, those agents are hardly junior engineers at best. It is equivalent or even worse than hiring a team of 3-5 junior engineers and letting them run rampant with your code.
As much as this sentiment is nice, it is completely divorced from reality, unless the competence is verifiably there. If you take a bunch of juniors and say "yeah I trust them to do everything well enough," you're going to have a disaster on your hands.
If slop is fine (and sometimes it is), the benefits are undeniable. If the dev was the kind that would have produced slop anyway - again, undeniable boost.
If the quality needs to be high I think it actually can slow you down, though.
The result is a whole bunch of dysfunctional systems unnecessarily dislodging perfectly acceptable processes.
So the mediocre-dev case may be worse than "10x more mediocre code." It's more code that also skews insecure by default, and that cost shows up downstream in review and incidents, not at the PR. Throughput goes up, and so does the liability per line.