Has your productivity objectively, measurably improved or does it just feel like it has improved? Recall the METR study which caught programmers self-reporting they were 20% faster with AI when they were actually 20% slower.
Has your productivity objectively, measurably improved or does it just feel like it has improved? Recall the METR study which caught programmers self-reporting they were 20% faster with AI when they were actually 20% slower.
If I could have something that said, "Here are some things that it looks like you're procrastinating on -- do you want me to get started on them for you?" -- that would probably be crazy useful.
The privacy implications are horrifying. But if done right, you’re taking about a kind of digital ‘executive function’ that could help a lot of kids that struggle with things like prioritization and time blindness.
I was diagnosed with ADHD and my interpretation of that diagnoses was not "I need something to take over this functionality for me," but "I need to develop this functionality so that I can function as a better version of myself or to fight against a system which is not oriented towards human dignity but some other end."
I guess I am reluctant to replace the unique faculties of individual children with a generic faculty approved by and concordant with the requirements of the larger society. How dismal to replace the unique aspects of children's minds with a cookie cutter prosthetic meant to integrate nicely into our bullshit hell world. Very dismal.
As someone with ADHD, I say: Please don't build this.
Open source transcription models are already good enough to do this, and with good context engineering, the base models might be good enough, too.
It wouldn't be trivial to implement, but I think it's possible already.
> If you were expecting iOS 19 after iOS 18, you might be a little surprised to see Apple jump to iOS 26, but the new number reflects the 2025-2026 release season for the software update.
For many people, it's easier to improve a bad first version of a piece of writing than to start from scratch. Even current mediocre LLM are great at writing bad first drafts.
Anyone is great at creating a bad first draft. You don’t need help to create something bad, that’s why that’s a common tip. Dan Harmon is constantly hammering on that advice for writer’s block: “prove you’re a bad writer”.
https://www.youtube.com/watch?v=BVqYUaUO1cQ
If you get an LLM to write a first draft for you, it’ll be full of ideas which aren’t yours which will condition your writing.
Famously not so! Writer's block is real!
Getting an LLM to produce a not-so-bad first draft is just another technique.
https://archive.ph/20250924025805/https://www.nytimes.com/20...
I am seeing a pattern here. It appears that AI isn't for everyone. Not everyone's personality may be a good fit for using AI. Just like not everybody is a good candidate for being a software dev, or police officer etc.
I used to think that it is a tool. Like a car is. Everybody would want one. But that appears not be the case.
For me, I used AI every day as a tool, for work and and home tasks. It is a massive help for me.
It's hard for me to imagine many. It's not doing the dishes or watering the plants.
If I wanted to rearrange the room I could have it mock up some images, I guess...
How can you verify the recommendations are sound, valid, safe, complete, etc., without trying them out? And trying out unsound, invalid, unsafe, incomplete, etc., recommendations might result in dead plants in a couple of weeks.
Such an odd complaint about LLMs. Did people just blindly trust Google searches before hand?
If it's something important, you verify it the same way you did anything else. Check the sources and use more than a single query. I have found the various LLMs to very useful in these cases, especially when I'm coming at something brand new and have no idea what to even search for.
I've found it immensely helpful for giving real world recommendations about things like this, that I know how to find on my own but don't have the time to do all the reading and synthesizing.
Use only actionable prompts, negations don't work on ai and they don't work on people.
Hell, in the past few days I started making something to help me write documents for work (https://www.writelucid.cc) and a viewer for all my blood tests (https://github.com/skorokithakis/bt-viewer), and I don't think I would have made either without an LLM.
Would have never done that without LLMs.
I'm tackling projects solo I never would have even attempted before but I could see people getting bad results and giving up.
This is a truism and, I believe, is at the core of the disagreement on how useful AI tools are. Some people keep talking about outlier success. Other people are unimpressed with the performance in ordinary tasks, which seem to take longer because of back-and-forth prompting.
IOW, can you redo it by yourself? If you can't then you did not learn it.
Knowing the abstract steps and tripwires yes, but details will always have to be looked up. If just not to miss any new developments.
Well, yes it is; you can't very well claim to have learned something if you are unable to do it.
I get that point, but the original post I replied to didn't say "Hey, I know have $THING set up when I never had it before", he said "I learned to do $THING", which is a whole different assertion.
I'm not contending the assertion that he now has a thing he did not have before, I'm contending the assertion that he has learned something.
One thing I've noticed is that I don't have a circle of people where I can discus programming with, and having an LLM to answer questions and wireframe up code has been amazing.
My job doesn't require programming, but programming makes my job much easier, and the benefits have been great.
I want to second your experience as I’ve had the same as well. Tackling SO many more tasks than before and at such a crazy pace. I’ve started entire businesses I wouldn’t have just because of AI.
But at the same time, some people have weird blockers and just can’t use AI. I don’t know what it is about it - maybe it’s a mental block? Wrong frame of mind? It’s those same people who end up saying “I spend more time fighting the ai and refining prompts than I would on the end task”.
I’m very curious what it is that actually causes this divide.
I've been using it for almost a year now, and it's definitely improved my productivity. I've reduced work that normally takes a few hours to 20 minutes. Where I work, my manager was going to hire a junior developer and ended up getting a pro subscription to Claude instead.
I also think it will be a concern for that 50-something developer that gets laid off in the coming years, has no experience with AI, and then can't find a job because it's a requirement.
My cousin was a 53 year old developer and got laid off two years ago. He looked for a job for 6 months and then ended up becoming an auto mechanic at half the salary, when his unemployment ran out.
The problem is that he was the subject matter expert on old technology and virtually nobody uses it anymore.
Ok, so subjective
Task: Walk to the shops & buy some milk.
Deliverables: 1. Video of walking to the shops (including capturing the newspaper for that day at the local shop) 2. Reciept from local store for milk. 3. Physical bottle of Milk.
milk (noun):
1. A whitish liquid containing proteins, fats, lactose, and various vitamins and minerals that is produced by the mammary glands of all mature female mammals after they have given birth and serves as nourishment for their young.
2. The milk of cows, goats, or other animals, used as food by humans.
3. Any of various potable liquids resembling milk, such as coconut milk or soymilk.
You get what you asked for, or you didn't sufficiently define it.
There's nothing worse than a task where you can deliver one item and then have to rely on someone else to be able to deliver a second. Was once in a role where performance was judged on closing tasks; getting the burn-down chart to 0, and also having it nicely stepped. Was given a good tip to make sure each task had one deliverable and where possible—be completed independent of any other task.
Why would you write down "Buy Milk", then go buy whatever thing you call milk, then come back home and be confused about it?
Only an imbecile would get stuck in such a thing.
Could you give some examples, and an indication of your level of experience in the domains?
The statement has a much different meaning if you were a junior developer 2 years ago versus a staff engineer.
I've been adding small features in a language I don't program in using libraries I'm not familiar with thhat meet my modest functional requirements in a couple minutes each. I work with an LLM to refine my prompt, put it into cursor, run the app locally, look at the diffs, commit, push and I'm live on vercel within a minute or two.
I don't have any good metrics for productivity, so I'm 100% subjective but I can say that even if I'd been building in Rails (it's been ~4 years but I coded in it for a decade) it would have taken me at least 8 hours to have an app where I was happy with both the functionality and the look and feel so a 10x improvement in productivity for that task feels about right.
And having a "buddy" I can discuss a project with makes activation energy lower allowing me to complete more.
Also, YC videos I don't have the time to watch, I get a transcript, feed into chatGTP, ask for the key take aways I could apply to my business (it's in a project where it has context on stage, industry, maturity, business goals, key challenges, etc) so I get the benefits of 90 minutes of listening plus maybe 15 minutes of summarizing, reviewing and synthesis in typically 5-6 minutes - and it'd be quicker if I built a pipeline (something I'm vibe coding next month)
Wouldn't want to do business without it.
And have another AI review your unit tests and code. It's pretty amazing how much nuance they pick up. And just rinse and repeat until the AI can't find anything anymore (or you notice it going in circles with suggestions)
Are the frontend folks having such great results from LLMs that they're OK with "just let the LLM check for security too" for non-frontend-engineer created projects that get hosted publicly?
This. 100x this.
Personally, I think it really shines at doing the boring maintenance and tech debt work. None of these are hard or complex tasks but they all take up time and for a buck or two in tokens I can have it doing simple but tedious things while I'm working on something else.
It shines at doing the boring maintenance and tech debt work for web. My experiences with it, as a firmware dev, have been the diametric opposite of yours. The only model I've had any luck with as an agent is Sonnet 4 in reasoning mode. At an absolutely glacial pace, it will sometimes write some almost-correct unit tests. This is only valuable because I can have it to do that while I'm in a meeting or reading emails. The only reason I use it at all is because it's coming out of my company's pocket, not mine.
If you're doing JS/Python/Ruby/Java, it's probably the best at that. But even with our stack (elixir), it's not as good as, say, React/NextJS, but it's definitely good enough to implement tons of stuff for us.
And with a handful of good CLAUDE.md or rules files that guide it in the right direction, it's almost as good as React/NextJS for us.
Random Postgres stuff:
- Showed a couple of Geo/PostGIS queries that were taking up more CPU according to our metrics, asked it to make it faster, it rewrote it in away that it actually used the index. (using the <-> operator for example for proximity). One-shotted. Whole effort was about 5 mins.
- Regularly asking for maintenance scripts (like give me a script that shows me the most fragmented tables, or highest storage, etc).
CSS:
Built a whole horizontal logo marquee with CSS animations, I didn't write a single line, then I asked for little things like "have the people's avatars gently pulsate" – all this was done in about 15 mins. I would've normally spent 8-16 hours on all that pixel pushing.
Elixir App:
- I asked it to look at my GitHub actions file and make it go faster. In about 2-3 iterations, it cut my build time from 6 minutes to 2 minutes. The effort was about an hour (most of it spent waiting for builds, or fiddling with some weird syntax errors or just combining a couple extra steps, but I didn't have to spend a second doing all the research, its suggestions were spot on)
- In our repo (900 files) we had created an umbrella app (a certain kind of elixir app). I wanted to make it a non-umbrella. This one did require more work and me pushing it, but I've been putting off this task for 3 YEARS since it just didn't feel like a priority to spend 2-3 days on. I got it done in about 2 hours.
- Built a whole discussion board in about 6 hours.
- There are probably 3-6 tickets per week where I just say "implement FND-234", and it one-shots a bugfix, or implementation, especially if it's a well defined smaller ticket. For example, make this list sortable. (it knows to reuse my sortablejs hook and look at how we implemented it elsewhere).
- With the Appsignal MCP, I've had it summarize the top 5 errors in production, and write a bug fix for one I picked (I only did this once, the MCP is new). That one was one-shotted.
- Rust library (It's just an elixir binding to a rust library, the actual rust is like 20 lines, so not at all complex)... I've never coded a day of rust in my life, but all my cargo updates and occasional syntax/API deprecations, I have claude do my upgrades and fixes. I still don't know how to write any Rust.
NextJS App:
- I haven't fixed a single typescript error in probably 5 months now, I can't be bothered, CC gets it right about 99% of the time.
- Pasted in a Figma file and asked it to implement. This rarely is one-shotted. But it's still about 10x faster than me developing it manually.
The best combination is if you have a robust component library and well documented patterns. Then stuff goes even faster.
All on the $100 plan in which I've hit the limit only twice in two months. I think if they raised the price to $500, it would still feel like a no-brainer.
I think Anthropic knows this. My guess is that they're going to get us hooked on the productivity gains, and we will happily pay 5x more if they raised the prices, since the gains are that big.
But I’m not even going to argue about that. I want to raise something no one else seems to mention about AI in coding work. I do a lot of work now with AI that I used to code by hand, and if you told me I was 20% slower on average, I would say “that’s totally fine it’s still worth it” because the EFFORT level from my end feels so much less.
It’s like, a robot vacuum might take way longer to clean the house than if I did it by hand sure. But I don’t regret the purchase, because I have to do so much less _work_.
Coding work that I used to procrastinate about because it was tedious or painful I just breeze through now. I’m so much less burnt out week to week.
I couldn’t care less if I’m slower at a specific task, my LIFE is way better now I have AI to assist me with my coding work, and that’s super valuable no matter what the study says.
(Though I will say, I believe I have extremely good evidence that in my case I’m also more productive, averages are averages and I suspect many people are bad at using AI, but that’s an argument for another time).
The problem is, there are very few if any other studies.
All the hype around LLMs we are supposed to just believe. Any criticism is "this study has serious problems".
> It’s like, a robot vacuum might take way longer
> Coding work that I used to procrastinate
Note how your answer to "the study had serious problems" is totally problem-free analogies and personal anecdotes.
Not at all, the METR study just got a ton of attention. There are tons out there at much larger scales, almost all of them showing significant productivity boosts for various measures of "productivity".
If you stick to the standard of "Randomly controlled trials on real-world tasks" here are a few:
https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4945566 (4867 developers across 3 large companies including Microsoft, measuring closed PRs)
https://www.bis.org/publ/work1208.pdf (1219 programmers at a Chinese BigTech, measuring LoC)
https://www.youtube.com/watch?v=tbDDYKRFjhk (from Stanford, not an RCT, but the largest scale with actual commits from 100K developers across 600+ companies, and tries to account for reworking AI output. Same guys behind the "ghost engineers" story.)
If you look beyond real-world tasks and consider things like standardized tasks, there are a few more:
https://ieeexplore.ieee.org/abstract/document/11121676 (96 Google engineers, but same "enterprise grade" task rather than different tasks.)
https://aaltodoc.aalto.fi/server/api/core/bitstreams/dfab4e9... (25 professional developers across 7 tasks at a Finnish technology consultancy.)
They all find productivity boosts in the 15 - 30% range -- with a ton of nuance, of course. If you look beyond these at things like open source commits, code reviews, developer surveys etc. you'll find even more evidence of positive impacts from AI.
> https://www.youtube.com/watch?v=tbDDYKRFjhk (from Stanford, not an RCT, but the largest scale with actual commits from 100K developers across 600+ companies, and tries to account for reworking AI output. Same guys behind the "ghost engineers" story.)
I like this one a lot, though I just skimmed through it. At 11:58 they talk about what many find correlates with their personal experience. It talks about easy vs complex in greenfield vs brownfield.
> They all find productivity boosts in the 15 - 30% range -- with a ton of nuance, of course.
Or 5-30% with "Ai is likely to reduce productivity in high complexity tasks" ;) But yeah, a ton nuance is needed
However even these levels are surprising to me. One of my common refrains is that harnessing AI effectively has a deceptively steep learning curve, and often individuals need to figure out for themselves what works best for them and their current project. Took me many months, personally.
Yet many of these studies show immediate boosts in productivity, hinting that even novice AI users are seeing significant improvements. Many of the engineers involved didn't even get additional training, so it's likely a lot of them simply used the autocompletion features and never even touched the powerful chat-based features. Furthermore, current workflows, codebases and tools are not suited for this new modality.
As things are figured out and adopted, I expect we'll see even more gains.
With ai code you have more loc and NEED more PRs to fix all its slop.
In the end you have increased numbers with net negative effect
> https://www.youtube.com/watch?v=tbDDYKRFjhk (from Stanford, not an RCT, but the largest scale with actual commits from 100K developers across 600+ companies, and tries to account for reworking AI output. Same guys behind the "ghost engineers" story.)
Emphasis added. They modeled a way to detect when AI output is being reworked, and still find a 15-20% increase in throughput. Specific timestamp: https://youtu.be/tbDDYKRFjhk?t=590&si=63qBzP6jc7OLtGyk
I did say I wasn’t going to argue the point that study made, and I didn’t.
In your particular case it sounds like you’re rapidly loosing your developer skills, and enjoy that now you have to put less effort and think less.
Pretty much like muscles decay when we stop using them.
>> Pretty much like muscles decay when we stop using them.
> Sure, but sticking with that analogy, bicycles haven’t caused the muscles of people that used to go for walks and runs to atrophy either ...
This is an invalid continuation of the analogy, as bicycling involves the same muscles used for walking. A better analogy to describe the effect of no longer using learned skills could be:
Asking Amazon's Alexa to play videos of people
bicycling the Tour de France[0] and then walking
from the couch to the your car every workday
does not equate to being able to participate in
the Tour de France[0], even if years ago you
once did.
0 - https://www.letour.fr/en/Then the citation served its purpose.
You're welcome.
Take my personal experience for whatever it is worth, but my knees do not lie.
I believe the same holds true for cognitive tasks. If you enjoy going through weird build file errors, or it feels like it helps you understand the build system better, by all means, go ahead!
I just don't like the idea of somehow branding it as a moral failing to outsource these things to an LLM.
But if you were to be literally chained to a bike, and could not move in any other way than surely you would "forget"/atrophy in specific ways that you wouldn't be able to walk without relearning/practicing.
A similar phenomena occurs when people see or hear information and whether they record it in writing or not. The act of writing the percepts, in and of itself, assists in short-term to long-term memory transference.
Same with LLMs. I am better with it, because I know how to solve things without the help of it. I understand the problem space and the limitations. Also I understand how hype works and why they think they need it (investors money).
In other words, no, just using google maps or ChatGPT does not make me dumb. Only using it and blindly trusting it would.
Applying all of this to LLMs has felt similar.
Consultancy A submit work, Consultancy B reviews/tests. As A increases the use of AI, B will have to match with more staff or more AI. More staff for B, mean higher costs, at slower pace. More AI for B, means higher burden of proof, an A vs B race condition is likely.
Ultimately clients will suffer from AI fatigue and inadvertently incur more costs at later stage (post-delivery).
It’s the same story with UI/UX. Previously, I’d often have to skip little UI niceties because they take time and aren’t that important. Now even relatively minor user flows can be very well polished because there isn’t much cost to doing so.
Well your perfectionism needs to be pointed towards this line. If you get truly large numbers of users this will either slow down token checking directly or your process for removing ancient expired tokens (I'm assuming there is such a process...) much slower and more problematic.
Just the other day I was complaining that no one knows how to use a slide rule anymore...
Also C++ is producing bytecode that's hot garbage. It's like no one understands assembly anymore...
Even simple tools are often misused (like hammering a screw). Sometimes they are extremely useful in right hands though. I think we'll discover that the actual writing of code isn't as meaningful as thinking about code.
Imagine telling someone with a typewriter that they’d be unable to write if they don’t write by hand all the time lol. I write by hand maybe a few times a year - usually writing a birthday card or something - but I haven’t forgotten.
That stuff kills my motivation to solve actual problems like nothing else. Being able to send off an agent to e.g. fix some build script bug so that I can get to the actual problem is amazing even with only a 50% success rate.
Otherwise, I’ll continue using what works for me now.
I feel like the past few decades of framework churn has shown that we're really never going to agree on what this means
I completely get this and I often have an LLM do boring stupid crap that I just don't wanna do. I frequently find myself thinking "wow I could've done it by hand faster." But I would've burned some energy that could be better put towards other stuff.
I don't know if that's a net positive, though.
On one hand, my being lazy may be less of a hindrance compared to someone willing to grind more boring crap for longer.
On the other hand, will it lessen my edge in more complicated or intricate stuff that keeps the boring-crap-grinders from being able to take my job?
I know it's not some amazing GDP-improving miracle, but in my personal life it's been incredibly rewarding.
I had a dozen domains and projects on the shelf for years and now 8 of them have significant active development. I've already deployed 2 sites to production. My github activity is lighting up like a Christmas tree.
on the other hand i believe my coworker may have taken it too far. it seems like productivity has significantly slipped. in my perception the approaches hes using are convoluted and have no useful outcome. im almost worried about him because his descriptions of what hes doing make no sense to me or my teammates. hes spending a lot of time on it. im considering telling him to chill out but who knows, maybe im just not as advanced a user as him? anyone have experience with this?
it started as an approach to a mass legacy code migration. sound idea with potential to save time. i followed along and understood his markdown and agent stuff for analyzing and porting legacy code
i reviewed results which apply to my projects. results were mixed bag but i think it saved some time overall. but now i dont get where hes going with his ai aspirations
my best attempt to understand is he wants to work entirely though chats, no writing code, and hes doing so by improving agents through chats. hes really swept up in the entire concept. i consider myself optimistic about ai but his enthusiasm feels misplaced
its to the point where his work is slipping and management is asking him where his results are. were a small team and management isnt savvy enough to see hes getting NOTHING done and i wont sell him out. however if this is a known delusional pattern id like to address it and point to a definition and/or past cases so he can recognize the pattern and avoid trouble
but I do recall seeing some Amazon engineer who worked on Amazon q and his repos and they were... something.
like making PRs that were him telling the ai that "we are going to utilize the x principle by z for this" and like 100s of lines of "principles" and stuff that obviously would just pollute the context and etc.
like huge amounts of commits but it was just all this and him trying to basically get magic working or something.
and to someone like me it was obvious that this was a futile effort but clearly he didn't seem to quite get it.
I think the problem is that people don't understand transformers, that they're basically huge datasets in a model form where it'll auto-generated based on queries from the context (your prompts and the models reponses)
so you basically are just getting mimicked responses
which can be helpful but I have this feeling that there's a fundamental limit, like a mathematical one where you can't get it really to do stuff unless you provide the solution itself in your prompt, that covers everything because otherwise it'd have to be in its training data (which it may have, for common stuff like boilerplate, hello world etc.)
but maybe I'm just missing something. maybe I don't get it
but I guess if you really wanna help him, I'd maybe play around with claude/gpt and see how it just plays along even if you pretend, like you're going along with a really stupid plan or something and how it'll just string you along
and then you could show him.
Orr.... you could ask management to buy more AI tools and make him head of AI and transition to being an AI-native company..
you put it nicely when you mention a fundamental limit and will borrow that if i think hes wasting a risky amount of time
i really like the sibling idea of having him try to explain again, then use claude to explain if he cant
genuine thanks to you and sibling for offering advice
The right AI, good patterns in the codebase and 20 years of experience and it is wild how productive I can be.
Compare that to a few years ago, when at the end of the week, it was the opposite.
I've seen a lot of people who previously touted that it doesn't work at all use that study as a way to move the goalpost and pretend they've been right all along.
I just recently had to rate whether I felt like I got more done by leaning more on Claude Code for a week to do a toy project and while I _feel_ like I was more productive, I was already biased to think so, and so I was a lot more careful with my answer, especially as I had to spend a considerable amount of time either reworking the generated code or throwing away several hours of work because it simply made things up.
There's a reason self-reported measures are questioned: they have been wildly off in different domains. Objectively verifying that a car is faster than walking is easy. When it's not easy to objectively prove something, then there are a lot that could go wrong, including the disagreements on the definition of what's being measured.
Again, people who were already highly productive without AI won't understand how profound the increase is.
If I showed them time gains, they’d just say “well you don’t know how much tech debt you’re creating”, they’d find a weasel way to ignore any methodology we used.
If they didn’t, they wouldn’t be conveniently ignoring all but that one study that is skeptical of productivity gains.
So - this thing would never be in existance and work without a 20 USD ClaudeAI subscription :)
I would ask, then, if you're qualified to evaluate that what 'you' are doing now is what you think it is? Writing off 'does it lead to other problems' with 'no doubt, but' feels like something to watch closely.
I imagine a would-be novelist who can't write a line. They've got some general notions they want to be brilliant at, but they're nowhere. Apply AI, and now there's ten million words, a series, after their continual prompts. Are they a novelist, or have they wasted a lot of time and energy cosplaying a novelist? Is their work a communication, or is it more like an inbox full of spam into which they're reading great significance because they want to believe?
You can currently go to websites and use character generators and plot idea generators to get unstuck from writers block or provide inspiration and professional writers already do this _all the time_.
I have always been a careful tester, so my UAT hasn't blown up out of proportion.
The big issue I see is rust it generates code using 2023-recent conventions, though I understand there is some improvement in thst direction.
Our hiring pipeline is changing dramatically as well, since the normal things a junior needs to know (code, syntax) is no longer as expensive. Joel Spolsky's mantra to higher curious people who get things done captures well the folks I find are growing well as juniors.
AI has not made me much more productive at work.
I can only work on my hobby project when I’m tired after the kids go to bed. AI has made me 3x productive there because reviewing code is easier than architecting. I can sense if it’s bad, I have good tests, the requests are pretty manageable (make a new crud page for this DTO using app conventions).
But at work where I’m fresh and tackling hard problems that are 50% business political will? If anything it slows me down.
Interesting to consider that if our first vibecode prompt isn't what we actually want; it can train on how we direct it further.
Offloading human intelligence is useful but... we're losing something.
As with many other technologies, AI can be an enabler of this, or it can be used as a tool to empower and enhance learning and personal growth. That ultimately depends on the human to decide. One can dramatically accelerate personal and professional growth using these tools.
Admittedly the degree to which one can offload tasks is greatly increased with this iteration, to the extent that at times you can almost seem like offloading your own autonomy. But many people already exist in this state, exclusively parroting other people's opinions without examining them, etc.
https://www.fightforthehuman.com/are-developers-slowed-down-...
The study gets so much attention since it's one of the few studies on the topic with this level of rigor on real-world scenarios, and it explains why previous studies or anecdotes may have claimed perceived increases in productivity even if there wasn't any actual increases. It clearly sets a standard that we can't just ask people if they felt more productive (or they need to feel massively more productive to clearly overcome this bias).
Yes, but most people don't seem aware of those caveats, and this is a good summary of them, and I think it does undercut the "level of rigour" of the study. Additionally, some of what the article points out is not explicitly acknowledged and connected by the study itself.
For instance, if you actually split up the tasks by type, some tasks show a speed up and some show a slowdown, and the qualitative comments by developers about where they thought AI was good/bad aligned very well with which saw what results.
Or (iirc) the fact that the task timing was per task, but developer's post hoc assessments were a prediction of how much they thought they were sped up on average across all tasks, meaning it's not really comparing the same things when comparing how developers felt vs how things actually went.
Or the fact that developers were actually no less accurate in predicting times to task completion overall wrt to AI vs non-AI.
> and it explains why previous studies or anecdotes may have claimed perceived increases in productivity even if there wasn't any actual increases.
Framing it that way assumes as an already established fact that needs to be explained that AI does not provide more productivity Which actually demonstrates, inadvertently, why the study is so popular! People want it to be true, so even if the study is so chock full of caveats that it can't really prove that fact let alone explain it, people appeal to it anyway.
> It clearly sets a standard that we can't just ask people if they felt more productive
Like we do for literally every other technological tool we use in software?
> (or they need to feel massively more productive to clearly overcome this bias).
All of this assumes a definition of productivity that's based on time per work unit done, instead of perhaps the amount of effort required to get a unit of work done, or the extra time for testing, documentation, shoring up edge cases, polishing features, that better tools allow. Or the ability to overcome dread and procrastination that comes from dealing with rote, boilerplate tasks. AI makes me so much more productive that friends and my wife have commented on it explicitly without needing to be prompted, for a lot of reasons.
The performance gains come from being able to ask specific questions about problems I deal with and (basically) have a staff engineer that I can bounce ideas off of.
I am way faster at writing tasks on problems I am familiar with vs an AI.
But me trying to figure out the database I should deeply look at for my usecase or debug android code when I don't know kotlin has saved me 5000x time.
Instead of getting overwhelmed doing to many things, I can offload a lot of menial and time-driven tasks
Reviews are absolutely necessary but take less time than creation
From the one random file I opened:
/// Real LSP server implementation for Lens pub struct LensLspServer
/// Configuration for the LSP server
pub struct LspServerConfig
/// Convert search results to LSP locations
async fn search_results_to_locations()
/// Perform search based on workspace symbol request
async fn search_workspace_symbols()
/// Search for text in workspace
async fn search_text_in_workspace()
etc, etc, etc, x1000.
I don't see a single piece of logic actually documented with why it's doing what it's doing, or how it works, or why values are what they are, nearly 100% of the comments are just:
function-do-x() // Function that does x
I find that people who dismiss LoC out of hand without supplying better metrics tend to be low performers trying to run for cover.
Oh no, you've caught me.
On a serious note: LoC can be useful in certain cases (e.g. to estimate the complexity of a code base before you dive in, even though it's imperfect here, too). But, as other have said, it's not a good metric for the quality of a software. If anything, I would say fewer LoC is a better indication of high quality software (but again, not very useful metric).
There is no simple way to just look at the code and draw conclusions about the quality or usefulness of a piece of software. It depends on sooo many factors. Anybody who tells you otherwise is either naive or lying.
There are none. All are various variant of bad. LoC is probably the worst metric of all. Because it says nothing about quality, or features, or number of products shipped. It's also the easiest metric to game. Just write GoF-style Java, and you're off to the races. Don't forget to have a source code license at the beginning of every file. Boom. LoC.
The only metrics that barely work are:
- features delivered per unit of time. Requires an actual plan for the product, and an understanding that some features will inevitably take a long time
- number of bugs delivered per unit of time. This one is somewhat inversely correlated with LoC and features, by the way: the fewer lines of code and/or features, the fewer bugs
- number of bugs fixed per unit of time. The faster bugs are fixed the better
None of the other bullshit works.
To clarify, people critical of the “productivity increase” argument question whether the productivity is of the useful kind or of the increased useless output kind.
If nobody is watching loc, it’s generally a good metric. But as soon as people start valuing it, it becomes useless.
and, in the case of "Lines of code" as a metric: https://en.wikipedia.org/wiki/Cobra_effect
Second, as you seem to be an entrepreneur, I would suggest you consider adopting the belief that you've not been productive until the thing's shipped into prod and available for purchase. Until then you've just been active.
I'll pass on this data point.
I truly don’t know how to account for the discrepancy, I can imagine many possible explanations.
But what really gets my goat is how political this debate is becoming. To the point that the productivity-camp, of which I’m a part, is being accused of deluding themselves.
I get that OpenAI has big ethical issues. And that there’s a bubble. And that ai is damaging education. And that it may cause all sorts of economic dislocation. (I emphatically Do Not get the doomers, give me a break).
But all those things don’t negate the simple fact that for many of us, LLMs are an amazing programming tool, and we’ve been around long enough to distinguish substance from illusion. I don’t need a study to confirm what’s right in front of me.
I work with many developers of varying skill levels, all of which use AI. The only ones who have attempted to turn in slop are ones that basically turned out that they can’t code at all and didn’t keep their job long. Those who know what they’re doing, use it as a TOOL. They carefully build, modify, review and test everything and usually write about half of it themselves and it meets our strict standards.
Which you would know if you’d listened to what we’ve been telling you in good faith.