AI adoption and Solow's productivity paradox
fortune.com
fortune.com
Which is that information technology similarly (and seemingly shockingly) didn't produce any net economic gains in the 1970's or 1980's despite all the computerization. It wasn't until the mid-to-late 1990's that information technology finally started to show clear benefit to the economy overall.
The reason is that investing in IT was very expensive, there were lots of wasted efforts, and it took a long time for the benefits to outweigh the costs across the entire economy.
And so we should expect AI to look the same -- it's helping lots of people, but it's also costing an extraordinary amount of money, and the few people it's helping is currently at least outweighed by the people wasting time with it and its expense. But, we should recognize that it's very early days, and that productivity will rise with time, and costs will come down, as we learn to integrate it with best practices.
The thing to note is, verifying if something got done is harder and takes time in the same ballpark as doing the work.
If people are serious about AI productivity, lets start by addressing how we can verify program correctness quickly. Everything else is just a Ferrari between two traffic red lights.
A Claude subscription is 20 bucks per worker if using personal accounts billed to the company, which is not very far from common office tools like slack. Onboarding a worker to Claude or ChatGPT is ridiculously easy compared to teaching a 1970’s manual office worker to use an early computer.
Larger implementations like automating customer service might be more costly, but I think there are enough short term supposed benefits that something should be showing there.
but at that point you could go for a bugger one and split amongst headcount
Claude Code has rate limits for a reason: I expect they are carefully designed to ensure that the average user doesn't end up losing Anthropic money, and that even extreme heavy users don't cause big enough losses for it to be a problem.
Everything I've heard makes me believe the margins on inference are quite high. The AI labs lose money because of the R&D and training costs, not because they're giving electricity and server operational costs away for free.
I'll be convinced they're actually making money when they stop asking for $30 billion funding rounds. None of that money is free! Whoever is giving them that money wants a return on their investment, somehow.
Training costs are fixed at whatever billions of dollars per year.
If inference is profitable they might conceivably make a profit if they can build a model that's good enough to sign up vast numbers of paying customers.
If they lose even more money on each new customer they don't have any path to profitability at all.
In theory they can increase prices once the customers will be hocked up. That's how many startups works.
I mean we just have to look at old discussions about Uber for the exact same arguments. Uber, after all these years, still is at a negative 10 % lifetime ROI , and that company doesn't even have to meaningfully invest in hardware.
IMO this will probably develop like the railroad boom in the first half of the 19th century: All the AI-only first movers like OpenAI and Anthropic will go bust, just like most railroad companies who laid the tracks, because they can't escape the training treadmill. But the tech itself will stay, and even become a meaningful productivity booster over the next decades.
He often gathers good information but his analysis of that information appears to be heavily influenced by the conclusions he's already trying to reach.
I do pay attention to him but I'd like to see similar conclusions from other analysts against the same data before I treat them as robust.
I don't personally have the knowledge or experience of company finance to be able to confidently evaluate his findings myself!
Once that happens, whomever is left standing can dial back the training investment to whatever their share of inference can bear.
Or, if there's two people left standing, they may compete with each other on price rather than performance and each end up with cloud compute's margins.
Which means that training needs to be ongoing. So the revenue covers the inference? So what? All that means is that it doesn't cover your costs and you're operating at a loss. Because it doesn't cover the training that you can't stop doing either.
>Training costs are fixed at whatever billions of dollars per year.
Which I think is the part people disagree with.
Capex is probably the biggest hurdle, but I can see how electricity cost might become a factor under heavy use.
I was downvoted big time. Ah, I love it when people provide an example so it can finally be exposed without me having to say anything.
Unfortunately this is a huge problem on here - many people step outside of their domains, even if on the surface it seems simple, but post gibberish and completely mangled stuff. How does this benefit people who get exposed to crap?
People form very strong opinions on topic they barely understand. I'd say since they know little the opinions come mostly from emotions, which is hardly a good path for objective and deeper knowledge.
Seems like a pretty dumb take. It’s like saying it only takes $X in electricity and raw materials to produce a widget that I sell for $Y. Since $Y is bigger than $X, I’m making money! Just ignore that I have to pay people to work the lines. Ignore that I had to pay huge amounts to build the factory. Ignore every other cost.
They can’t just fire everyone and stop training new models.
Gross profit = revenues - cost of goods sold
Operating profit = Gross profit - operating expenses including depreciation & amortisation
Net profit = Operating profit - net interest expense - taxes
If I am on a roll, I will flip on Extra Usage. I prototyped a fully functional and useful niche app in ~6 total hours and $20 of extra usage, and it's solid enough and proved enough value to continue investing in and eventually ship to the App store.
Without Claude I likely wouldn't have gotten to the finished prototype version to use in the real world.
For Indy dev, I think LLMs are a new source of solutions. This app is too niche to justify building and marketing without LLM assistance. It likely won't earn more than $25k/year but good enough!
For people doing work with LLMs as an assistant for codebase searching, reviews, double checks, and things like that the $20/month plan is more than fine. The closer you get to vibecoding and trying to get the LLM to do all the work, the more you need the $100 and $200 plans.
On the ChatGPT side, the $20/month subscription plan for GPT Codex feels extremely generous right now. I tried getting to the end of my window usage limit one day and could not.
> so the "just another $20 SaaS" argument doesn't sound too good
Having seen several company's SaaS bills, even $100/month or $200/month for developers would barely change anything.
Not recognizing the essential role of sales seemed to be a common mistake.
It identified advertising as part of the category that it classed as heavily-bullshit-jobs for reason of being zero-sum—your competitor spends more, so you spent more to avoid falling behind, standard red queen’s race. (Another in this category was the military, which is kinda the classic case of this—see also, the Missile Gap, the dreadnought arms race, et c.) But not sales, IIRC.
Like when a competing country builds their tenth battleship, so you commission another one to match them. The world would have been better if neither had been build. Money changed hands (one supposes) but the aim of the whole exercise had no effect. It was similar to paying people to dig holes a fill them back in again, to the tune of serious money. This was so utterly stupid and wasteful that there was a whole treaty about it, to try to prevent so many bullshit jobs from being created again.
Or when Pepsi increases their ad spending in Brazil, so Coca Cola counters, and much of the money ends up accomplishing little except keeping things just how they were. That component or quality of the ad industry, the book claims, is bullshit, on account of not doing any good.
The book treats of several ways in which a job might be bullshit, and just kinda mentions this one as an aside: the zero-sum activity. It mostly covers other sorts, but this is the closest I can recall it coming to declaring sales “bullshit” (the book rarely, bordering on never, paints even most of an entire industry or field as bullshit, and advertising isn’t sales, but it’s as close as it got, as I recall)
Is it a waste of effort when two companies try to make the best electric car?
If it is wasteful what does aw orld look like where nobody ever spends resources on a goal which overlaps with someone else?
- Which products get included in the candidate list? Every product in existence which claims use? - how many results can it return? And in what order? - which attributes or description of the product is provided to the llm? Who provides it? - how are the claims in those descriptions verified? - what if my business believes the claims or description of our product is false? - how will the llm change its relative valuations based on demand?
> The only way advertising won't exist or won't be needed is when humanity becomes a hive mind and removes all competition.
I don't need advertisement to pick the best product for myself. I have a list of requirements that I need fulfilled – why do I need advertisement for it?
No. They can ban particular modes. They can’t stop people from using power and money to spread ideas.
In the US hedge funds are banned from advertising and all they did is change their forms of presentation to things like presenting at conferences or on podcasts.
If there was a socialist fantasy of a government review board for which all products were submitted before being listed in a government catalog. Then advertising would be lobbying and jockeying that review board to view your product in a particular way. Or merely to go through the process and ensure correct information was kept.
It says stuff like why can’t a customer just order from an online form? The employee who helps them doesn’t do anything except make them feel better. Must be a bullshit job. It talks specifically about my employees filling internal roles like this.
> advertising
I understand the arms race argument, but it’s really hard to see what an alternative looks like. People can spend money to make you more aware of something. You can limit some modes, but that kind of just exists.
I don’t see how they aren’t performing an important function.
There's nothing inherent to socialism that would preclude advertising. It's an economic system where the means of production (capital) is owned by the workers or the state. In market socialism you still have worker cooperatives competing on the market.
If you've read much else you should be able to engage with text properly, and construct charitable interpretations of author's claims or arguments.
Did I miss something?
The example used here was advertising. And then when we push on the example the fallback is to the subjective - feeling unfilled, definition.
So I am still look for concrete examples of bullshit jobs to justify the original comment that AI will find efficiencies by letting us throw these away.
Got any? You are an expert on the text so I’m hoping you can identify one.
See what I mean? We push on where these fake jobs are and you fallback to a subjective internal definition we can’t inspect.
And now let me remind you of the context. If the real definition of bullshit isn’t economic slack, but internal dissatisfaction then this comment would be false:
> What if LLMs are optimizing the average office worker's productivity but the work itself simply has no discernable economic value? This is argued at length in Grebber's Bullshit Jobs essay and book.
Ever actually lived in anything approaching one? Yeah, if the stores are empty, it does not make sense to produce ads for stuff that isn't there ...
... but we still had ads on TV, surprisingly, even for stuff that was in shortage (= almost everything). Why? Because the Plan said so, and disrespecting the Plan too openly would stray dangerously close to the crime of sabotage.
You have no idea.
And people won't give up their shops and fields and other means of production to the government voluntarily, at least not en masse. Thus they have to be forced at a gunpoint, and they always were.
All the subsequent horror is downstream from that. This is what is inherent to building a socialist economy: mass expropriation of the former "exploitative class". The bad management of the stolen assets is just a consequence, because ideologically brainwashed partisans are usually bad at managing anything including themselves.
Yugoslavia was extremely successful, with economic growth that matched or exceeded most capitalist European economies post-WW2. In some ways it wasn't as free as western societies are today but it definitely wasn't totalitarian, and in many ways it was more free - there's a philosophical question in there about what freedom really is. For example Yugoslavia made abortion a constitutionally protected right in the 70s.
I don't want to debate the nuances of what's better now and what was better then as that's beside the point, which is that the idiosyncrasies of the terrible Soviet economy are not inherent to "socialism", just like the idiosyncrasies of the US economy aren't inherent to capitalism.
It is the model, introduced basically everywhere where socialism was taken seriously. It is like saying that cars with four wheels are just one terrible model, because there were a few cars with three wheels.
Yugoslavia was a mixed economy with a lot of economic power remaining in private hands. You cannot point at it and say "hey, successful socialism". Tito was a mortal enemy of Stalin, stroke a balanced neither-East-nor-West, but fairly friendly to the West policy already in 1950, and his collectivization efforts were a fraction of what Marxist-Leninist doctrine demands.
You also shouldn't discount the effect of sending young Yugoslavs to work in West Germany on the total balance sheet. A massive influx of remittances in Deutsche Mark was an important factor in Yugoslavia getting richer, and there was nothing socialist about it, it was an overflow of quick economic growth in a capitalist country.
> You cannot point at it and say "hey, successful socialism"
Yes I can because ideological purity doesn't exist in the real world. All of our countries are a mix of capitalist and socialist ideas yet we call them "capitalist" because that's the current predominant organization.
> Tito was a mortal enemy of Stalin, stroke a balanced neither-East-nor-West, but fairly friendly to the West policy already in 1950, and his collectivization efforts were a fraction of what Marxist-Leninist doctrine demands.
You're making my point for me, Yugoslavia was completely different from USSR yet still socialist. Socialism is not synonymous with Marxist-Leninist doctrine. It's a fairly simple core idea that has an infinite number of possible implementations, one of them being market socialism with worker cooperatives.
Aside from that short period post-WW2, no socialist or communist nation has been allowed to exist without interference from the US through oppressive economic sanctions that would cripple and destroy any economy regardless of its economic system, but people love nothing more than to draw conclusions from these obviously-invalid "experiments".
"You" (and I mean the collective you) are essentially hijacking the word "socialism" to simply mean "everything that was bad about the USSR". The system has been teaching and conditioning people to do that for decades, but we should really be more conscious and stop doing that.
That is what COMECON was supposed to solve, but if you aggregate a heap of losers, you won't create a winning team.
"Socialism is not synonymous with Marxist-Leninist doctrine. It's a fairly simple core idea that has an infinite number of possible implementations, one of them being market socialism with worker cooperatives."
Of that infinite number, the violent Soviet-like version became the most widespread because it was the only one that was somewhat stable when implemented on a countrywide scale. That stability was bought by blood, of course.
No one is sabotaging worker cooperatives in Europe and lefty parties used to given them extra support, but they just don't seem to be able to grow well. The largest one is located in Basque Country and it is debatable if its size is partly caused by Basque nationalism, which is not a very socialist idea. Aside from that one, worker cooperatives of more than 1000 people are rare birds.
"The system has been teaching and conditioning people to do that for decades, but we should really be more conscious and stop doing that."
No one in the former socialist bloc will experiment with that quagmire again. For some reason, socialism is a catnip of intellectuals who continue to defend it, but real-world workers dislike it and defect from various attempts to build it at every opportunity.
We should stop trying to ride dead horses. Collective ownership of means of production on a macro scale is every bit as dead as divine right of kings to rule. There are still Curtis Yarvin types of intellectual who subscribe to the latter idea, but it is pining for the fjords. So is socialism.
What kind of disingenuous argument is that? Existence of COMECON doesn't neutralize the enormous disadvantage and economic pressure of having sanctioned imposed on you.
> Of that infinite number
I'm glad we agree that Soviet communism is not synonymous with "socialism".
> Aside from that one, worker cooperatives of more than 1000 people are rare birds.
You're applying pointless capitalist metrics to non-capitalist organizations and moralizing about how they don't live up to them.
> No one in the former socialist bloc will experiment with that quagmire again.
You're experimenting with socialist policies and values right now, you just don't want to call it by that name because of your weird fixation. Do public healthcare, transport, education, social security benefits ring any bells?
If you talked to people from ex-Yugoslavia, you'd know that many would be happy to return to that time.
> We should stop trying to ride dead horses.
We should stop declaring horses extinct when it's just your own horse that has died.
The Soviets and their satellites (like the DDR), had another problem related to arbitrage, and that is that their professionals (such as doctors and engineers and scientists, all of whom received high quality, free, state-subsidized education), were being poached by the Western Bloc countries (a Soviet or East German engineer would work for half the local salary in France or West Germany, _and_ they would be a second class citizen, easy to frighten with deportation -- the half-salary was _much_ greater than what they could earn in the Eastern Bloc). The iron curtain was erected to prevent this kind of arbitrage (why should the Soviets and satellites subsidize Western medicine and engineering? Shouldn't a capitalist market system be able to sustain itself? Well no, market systems are inefficient by design, and so they only work as _open_ systems and not _closed_ systems -- they need to _externalize_ the costs and _internalize_ the gains, which is why colonialism was a thing to begin with, and why the "third world" is _still_ a thing).
Note that after the Berlin Wall fell, the first thing to happen was mass migrations of all kinds of professionals (such as architects and doctors) and semi-professionals (such as welders and metal-workers), creating an economic decline in the East, and an economic and demographic boom in the West (the reunification of Germany was basically a _demographic_ subsidy -- in spite of the smaller size, East Germany had much higher birth rates for _decades_; and after the East German labor pool was integrated, Western economies sought to integrate the remaining Eastern labor pools (more former Yugoslavs live abroad in Germany than in any other non-Yugo part of the world [the USA numbers are iffy, but if true Croatians are the only exception, with ~2M residents in USA, which seems unlikely]).
The problem, in the end, is that all of these countries are bound by economic considerations (this is thesis of Marx, by the way), and they cannot escape the vicious arbitrage cycle (I mean, here in the USA, we have aggressively been brain-draining _ourselves_ since at least 1980, which is why we have the extreme polarization, stagnation, and instability _today_ -- it is reminiscent of the Soviet situation in the mid 1980s to late 1990s). Not without something like a world government (if there is only one account to manage, there is no possibility of deficit or surplus, unless measured inter-temporally), or an alternative flavor of globalization.
Internationalism is a wonderful ideology, and one that I support. You can make the case that Yugoslavia, the USSR, etc, were an early experiment in Internationalism, that each succumbed to corruption and unclear thinking (a citizenry that is _inclusive_ by nature and can _think_ clearly is a hard requirement for any successful polity). Globalization, on the other hand, has a bit of an Achilles Heel: when countries asked why they should open their borders and economies to outsider/foreigners, they were told, "so that we can all get rich!". The problem is that once the economic gains get squeezed out of globalization, countries will start looking for new ways to rich, even if it means reversing decades of integration. Appealing to people's greed only works to the extent that you can placate their appetites. We should have justified Internationalism using _intrinsic_ arguments: "we should integrate because learning how others see and experience the world is intrinsically beautiful, and worth struggling for".
Note that most of these economic pathologies disappear, when the reserve currency (dollar) is replaced with a self-balancing currency (like Keynes' Bancor: https://en.wikipedia.org/wiki/Bancor). We have the tools, but everyone wants to feel like the only/greatest winner. These are the first people that have to be exiled.
> There’s not much of value to obtain from the book.
Anthropological insight has much more value than anything economists may produce on economy.
Modern economics is literally a bullshit job generating process or complex system.
Assistant is dispatching a courier to get medical records. AI auto completes to include the address. Normally they wouldn't put the address, the courier knows who we work with, but AI added it so why not. Except it's the wrong address because it's for a different doctor with the same name. At least they knew to verify it, but still mistakes like this happening at scale is making the other time savings pretty close to a wash.
See also: AI-Generated “Workslop” Is Destroying Productivity [1]
[1] https://hbr.org/2025/09/ai-generated-workslop-is-destroying-...
What AI has done is accelerate and magnify both the positives and the negatives.
There are a lot of white-collar tasks that have far lower quality and correctness bars. "Researching" by plugging things into google. Writing reports summarizing how a trend that an exec saw a report on can be applied to the company. Generating new values to share at a company all-hands.
Tons of these that never touch the "real world." Your assistant story is like a coding task - maybe someone ran some tests, maybe they didn't, but it was verifiable. No shortage of "the tests passed, but they weren't the right test, this broke some customers and had to be fixed by hand" coding stories out there like it. There are pages and pages of unverifiable bullshit that people are sleepwalking through, too, though.
Nobody already knows if those things helped or hurt, so nobody will ever even notice a hallucination.
But everyone in all those fields is going to be trying really really hard to enumerate all the reasons it's special and AI won't work well for them. The "management says do more, workers figure out ways to be lazier" see-saw is ancient, but this could skew far towards "management demands more from fewer people" spectrum for a while.
In all areas where there's less easy ways to judge output there is going to be correspondingly more value to getting "good" people. Some AI that can produce readable reports isn't "good" - what matters is the quality of the work and the insight put into it which can only be ensured by looking at the workers reputation and past history.
That's not obvious at all if the AI writing the tests is different than the AI writing the code being tested. Put into an adversarial and critical mode, the same model outputs very different results.
Someone needs to build an agentic tool that does strict, enforced TDD.
Obviously this is only partially true but it's true enough.
It takes humans quite a long time to learn the external context that lets them write good tests IMO. We have trouble feeding enough context into AIs to give them equal ability. One is often talking about companies where nobody bothers to write down more than 1/20th of what is needed to be an effective developer. So you go to some place and 5 years later you might be lucky to know 80% of the context in your limited area after 100s of meetings and talking to people and handling customer complaints etc.
I have been doing this with coding agents across LLM providers for a while now, with very successful results. Grok seems particularly happy to tell Anthropic where it’s cutting corners, but I get great insights from O3 and Gemini too.
Except the test suite isnt just something that appears and the bugs dont necessarily get covered by the test suite.
The bugginess of a lot of the software i use has spiked in a very noticeable way, probably due to this.
>But everyone in all those fields is going to be trying really really hard to enumerate all the reasons it's special and AI won't work well for them.
No, not everyone. Half of them are trying to lean in to the changing social reality.
The gaslighting from the executive side, on the other hand, is nearly constant.
Consider, for example, the following python code:
x = (5)
vs x = (5,)
One is a literal 5, and the other is a single element tuple containing the number 5. But more importantly, both are valid code.Now imagine trying to spot that one missing comma among the 20kloc of code one so proudly claims AI helped them "write", especially if it's in a cold path. You won't see it.
Disagree.
Even though performing checks on dynamic PLs is much harder than on static ones, PLs are designed to be non-ambiguous. There should be exactly 1 interpretation for any syntactically valid expression. Your example will unambiguously resolve to an error in a standard-conforming Python interpreter.
On the other hand, natural languages are not restricted by ambiguity. That's why something like Poe's law exists. There's simply no way to resolve the ambiguity by just staring at the words themselves, you need additional information to know the author's intent.
In other words, an "English interpreter" cannot exist. Remove the ambiguities, you get "interpreter" and you'll end up with non-ambiguous, Python-COBOL-like languages.
With that said, I agree with your point that blindly accepting 20kloc is certainly not a good idea.
Those are both syntactically valid lines of code. (it's actually one of python's many warts). They are not ambiguous in any way. one is a number, the other is a tuple. They return something of a completely different type.
My example will unambiguously NOT give an error because they are standard conforming. Which you would have noticed had you actually took 5 seconds to try typing them in the repl.
> Those are both syntactically valid lines of code. (it's actually one of python's many warts). They are not ambiguous in any way. one is a number, the other is a tuple. They return something of a completely different type.
You just demonstrated how hard it is to "check" an email or text message by missing the point of my reply. > "Now imagine trying to spot that one missing comma among the 20kloc of code"
I assume your previous comment tries to bring up Python's dynamic typing & late binding nature and use it as an example of how it can be problematic when someone tries to blindly merge 20kloc LLM-generated Python code.My reply, "Your example will unambiguously resolve to an error in a standard-conforming Python interpreter." tried to respond to the possibility of such an issue. Even though it's probably not the program behavior you want, Python, being a programming language, will be 100% guaranteed to interpret it unambiguously.
I admit, I should have phrased it a bit more unambiguously than leaving it like that.
Even if it's hard, you can try running a type checker to statically catch such problems. Even if it's not possible in cases of heavy usage of Python's dynamic typing feature, you can just run it and check the behavior at runtime. It might be hard to check, but not impossible.
On the other hand, it's impossible to perform a perfectly consistent "check" on this reply or an email written in a natural language, the person reading it might interpret the message in a completely different way.
The few times I've tried giving LLMs a shot I've had them warning me of not putting some validations in, when that exact validation was exactly 1 line below where they stopped looking.
And even if it did pass an AI code review, that's meaningless anyway. It still needs to be reviewed by an actual human before putting it into production. And that person would still get scrolling blindness whether or not the ai "reviewer" actually detected the error or not.
I didn't say they were guaranteed to find it: I said they were really good at finding these sorts of errors. Not perfect: just really good. I also didn't make any assumption: I said in my experience, by which I mean the code you shared is similar to a portion of the errors that I've seen LLMs find.
Which LLMs have you used for code generation?
I mostly use claude-opus-4-6 at the moment for development, and have had mostly good experiences. This is not to say it never gets anything wrong, but I'm definitely more productive with it than without it. On GitHub I've been using Copilot for more limited tasks as an agent: I find it's decent at code review, but more variable at fixing problems it finds, and so I quite often opt for manual fixes.
And then the other question is, how do you use them? I tend to keep them on quite a short leash, so I don't give them huge tasks, and on those occasions where I am doing something larger or more complex, I tend to write out quite a detailed and prescriptive prompt (which might take 15 minutes to do, but then it'll go and spend 10 minutes to generate code that might have taken me several hours to write "the old way").
Most "Bullshit Jobs" can already be automated, but can isnt always should or will. Graeber is a capex thinker in an opex world.
That book was very different than what I expected from all of the internet comment takes about it. The premise was really thin and did't actually support the idea that the jobs don't generate value. It was comparing to a hypothetical world where everything is perfectly organized, everyone is perfectly behaved, everything is perfectly ordered, and therefore we don't have to have certain jobs that only exist to counter other imperfect things in society.
He couldn't even keep that straight, though. There's a part where he argues that open source work is valuable but corporate programmers are doing bullshit work that isn't socially productive because they're connecting disparate things together with glue code? It didn't make sense and you could see that he didn't really understand software, other than how he imagined it fitting into his idealized world where everything anarchist and open source is good and everything corporate and capitalist is bad. Once you see how little he understands about a topic you're familiar with, it's hard to unsee it in his discussions of everything else.
That said, he still wasn't arguing that the work didn't generate economic value. Jobs that don't provide value for a company are cut, eventually. They exist because the company gets more benefit out of the job existing than it costs to employ those people. The "bullshit jobs" idea was more about feelings and notions of societal impact than economic value.
> Jobs that don't provide value for a company are cut, eventually.
Uhm, seems like Greaber is not the only one drawing conclusions from a hypothetical perfect world
I tried to respond to the specific conversation about Bullshit Jobs above. In my experience, the way this book is brought up so frequently in online conversations is used as a prop for whatever the commenter wants it to mean, not what the book actually says.
I think Graeber did a fantastic job of picking "bullshit jobs" as a topic because it sounds like something that everyone implicitly understands, but how it's used in conversation and how Graeber actually wrote about the topic are basically two different things
I thought he made a case for both societal and economic impact.
Not necessarily, I’ve seen a lot of jobs that were just flying under the radar. Sort of like a cockroach that skitters when light is on but roams freely in the dark.
I don't know if maybe he wasn't explaining it well enough, but that kind of reasoning makes some sense.
A lot of code is written because you want the output from Foo to be the input to Bar and then you need some glue to put them together. This is pretty common when Foo and Bar are made by different people. With open source, someone writes the glue code, publishes it, and then nobody else has to write it because they just use what's published.
In corporate bureaucracies, Company A writes the glue code but then doesn't publish it, so Company B which has the same problem has to write it again, but they don't publish it either. A hundred companies are then doing the work that only really needed to be done once, which makes for 100 times as much work, a 1% efficiency rate and 99 bullshit jobs.
Sure, but there's no such thing as "the company." That's shorthand - a convenient metaphor for a particular bunch of people doing some things. So those jobs can exist if some people - even one person - gets more benefit out of the job existing than it costs that person to employ them. For example, a senior manager padding his department with non-jobs to increase headcount, because it gives him increased prestige and power, and the cost to him of employing that person is zero. Will those jobs get cut "eventually"? Maybe, but I've seen them go on for decades.
But he states that expressis verbis, so your discovery is not that spectacular.
Although he gives examples of jobs, or some aspects of jobs, that don't help to deliver what specific institutions aim to deliver. Example would be bureaucratization of academia.
This is not true at all. You can find plenty of examples going either way but it’s far from truth from being a universal reality
Maybe you're lucky enough to be doing cutting edge research or do something that really seriously impacts human beings, but I've done plenty of "mission critical right fucking now" work that a week from now (or even hours from now, when I worked for a content marketing business) is beyond irrelevant. It's an amazing thing watching marketing types set money on fire burning super expensive developer time (but salaried, so they discount the cost to zero) just to make their campaigns like 2-3% more efficient.
I've intentionally sat on plenty of projects that somebody was pushing really hard for because they thought it was the absolute right necessary thing at the time and the stakeholder realized was pointless/worthless after a good long shit and shower. This one move has saved literally man years of work to be done and IMO is the #1 most important skill people need to learn ("when to just do nothing").
Companies are obviously not hesitant to lay off anyone, especially for cost saving. It is interesting how you think that people are laid off because they’re unproductive.
Was a LLM used during that optimization? Yes.
Who will correlate the sudden productivity improvement with our optimization of the data flow with the availability of a LLM to do such optimizations fast enough that no project+consultants+management is needed ?
No one.
Just like no one is evaluating the value of a hammer or a ladder when you build a house.
This is where the whole "show me what you built with AI" meme comes from, and currently there's no substitute for SWEs. Maybe next year or next next year, but mostly the usage is generating boring stuff like internal tool frontends, tests, etc. That's not nothing, but because actually writing the code was at best 20% of the time cost anyway, the gains aren't huge, and won't be until AI gets into the other parts of the SDLC (or the SDLC changes).
It’s easy to convince yourself that it is, and anyone can massage some internal metric enough to prove their desired outcome.
I think broadly that's a paradoxical statement; improving office productivity should translate to higher gdp; whatever it is you're doing in some office - even if you're selling paper or making bombs, if you're more productive it means you're selling more (or using less resources to sell the same amount); that should translate to higher gdp (at least higher gdp per worker, there's the issue of what happens to gdp when many workers get fired).
Society as a whole is no better off since no value or wealth was generated, but the number did go up.
A whole bunch of our economy is broken windows like this to varying degrees.
Edit: If anyone haven't watched Yes Minster, you should go and watch it, it is a documentary on UK Government that is still true today as it was 40-50 years ago.
i.e. it's not the LLM, it's that they're not being used properly.
I get accused of the no true Scotsman argument because I think agile can be done right, for example. Is work bullshit because an LLM doesn't help it?
i'm going to need you to go ahead and come in on sunday
And it is of this lowly commenter's opinion that proofreading for accuracy and clarity is harder than writing it yourself and defending it later.
As measured by whom? The same managers who demanded we all return to the office 5 days a week because the only way they can measure productivity is butts in seats?
We do have a way to see the financial impact - just add up Anthropic and oAI's reported revenues -> something like $30b in annual run rate. Given growth rates, (stratospheric), it seems reasonable to conclude informed buyers see economic and/or strategic benefit in excess of their spend. I certainly do!
That puts the benefits to the economy at just around where Mastercard's benefits are, on a dollar basis. But with a lot more growth. Add something in there for MS and GOOG, and we're probably at least another $5b up. There are only like 30 US companies with > $100bn in revenues; at current growth rates, we'll see combined revenues in this range in a year.
All this is sort of peanuts though against 29 trillion GDP, 0.3%. Well not peanuts, it's boosting the US GDP by 10% of its historical growth rate, but the bull case from singularity folks is like 10%+ GDP growth; if we start seeing that, we'll know it.
All that said, there is real value being added to the economy today by these companies. And no doubt a lot of time and effort spent figuring out what the hell to do with it as well.
Does profitable always equal useful? Might other cultures justifiably think differently, like the Amish?
I also didn’t talk profitable. Upshot, though, I don’t think it’s just a US thing to say that when money exchanges hands, generally both parties feel they are better off, and therefore there is value implied in a transaction.
As to what it will be used for: yes.
And that's the point here: value is handicapped by the web interface, and we are stuck there for the foreseeable future until the tech teams get their priorities straight and build decent data integration layers, and workflow management platforms.
Real life example: A client came to me asking how to compare orders against order confirmation from the vendor. They come as PDF files. Which made me wonder: Wait, you don't have any kind of API or at least structured data that the vendor gives you?
Nope.
And here you are. I am not talking about a niche business. I assume that's a broader problem. Tech can probably automate everything and this since 30 years. Still business lack of "proper" IT processes, because at the end every company is unique and requires particular measures to be "fully" onboarded to IT based improvements like that.
I've seen this play out with time sheets. Every day factory floor workers write up every job they do on time sheets, paper and pencil, but management wants excel spreadsheets with pretty plots. Solution? One old typist who can type up all the time sheets every day. No OCR trouble, no expensive developer time, if she encounters illegible numbers she just cross references against the job fliers to figure it out instead of throwing errors the first time a 0 looks like a 6.
- protect yourself from data loss / secret leaks - what it can and can't do - trust issues & hallucinations - Can't just enable Claude for Excel and expect people to become Excel wizards.
I think it could definitely already create economic benefit, after someone instructed clearly how to use it and how to integrate it in your work. Most people are really not good at figuring that out on their own, in a busy workday, when left to their own devices and companies are just finding out where the ball is moving and what to organize around too.
So I can totally see a lot of failed experiments and people slowly figuring stuff out, and all of that not translating to measurable surpluses in a corp, in a setup similar to what OP laid out.
This shit will take like 10 years to adopt properly, at least in most boomer companies.
I use these all the time nowadays and they are great tools when utilized properly but I have hard time seeing it replace functions completely due to humans having limited cognitive capacity to multitask and still needing to review stuff and build infra to actually utilize all this..
Talking about macro economics, I don’t think that number is correct.
Getting people to figure out how to enter questions is easy. Getting people to a point where they don't burn up all the savings by getting into unproductive conversations with the agent when it gets something wrong, is not so easy.
Only until the loans come due. We're still in the "uber undercutting medalian cabs" part of the game.
ChatGPT just lets you generate slop, that may be helpful. For the vast majority of industries it doesn’t actually offer much. Your meme departments like HR might be able to push out their slop quicker, but that doesn’t improve profitability.
InfoSec and Legal would like a word with you...
I read an article yesterday about people working insane hours at companies that have bet heavily on AI. My interpretation is that a worker runs out of juice after a few hours, but the AI has no limit and can work its human tender to death.
I mean, sure. If you want to use it for 20 minutes and wait for two hours at a time.
My company is spending 20-50x that much per head easily from the firmwide cost numbers being reported.
They had to set circuit breakers because some users hit $200/day.
Is it fair to say that wall street is betting America's collective pensions on AI...
[0]: https://www.almendron.com/tribuna/wp-content/uploads/2018/03...
https://www.nber.org/system/files/working_papers/w25148/w251...
I don't agree that real GDP measures what he thinks it measures, but he opines
>Data released this week offers a striking corrective to the narrative that AI has yet to have an impact on the US economy as a whole. While initial reports suggested a year of steady labour expansion in the US, the new figures reveal that total payroll growth was revised downward by approximately 403,000 jobs. Crucially, this downward revision occurred while real GDP remained robust, including a 3.7 per cent growth rate in the fourth quarter. This decoupling — maintaining high output with significantly lower labour input — is the hallmark of productivity growth.
https://www.ft.com/content/4b51d0b4-bbfe-4f05-b50a-1d485d419...
[0] on the basis that IT and AI are not general technologies in the mold of the dynamo, keyword "intangibles", see section 4 p21, A method to measure intangibles
Is consumer spending measured in total dollars spent? If so, isn't that curious wrinkle in an economy of rising prices, and decreasing purchasing power?
If true, I believe less quantity could be purchased at a higher cost per person, making it appear that consumer spending is up.
Presumably these numbers are benchmarked/peg to some sort of constant and/or standardization
https://en.wikipedia.org/wiki/Personal_consumption_expenditu...
https://fortune.com/2026/02/15/ai-productivity-liftoff-doubl...
Source of the Stanford-approved opinion: https://www.ft.com/content/4b51d0b4-bbfe-4f05-b50a-1d485d419...
Is that somewhat substantiated assumption? I recall learning on University in 2001 the history of AI and that initial frameworks were written in 70's and that prediction was we will reach human-like intelligence by 2000. Just because Sama came up with this somewhat breakthrough of an AI, it doesn't mean that equal improvement leaps will be done on a monthly/annual basis going forward. We may as well not make another huge leaps or reach what some say human intelligence level in 10 years or so.
The 1990s boom was in large part due to connectivity -- millions[1] of computers joined the Internet.
[1] _ In the 1990s. Today, there are billions of devices connected, most of them Android devices.
On hacker news, a very tech literate place, I see people thinking modern AI models can’t generate working code.
The other day in real life I was talking to a friend of mine about ChatGPT. They didn’t know you needed to turn on “thinking” to get higher quality results. This is a technical person who has worked at Amazon.
You can’t expect revolutionary impact while people are still learning how to even use the thing. We’re so early.
As for the last example, for all the money being spent on this area, if someone is expected to perform a workflow based on the kind of question they're supposed to ask, that's a failure in the packaging and discoverability aspect of the product, the leaky abstraction only helps some of us who know why it's there.
1. People who only think of using AI in very specific scenarios. They don’t know when you use it outside of the obvious “to write code” situations and they don’t really use AI effectively and get deflated when AI outputs the occasional garbage. They think “isn’t AI supposed to be good at writing code?”
2. People who let AI do all the thinking. Sometimes they’ll use AI to do everything and you have to tell them to throw it all away because it makes no sense. These people also tend to dump analyses straight from AI into Slack because they lack the tools to verify if a given analysis is correct.
To be honest, I help them by teaching them fairly rigid workflows like “you can use AI if you are in this specific situation.” I think most people will only pick up tools effectively if there is a clear template. It’s basically on-the-job training.
I think this is the prior you should investigate. That may be what HN used to be. But it's been a long time since it has been an active reality. You can still see actual expert opinions on HN, but they are the minority more and more.
From personal experience, I've also noticed that some of the most toxic discourse and responses I've received on this platform are overwhelmingly from post-2022 users.
But from my social feed the impression was that it is taking over the world:)
I asked it because I am building something similar since some tome and I thought its over they were faster than me but as it appears there’s no real adoption yet. Maybe there will be some once they release it as part of ChatGPT but even then it looks like too early as actually few people are using the more advanced tools.
It’s definitely in very early stage. It appears that so far the mainstream success in AI is limited to slop generation and even that is actually small number of people generating huge amounts of slop.
If you have been working on a usecase similar to OpenClaw for sometime now I'd actually say you are in a great position to start raising now.
Being first to market is not a significant moat in most cases. Few people want to invest in the first company in a category - it's too risky. If there are a couple of other early players then the risk profile has been reduced.
That said, you NEED to concentrate on GTM - technology is commodified, distribution is not.
> It appears that so far the mainstream success in AI is limited to slop generation and even that is actually small number of people generating huge amounts of slop
The growth of AI slop has been exponential, but the application of agents for domain specific usecases has been decently successful.
The biggest reason you don't hear about it on HN is because domain-specific applications are not well known on HN, and most enterprises are not publicizing the fact that they are using these tools internally.
Furthermore, almost anyone who is shipping something with actual enterprise usage is under fairly onerous NDAs right now and every company has someone monitoring HN like a hawk.
On my app the tech is based on running agent generated code on JavaScriptCore to do things like OpenClaw, I’m wrapping the JS engine with the missing functionality like networking, file access and database access so I believe I will not have a problem with releasing it on Apple AppStore as I use their native stack. Then since this stack is also OS, I’m making a version that will run on Linux, the idea being users develops their solution on their device(iOS&Mac currently) see it working and and then deploys on a server with a tap of a button, so it keeps running.
You need to answer these questions in order to decide whether a Show HN makes sense versus a much more targeted launch.
If you do not know how to answer these questions you need to find a cofounder asap. Technology is commodified. GTM, sales, and packaging is what turns technology into products. Building and selling and fundraising as 1 person is a one-way ticket to burnout, which only makes you and your product less attractive.
I also highly recommend chatting with your network to understand common types of problems. Once you've identified a couple classes of problems and personas for whom your story resonates, then you can decide what approach to take.
Best of luck!
Thanks for the advice! I’m at a stage where I want to have such tool and see who else wants it. Not sure yet about it’s viability as a business and what is the exact market. Maybe I will find out by putting it into the wild and that’s why I consider to release it as a mobile app first.
But even with that persona, it should already answer your question whether posting on HN and producthunt should be a core part of your strategy. Not a lot of social media managers or compliance people around here. And even for crypto traders there are better places to pitch products to them
You need to narrow it down to a single and specific persona and business domain.
This is because it takes years to fully flesh out and productionize a workflow from scratch, so concentrating on a business domain you know intimately well helps you build that muscle, which you can then repeat if you are able to hit revenue metrics for a Series A/B.
Monitoring specific user accounts or keywords? Is this typically done by a social media reputation management service?
Shocking results, I say!
Most people are just "not that deep in there" as most people on HN.
A guy attached Claude to his socials, groundbreaking tech.
While they were deep in software development in general, no body of them read any of the essential/required daily industrial news (also not that one related to doing software development in sector ABC)
:-)
So no, even people somehow attached to a topic are not necessarily somehow deeper involved.
If one lets AI FOMO since the release of chatgpt drive them they'd be glued to their screen 24/7.
OAI wants to keep the hype train going. That is all. OpenClaw is just a project that attracted the interests of people messing about with LLMs. Which as a proportion of economically active people is.... tiny.
They brought him (Pete) over as he seems to have some way of thinking about LLMs in the form of a product. Will he have repeatable success on a large scale? Who knows. I doubt it personally.
Really? Can you show any examples of someone claiming AI models cannot generate working code? I haven't seen anyone make that claim in years, even from the most skeptical critics.
I started working today on a project I hadn’t touched in a while but I now needed to as it was involved in an incident where I needed to address some shortcomings. I knew the fix I needed to do but I went about my usual AI assisted workflow because of course I’m lazy the last thing I want to do is interrupt my normal work to fix this stupid problem.
The AI doesn’t know anything about the full scope of all the things in my head about my company’s environment and the information I need to convey to it. I can give it a lot of instructions but it’s impossible to write out everything in my head across multiple systems.
The AI did write working code, but despite writing the code way faster than me, it made small but critical mistakes that I wouldn’t have made on my first draft.
For example, it just added in a command flag that I knew that it didn’t need, and it actually probably should have known it, too. Basically it changed a line of code that it didn’t need to touch.
It also didn’t realize that the curled URL was going to redirect so we needed an -L flag. Maybe it should have but my brain knew it already.
It also misinterpreted some changes in direction that a human never would have. It confused my local repository for the remote one because I originally thought I was going to set a mirror, but I changed plans and used a manual package upload to curl from. So it out the remote URL in some places where the local one should have been.
Finally, it seems to have just created some strange text gore while editing the readme where it deleted existing content for seemingly no reason other than some kind of readline snafu.
So yes it produced very fast great code that would have taken me way longer to do, but I had to go back and consume a very similar amount of time to fix so many things that I might as well have just done it manually.
But hey I’m glad my company is paying $XX/month for my lazy workday machine.
This is your problem: How should it know if you do not provide it?
Use Claude - in the pro version you can submit files for each project which are setting the context: This can be files, source code, SQL scripts, screenshots whatever - then the output will be based on your context given by providing these files.
If I was truly going to automate this one-time task I would have to give the AI access to my browser or an API token for the repository provider, so I’m either giving it dangerous modification capability via browser automation or I’m spending even more time setting up API access and trusting that it actually knows how to interact with the service via API calls.
My company doesn’t provide Claude, they give me GitHub Copilot Pro or whatever it’s called, and when I provided the website it needed to get the RPM files I was working with it didn’t actually do anything with it. It just wrote a readme file that told me what to do. Like I mention it also just eventually mistook the remote repository as my local internal repository.
And one of the specific commands it screwed up was in my existing script and was already correct, it just decided to change it for no discernible reason. I didn’t ask it to do anything related to that particular line.
With such a high error rate, I would be hesitant to actually integrate AI to other systems to try to achieve a more fully automated workflow.
They can do a tolerable job with super popular /simple things like web dev and Python. It really depends on what you're doing.
> Github Copilot has been great in getting that code coverage up marginally but ass otherwise.
When someone claims that AI can't generate working code I assume that it means consistently generating working code. We're talking about a tool. It has to work more often than not and on codebases that we tend to work with, i.e legacy code.
Personally I don't claim that because I'm using everyday to generate working code.
It takes time to get an intuition for the kinds of problems they've seen in pre-training, what environments it faced in RL, and what kind of bizarre biases and blindspots it has. Learning to google was hard, learning to use other peoples libraries was hard, and its on par with those skills at least.
If there is a well known design pattern you know, thats a great thing to shout out. Knowing what to add to the context takes time and taste. If you are asking for pieces so large that you can't trust them, ask for smaller pieces and their composition. Its a force multiplier, and your taste for abstractions as a programmer is one of the factors.
In early usenet/forum days, the XY problem described users asking for implementation details of their X solution to Y problem, rather than asking how to solve Y. In llm prompting, people fall into the opposite. They have an X implementation they want to see, and rather than ask for it, they describe the Y problem and expect the LLM to arrive at the same X solution. Just ask for the implementation you want.
Asking bots to ask bots seems to be another skill as well.
If not, you don't know how to use it efficiently.
A large part of using AI efficiently is to significantly lower that review burden by having it do far more of the verification and cleanup itself before you even look at it.
- PRD and spec fulfillment review
- code review + correction loops
- security review + corrections
- addl. test coverage and tidying
- addl. type checks and tidying
- addl. lint checks and tidying
- maybe more I haven't listed
And these are run after each commit, so you can only imagine the costs per engineer doing this 10, 20, 50+ times per day depending on how much work they're knocking out.
The question is what your time is worth for the company, and which tasks costs less to have an agent automate than having you do.
The other constraint is, for those who are being laid off (maybe because of cost reduction to support an AI budget for a smaller team to use), engineers wanting to expand their skill set and practice these levels of usage + efficiency are effectively unable to with their own funding, making it more difficult to find employment as expectations heighten.
Prior to AI entering the fray, software development was largely free for everyone, allowing anyone with enough time and motivation to build the skills towards gainful employment. As AI becomes more prevalent and expectations around how it's used become higher, fewer and fewer applicants will be able to claim they have the experience necessary because it was out of reach due to costs.
Right now you need to be Uncle Moneybags to do this in your personal life.
If you're lucky, your employer is footing the bill but otherwise... Ugh. It's like converting your app running perfectly fine on a cheap VPS to AWS Lambda. In theory, it's fine but in reality the next bill you get could make you faint.
I'm unconvinced AI reviewing AI is the answer here, because all LLMs have the same flaws. To me, the harness/guard rails for AI should be different technologies that work differently and in a more formal sense. IE, static code analysis, linters, tests, etc.
(Linting has actually been, by far, the BEST code quality enforcers for the agents I've run so far, and it's a lot cheaper and more configurable than running more agents.)
My current project that I started this weekend is a rust client server game with the client compiled into web assembly.
I do these projects without reading the code at all as a way to gauge what I can possibly do with AI without reading code, purely operating as a PM with technical intuition and architectural opinions.
So far Opus 4.6 has been capable of building it all out. I have to catch issues and I have asked it for refactoring analysis to see if it could optimize the file structure/components, but I haven't read the code at all.
At work I certainly read all the code. But would recommend people try to build something non trivial without looking at the code. It does take skill though, so maybe start small and build up the intuition on how they have issues, etc. I think you'll be surprised how much your technical intuition can scale even when you are not looking at the code.
I have friends in security on major platforms who are impressed by the security review of the SOT models. Certainly better than the average bootstrapped founder.
I'm building to play with my friends online.
They vibe-coded a complete rewrite of their products in a few months without any human review. Hundreds of thousands LOC. I feel sorry for the remaining engineers having to learn everything they just generated, and are now having customers use.
I agree. I'm constantly correcting the code it generates. But then, I do the same for humans when I review their PRs, and the LLM generated the code in a 100th of the time (or whatever figure you prefer).
Last time he said: "yes yes I know about ChatGPT, but I do not use it at work or home."
Therefore, most people wont even know about Gemini, Grok or even Claude.
I am completely flooded with comments and stories about how great LLMs are at coding. I am curious to see how you get a different picture than this? Can you point me to a thread or a story that supports your view? At the moment, individuals thinking AI cannot generate working code seem almost inexistent to me.
Folks like this have never used AI inside of an IDE or one of the CLI AI tools. Without that perspective, AI seems mostly like a gimmick.
What makes you think people this will ever change? Have you seen how well people know their already existing tools?
(Think about this old black/white or green mainframe screens - horrible looking but it gets their job done)
I think it's more likely that people "feel" more productive, and/or we're measuring bad things (lines of code is an awful way to measure productivity -- especially considering that these agents duplicate code all the time so bloat is a given unless you actively work to recombine things and create new abstractions)
So we are in the process of "adapting a technology". Welcome, keep calm, observe, don't be ashamed to feel emotions like fear, excitement, anger and all else.
While adapting, we learn how to use it better and better. At first, we try "do all the work for me", then "ok, that was bad, plan what you would do, good, adjust, ok do it like this" etc etc.
A couple of years into the future this knowledge is just "passed on". If productivity grew and we "figured out how to get more out of the universe", then no jobs had to be lost, just readapted. And "investors" get happy not by "replacing workers", but by "reaping win-win rewards" from the universe at large.
There are dangers of course, like "maybe this is truly a huge win-win, but some loses can be hidden, like ecology", but "I hope there are people really addressing these problems and this win-win will help them be more productive as well".
Wide spread internet access turned expensive toys (PCs) into useful assets.
Maybe! Or it might never pan out, or it may pan out way better. Complicated things like this rarely turn out the way people expect, no matter how smart.
But I can say that, judging by historical artifacts, a lot of it was along the same broad lines as AI. And we maybe don’t realize how serious people were about it back then. The technology that actually changed the world was so comparatively boring and pragmatic that the stuff that was being hyped back then seems comically overwrought. It’s easy to assume it must have been a joke all along.
The combination of attention-draining social media walled gardens, and the high performance pocket-computers (which are really designed for consumption instead of productivity), created a positive feedback loop that helped destroy the productivity that we won by defeating the paradox in the 1990s. And we have been struggling against this new paradox for twenty years, since. AI seems like it should defeat the paradox because it is a kind of hands-free system, perfect for mobile phones -- but this is really just a very expensive solution to a problem that we have created and allowed to fester. We could just shun the walled gardens, and demand to be paid for our attention and data.
The new productivity paradox (which I do not think AI in its current form can fix[1][2]), is the price that we pay for a prosperous and valuable advertising industry. And as long as the web is seen as an ad-channel, and as long as the web is always vibrating in your pocket, we will keep paying this price. We will eventually end up (metaphorically) lobotomizing our children, and families, and communities, so that the grand-children of ad-executives and tech-bros and frat-bros can grow up healthy, psychologically stable, educated, and comfortably wealthy. (Brain drain: now available literally everywhere).
[1]: It is telling that most LLMs are centralized, and are most useful as search-engines/information-retrieval-systems. The centralization makes them _spyware_, and their ability to directly answer any question, encourages users to actually ask direct questions, instead of stringing search-terms together. This makes the prompts high-signal advertising data (i.e. instead inferring what you are looking for from the search-string, these companies can see _exactly_ what you are looking for and why -- and with LLMs, they can probably turn these promps into joint-probability-tables or whatever other kind of serialization they need to figure out which products to sell you (either on the web or directly in the response to your prompt)).
[2]: As far as copyright infringement goes, LLM outputs may require mass clean-room rewrites (so your productivity, as pathetic as it already is, now gets _halved_ long term) of text, prose, code, and anything else that is produced with them, because of how copyright law works. In legal arts this is called _the fruit of the poison tree_, and any short-term productivity gains, may become long term liabilities that need to be replaced due to _legal mandate_ -- so even if LLMs can eventually produce perfect and faultless outputs, the copyright laws _in all 200+ countries_ would have to be torn down and rebuilt (and this will certainly come at great expense).
* If I don't know how to do something, llms can get me started really fast. Basically it distills the time taken to research something to a small amount.
* if I know something well, I find myself trying to guide the llm to make the best decisions. I haven't reached the state of completely letting go and trusting the llm yet, because the llm doesn't make good long term decisions
* when working alone, I see the biggest productivity boost in ai and where I can get things done.
* when working in a team, llms are not useful at all and can sometimes be a bottleneck. Not everyone uses llms the same, sharing context as a team is way harder than it should be. People don't want to collaborate. People can't communicate properly.
* so for me, solo engineers or really small teams benefit the most from llms. Larger teams and organizations will struggle because there's simply too much human overheads to overcome. This is currently matching what I'm seeing in posts these days
I think companies will need fewer engineers but there will be more companies.
Now: 100 companies who employ 1,000 engineers each
What we are transitioning to: 1000 companies who employ 10 engineers each
What will happen in the future: 10,000 companies who employ 1 engineer each
Same number of engineers.
We are about to enter an era of explosive software production, not from big tech but from small companies. I don't think this will only apply to the software industry. I expect this to apply to every industry.
When Engineering Budget Managers see their AI bills rising, they will fire the bottom 5-10% every 6-12 months and increase the AI assistant budget for the high performers, giving them even more leverage.
Our team shrunk by 50% but we are serving 200% more customers. Every time a dev left, we thought we're screwed. We just leveraged AI more and more. We are also serving our customers better too with higher retention rates. When we onboard a customer with custom demands, we used to have meetings about the ROI. Now we just build the custom demands in the time we took to meet to discuss whether we should even do it.
Today, I maintain a few repos critical to the business without even knowing the programming language they are written in. The original developers left the company. All I know is what is suppose to go into the service and what is suppose to come out. When there is a bug, I ask the AI why. The AI almost always finds it. When I need to change something, I double and triple check the logic and I know how to test the changes.
No, a normal person without a background in software engineering can't do this. That's why I still have a job. But how I spend my time as a software engineer has changed drastically and so has my productivity.
When a software dev say AI doesn't increase their productivity, it truly does feel like they're using it wrong or don't know how to use it.
Of course they can - if you don't know any of the tech-stack details (i.e. a "normal" user), why can't someone else who also doesn't know the tech-stackc details replace you?
What magic sauce do you possess other than tech-stack chops?
When a non software engineer can build a production app as well as I can, I know I won’t be working as a software engineer anymore. In that world, having great ideas, data, compute, and energy will be king.
I don’t think we will get there within the next 3-4 years. Beyond that, who knows.
How big is your team? How many customers? What’s your product? Can we see the code? How do you track defects? Etc.
Part of the reason I’m struggling with this is because we’d be seeing OpenAI, Anthropic, etc. plastering these case studies everywhere if they existed. Instead, I’m stuck using CC and all its poorly implemented warts.
Companies are charged per token, which means heavy AI users deliver more and stress budgets. They recently announced significant payroll costs over the past ~3 years.
Those savings I think will partially be reclaimed by AI companies, enabling the high performers more ai model usage.
During the promo review, people will look how many projects were done and the impact of those projects.
I'm not saying it is the case, just making it apparent how unreliable it is to measure productivity by comparing what's happening at the lowest level in a company to its financials.
Imagine if Microsoft didn't invest in AI? Maybe they'd be down 50% now.
The seniors today who have got to senior status by writing code manually will be different than seniors of tomorrow, who got to senior status using AI tools.
Maybe people will become more of generalists rather than specialists.
That’s putting it mildly. I think it’s going to be interesting to see what happens when an entire generation of software developers who’ve only ever known “just ask the LLM to do it” are unleashed on the world. I think these people will have close to no understanding of how computing works on a fundamental level. Sort of like the difference between Gen-X/millenial (and earlier) developers who grew up having to interact with computers primarily through CLIs (e.g., DOS), having to at least have some understanding of memory management, low-level programming, etc. versus the Gen-Z developers who’ve only ever known computers through extremely high level interfaces like iPads.
Sure, maybe you would have caught the bug if you wrote assembly instead of C. But the C programmer still released much better software than you faster. By the time you shipped v1 in assembly, the C program has already iterated 100 times and found product market fit.
Someone who is good at writing code isn't always good at making money.
we only have to look today at how different software quality is compared to the "old days" - when compilers were not as good, and people wrote in assembly by hand.
Old software were fast and optimized. Hand written assembly used minimal resources. Today, people write bloated electron webapps packaged into a bundle.
And yet, look who is surviving in the competitive land of software darwinian natural selection?
And large companies. The first half of my career was spent writing internal software for large companies. I believe it's still the case that the majority of software written is for internal software. AI will be a boon for these use cases as it will make it easier for every company big and small to have custom software for its exact use case(s).
Big cos often have the problem of defining the internal problems they’re trying to solve. Once identified they have to create organizational permission structures to allow the solutions. Then they need to stay on tasks long enough to build and use the software to solve the problem.
Only one of these steps is easily improved with AI.
The benefit of using off the shelf software is that many of the integration problems get solved by other people. Heck you may not even know you have a problem and they may already have a solution.
Custom software on the other hand could just breed more demand for custom software. We gotta be careful how much custom stuff we do lest it get completely out of hand
Not to mention that you'd need to integrate it with lots of other vibe-coded products. It can be great for some use cases for sure, though, but identifying them can be tricky, as big orgs are pretty terrible at formulating what they need clearly.
This would be strange, because all other technology development in history has taken things the exact opposite direction; larger companies that can do things on scale and outcompete smaller ones.
This would be strange, because all other technology development in history has taken things the exact opposite direction; larger companies that can do things on scale and outcompete smaller ones.
I don't think this has always been true.Youtube allowed many more small media production companies - sometimes just one person in their garage.
Shopify allowed many more small retailers.
Steam & cheap game engines allowed many more indie game developers instead of just a few big studios.
It likely depends on the stage of the tech development. I can see Youtube channels consolidating into a few very large channels. But today, there are far more media production companies than 30 years ago.
I think many developers, especially ones who come from EE backgrounds, grossly overestimate the number of people needed who understand what is going on underneath.
“Going on underneath” is a lot of interesting and hard problems, ones that true hackers are attracted to, but I personally don’t think that it’s a good use of talented people to have 10s or 100s of thousands of people working on those problems.
Let the tech geniuses do genius work.
Meanwhile, there is a massive need for many millions of people who can solve business problems with tech abstractions. As an economy (national or global), supply is nowhere close to meeting demand in this category.
LLMs just accelerated this trend.
Or magically 9900 more products or markets will be created, all of them successful?
I already feel spammed to death by desperate requests for my consumption as is.
Then companies won't need to spam you to convince you that you need something you don't. Or that their product will help you in ways it can't.
Once person companies will not have a 100 person marketing team trying to inject ads into ever corner of your life.
Because these one person companies will scale up everything with AI except marketing/advertising? Consider me skeptical.
But they could have a thousand-agent swarm connected via MCP to everything within our field of vision to bury us with ads.
It's been a long time since I read "The Third Wave" and up until 2026, not much has reminded me about its "Small is beautiful", and "rise of the prosumer" themes besides the content creator economy which is arguably the worst thing to ever happen to humanity's information environment, and LLM agent discussions.
This is exactly one of the things I find maddening at the moment. "Everyone" (except my actual friends) on social media is trying to sell me something.
Eg: I like dogs. It's becoming increasingly hard to follow dog accounts without someone trying to sell me their sponsor's stuff.
Now answer these questions:
And what will those people who make videos on Youtube do? Produce videos in uber-saturated categories?
Or magically 9900 more media channels will be created, all of them successful?
Many large production companies in the 2000s would have been extremely happy with that many views. They would have laughed you out of the building if you told them a single person could ever produce compelling enough video content and get at many viewers.
Claims 0.25% of channels makes any money at all. The amount that make a decent living is realistically even smaller, possibly < 0.1%.
To me the YouTube example seems to be the exact demonstration that markets saturate and market distribution is still a winner-takes-all kind of deal.
0.25% of how many?
Average size of YouTube channel team that makes money vs TV channel team in the 2000s?
Channels that make money consistently also have teams behind. Sure, probably they are smaller then TV studios, but TV studios do also other jobs compared to youtubers.
Anyway, these are the only numbers available. If there are numbers that show that masses of individuals can make a living in a market with so many competitors like YouTube I am happy to look at them. Until then, I will observe what is known for almost everything: a small % takes the vast majority of resources.
A quick search leads to different answer, but https://alanspicer.com/what-percentage-of-youtubers-make-mon... suggests that 0.25% of all YouTube channels makes any money (not good money, any money). Which means 99.75% earns 0$.
Basically I would flip the question and ask: if you could produce videos now very simply with AI and so could other 10000 people, how many of the new channels do you think will be successful? If anything, the YouTube example shows you exactly that it doesn't matter than 1000000 people now can produce content with low overhead, just a handful of them will be successful, because the market of companies available to spend money to sponsor through channels and the men hours of eyeballs on videos are both limited.
Talking about companies that just produce products, either you come up with something new (and create a new market), or you come up with something better (you take shares of an existing market). Having 10000 companies producing - say - digital editing software won't make suddenly increase by 10000x the number of people in need of digital editing software. Which means that among those 10000 there will be a few very good products which will eat all the market (like it is now), with the usual Paretian distribution.
The idea that many companies with smaller overhead can split the market evenly and survive might (and it's a big hopeful might) work on physical companies selling local and physical products (I.e., splitting the market geographically), but for software products I cannot even imagine it happening.
New markets are created all the time, and it's great if maybe smaller companies (or co-ops) could take over those markets rather than big corporations, but the way the market distribution happens I don't think will be affected. I don't see any reason why this should change with many more companies in the same market. I also don't think that 10000 new companies will create 10000 new markets, because that depends on ideas (and luck, and context, and regulation, etc.), not necessarily on resources available,
Now count how many TV stations there were in the 90s and 2000s.
This is just Youtube alone. What about streamers on Twitch? Youtube? Tiktokers? IG influencers?
Plenty of people are making money creating video content online.
The internet, ease and advancements in video editing tools, and cheap portable cameras all came together to allow millions of content creators instead of 100 TV channels.
I am not going to deny that YouTube (and all social media) created new markets. But how is this not an argument that shows that when N people suddenly do some activity, only a tiny minority is successful and gains some market share?
If tomorrow a product that is made by 3 companies will see competitors by 10000 1-man operations, maybe you will have 30 different successful products, or 100. 9900 of those 10000 will still be out of luck. I
YouTube is not an example of a market that being exposed to a flood of players gets shares somewhat equally between those players or that allows a significant number of the to survive with it. Nor is twitch or any of the other platforms.
> Or magically 9900 more products or markets will be created, all of them successful?
Yes. Products will become more tailored/bespoke rather than a lot of the one size fits all approach that is pervasive now.
Do you spec software for a variety of businesses?
I do.
It’s rare that one SaaS or software package does what the people paying want it to do. Either they have to customize internally (expensive and limited to larger orgs with a tech department) or Frankenstein a solution like Salesforce or WordPress with a lot of add ons. And even then, it’s not hitting all pain points.
Being able to spin up or modify an app cheaply and easily will be a massive boon for businesses.
To me it seems the reality works in the opposite way. Among the many products built, some will be successful and will swallow the whole market, like now with basically any software or SaaS product.
0. Sure, some products will be made in house. That said, being able to spec a product well is a skill that is not as common as some folks seem to think. It also assumes that an org is large enough to have a good internal dev team, which is both rare and relatively expensive.
1. It sloughs responsibility, which many folks want to do.
2. It allows for creation to be done not by committee and/or with less impact from internal politics.
3. It facilitates JIT product/tool development while minimizing costs.
That’s off the top of my head.
The realities of business often point to internal development not being ideal.
These tasks become your prompt once refined. I basically braindump to Claude, have it make tasks from my brain dump. Then I tell Claude to ask me clarifying questions, it updates the tasks and then I have Claude do market research for some or all tasks to see what the most common path is to solve a given problem and then update the tasks.
> the llm doesn't make good long term decisions
What could possibly go wrong, using something you know makes bad decisions, as the basis of your learning something new.
It's like if a dietician instructed a client to go watch McDonald's staff, when they ask how to cook the type of meals that have been recommended.
AI is great at exposing you to what you don’t even know you don’t know: your personal unknown unknowns, the complexity you’re completely unaware of.
Are you somehow equating basic multiplication to a "bad long term decision"?
> llms can get me started really fast. Basically it distills the time taken to research
I was saying that it speeds up research by exposing you to the things you don’t know that you don’t know.
Or the things that it has hallucinated, or just referenced incorrectly.
That's my point. You're talking about learning basic maths from a school teacher who isn't a Calculus expert, but the thread is talking about learning maths from a kid with ADHD who completes the homework before the teacher has finished describing what needs to be done, but sometimes returns homework in an invented language with references to Cthulhu throughout it.
https://newsletter.semianalysis.com/p/claude-code-is-the-inf...
Asking for "amazing" open source projects in this case is not asking out of genuine curiosity or want for debate, it is a rhetorical question asked out of frustration at the general trajectory of AI and who profits off of it -- namely the boot-wearers.
- https://github.com/simonw/sqlite-history-json
- https://github.com/simonw/sqlite-ast
- https://github.com/simonw/showboat - 292 stars
- https://github.com/simonw/datasette-showboat
- https://github.com/simonw/rodney - 290 stars and 4 contributors who aren't me or Claude
- https://github.com/simonw/chartroom
Noting the star counts here because they are a very loose indication that someone other than me has found them useful.
It does use transactions in the form of savepoints which means they can be nested: https://github.com/simonw/sqlite-history-json/blob/53e66b279...
Transactions are tested here: https://github.com/simonw/sqlite-history-json/blob/53e66b279...
I lead with sqlite-history-json because I think it's the most impressive of the bunch - it solves a difficult problem in an elegant way with code I would have been proud to write by hand.
I wouldn't call these toys either. If you want toys take a look at most of https://tools.simonwillison.net/ - these six are all real projects on GitHub with tests and documentation and release notes.
You rebutted by claiming 4% of open source contributions are AI generated.
GP countered (somewhat indirectly) by arguing that contributions don’t indicate quality, and thus wasn’t sufficient to qualify as “amazing AI-generated open source projects.”
Personally, I agree. The presence of AI contributions is not sufficient to demonstrate “amazing AI-generated open-source projects.” To demonstrate that, you’d need to point to specific projects that were largely generated by AI.
The only big AI-generated projects I’ve heard of are Steve Yegge’s GasTown and Beads, and by all accounts those are complete slop, to the point that Beads has a community dedicated to teaching people how to uninstall it. (Just hearsay. I haven’t looked into them myself.)
So at this point, I’d say the burden of proof is on you, as the original goalposts have not been met.
Edit: Or, at least, I don’t think 4% is enough to demonstrate the level of productivity GP was asking for.
4% for a single tool used in a particular way (many are out there using AI tools in a way that doesn't make it clear the code was AI authored) is an incredible amount. Don't see how you can look at that and see 'not enough'.
The vast majority of people using these tools aren't announcing it to the world. Why would they ? They use it, it works and that's that.
Just because people are shitting out endless slop code that they never bothered to throw a 2nd glance at doesn't mean it'sgood or that it's leading to better projects or tools, it literally just means people are pushing code out haphazardly . If I made a python script that everyone started using and all it did was create a repo, commit a README and push it every 5 seconds we'd be seeing billions of lines of code added! But none of it is useful in any way.
Same with AI, sure we're generating endless piles of code, but how much of it is actually leading to better software?
>If I made a python script that everyone started using and all it did was create a repo, commit a README and push it every 5 seconds we'd be seeing
1. Well you can't do that
2. Something like that won't register as Claude Code (or any other AI tool) usage anyway
3. Something like that won't come anywhere near 4%
But that's what these tools are doing, in a large number of cases? At least the end result is basically the same, like that Clawdbot or whatever name they've decided to try ride the coattails of guy who has 70k commits in the last few months that I saw being touted as an impressive feat on HN the other day. How much broken, unusable code exists within those 70k commits that ultimately would've had the same effect as if he had just pushed a `--allow-empty` commit thousands of times?
Now whatever, if it's people pushing slop into their own codebase that they own, more power to them, my issue stems from OSS projects being inundated with endless spam MR/PRs from AI hypesters. It's just making maintainer's lives more difficult, and the most annoying part of it all is that they have to put up with people who don't see the effort disparity between them prompting their chatbot to write up some broken bullshit vs the effort required for maintainers to keep up with the spam. It hurts the maintainers, it hurts genuine beginners who would like to learn and contribute to projects, it hurts the projects themselves since they have to waste what precious little time and resources they already have digging through crap, it hurts quite literally everyone who has to deal with it other than the selfish AI-using morons who just take a huge dump over everyone and spouts shit like "Well 4% of all code on Github is now AI-generated!" as if more of that is somehow a good thing.
I mean No not really. I'm not sure why you think that.
>How much broken, unusable code exists within those 70k commits that ultimately would've had the same effect as if he had just pushed a `--allow-empty` commit thousands of times?
How much stable usable code exists within those 70k commits ?
This is pretty much exactly why I said the original question was not a great ask. You have your biases. Show an example and the default response for some almost like a stochastic parrot is, 'Must be slop!". How do you know ? Did you examine it? No, you didn't. That's just what you want to believe so it must be true. It makes for a very frustrating conversation.
> amazing
Nobody moved the goal posts.
> If you want to help, more funding so we can pay more maintainers to deal with the slop (on top of everything we do already) is the only viable solution I can think of
https://www.pcgamer.com/software/platforms/open-source-game-...
This is exactly the wrong approach! Funnel even more money away from productive tasks and into AI? Madness![1]
The only viable solution is being quick with a banhammer - maybe someone should start up a spamhaus type list of every github user who submitted AI slop.
Force them to burn these accounts on the very first spam.
------------
[1] Imagine if we chose this approach to deal with spam - we ask people for more money to hire a warm body to individually verify each email. Do you think spam would be the solved problem it is today?
Also small businesses aren't going to publish blog posts saying "we saved $500 on graphic design this week!"
And make your own brushes.
Before the printing presses came along, putting up flyers was not even imaginable.
Signs for businesses used to hand carved.
Then printed. A store sign was still produced by a team of professionals, but small businesses coils reasonably afford to print a sign. Not often updated, but it existed.
Then desktop publishing took off. Now lone graphic designers could design and send work off to a print shop. Small businesses could now afford regularly updated menus, signage, and even adverts and flyers.
Now small businesses can make their own creatives. AI can change stylesheets, write ad copy, and generate promotional photos.
Does any of this have the artistry of hand carved signs from 600 years ago? Of course not.
But the point is technology gives individuals control.
People have been painting with red and yellow ochre and soot for at least 50K years for sure, and probably several hundred thousand years in truth. You don't need a brush, you have fingers or a twig.
The walls on the streets of Pompeii are full of advertising -- they had an election going on and people just scribbled slogans and such on walls. You don't need flyers lol.
The idea that signs or advertising was "artistry" is deeply ahistorical. The reason old stuff looks real fancy is because labor was extremely cheap and materials were expensive.
Compare those to the pigments used (mixed up!) by professional painters, and then to what printers could make.
If you wanted to paint fine art in the 1400s you were possibly making your own canvases, your own paint brushes, and your own paints.
And on top of that you had to be a skilled painter!
> The walls on the streets of Pompeii are full of advertising -- they had an election going on and people just scribbled slogans and such on walls. You don't need flyers lol.
The American revolution included a lot of propaganda courtesy of printing presses and some very rich financers who had a vested interest in a revolution occuring.
Pamphlets everywhere. It is one thing to scribble on a wall, it is another to produce messages at a mass scale.
That sense of scale has been multiplied yet again by AI.
..a month
..multiplied by how many small businesses globally?
Once you do have a billion dollar product protecting it requires spending time, money and people to keep running. Because building a new one is a lot more effort than protecting existing one from melting.
Once you have revenue you have downside to protect. Pre-revenue the worst that can happen is that you have to start again knowing more than you did.
But I like your and OP's analogy. Also, the productivity claims are coming from the guys in main memory or even disk, far removed from where the crunching is taking place. At those latency magnitudes, even riding a turtle would appear like a huge productivity gain.
Privacy is non existent, every word said and message sent at the office is recorded but the benefits we saw were amazing.
He was correct though. For example, I’ve been waiting over a month for another team to set me up so I can test something they wanted me to develop. I’ve followed up multiple times. AI coding tools aren’t going to solve my blocker.
Meetings are work, as much as IPC and network calls are work. Just because they're not fun, or what you like to do, it doesn't mean they're any less of a work.
I think you're analyzing things from a tactical perspective, without considering strategic considerations. For example, have you considered that it might not be desirable for CPUs to be just fast, or fast at all? is CISC faster than RISC? different architectural considerations based on different strategic goals right?
If you're an order picker at an amazon warehouse, raw speed is important. being able to execute a simpler and more fixed set of instructions (RISC), and at greater speed is more desirable. if you're an IT worker, less so. IT is generally a cost-center, except for companies that sell IT services or software. if you're in a cost center, then you exist for non-profit-related strategic reasons, such as to help the rest of the company work efficiently, be resilient, compete, be secure. Some people exist in case they're needed some day, others are needed critically but not frequently, yet others are needed frequently but not critically. being able to execute complex and critical tasks reliably and in short order is more desirable for some workers. Being fast in a human context also means being easily bored, or it could mean lots of bullshit work needs to be invented to keep the person busy and happy.
I'd suggest taking that compsci approach but considering not just the varying tasks and workloads, but also the diversity of goals and user cases of users (decision makers/managers in companies). There are deeper topics with regards or strategy and decision making surrounding the state machines of incentives and punishments, and decision maker organization (hierarchical, flat, hub-and-spoke,full-mesh,etc..).
That said, often, meetings are much more efficient means of syncing information than slack/chat or emails. call it "real-time active communication with rich context" if it sounds more technical. you can communicate in voice tones, body language, timing,etc.. what you can't using other means. and communicate doesn't mean just talk or listen for the sake of it, it can me brainstorm, understand requirements and expectations better, prevent misunderstandings and other wasted effort.
In my experience, things that exist as patterns like this in systems are always important, but it's also important to use them as intended, and not abuse them excessively.
Simply extracting the most value out of individual contributors isn't typically the goal of white collar management. as in my earlier example, you won't see order pickers at amazon warehouses attend meetings all day. their time at work is valued differently than a white collar workers'.
- reviews for code
- asking stakeholders opinions
- SDLC latency (things taking forever to test)
- tickets
- documentations/diagrams
- presentations
Many of these require review. The review hell doesn't magically stop at Open source projects. These things happen internally too.
Do they also make you write your own performance review and set your own objectives?
This is basically the same story I have heard both my own place of employment and also from a number of friends. There is a "need" for AI usage, even if the value proposition is undefined (or, as I would expect, non-existent) for most businesses.
Not to get off on a tangent but this has got to be a "tell" for how much a company is managed by formula and how much it's actually got thinking people running things. Every time I've had to write my own review I fill out the form with some corporatese bullshit, my supervisor approves it and adds some more bullshit, it disappears into HR and I never hear anything about it until it's time for the next review, and it starts over again. There isn't even reference to any of my "objectives" from the last review, because that review has simply disappeared.
But I'm sure some HR exec is checking boxes for following "best practices" in employee evaluation.
In my first year I didn’t know any better, so I tried to set myself some actual objectives (learn to use XYZ, improve test coverage by X%, measurable stuff that would actually help).
Fortunately my manager showed me how to do it correctly, so now my goals are to “differentiate with expertise” and to “empower through better solutions”.
Every year I open up the self-review, grade myself a 5/5 on these absurd, unmeasurable goals, my manager approves it, and it disappears off somewhere into the layers and layers of ever-higher management where nobody cares to look at it.
I think LinkedIn is in the dataset, right?
1. too old/expensive
2. not using AI
3. using AI but not productive
4. productive using AI but not within AI budget
5. reduce AI budget and GOTO 3
Dateline ~2010. Location: NYC Why:Indian outsourced shops.
Now the zinger, dear hn, is this: He actually said to us (we ran a more boutique consulting firm) that "everything has to be done 3 times" and "their work is crap". But "we're getting rid of this floor".
That, imho, was due to geopolitical machinations of inducing India to become part of the West. The immediate equation of "money for quality work" wasn't working but the 'our higher ups' had more grand plans and sacrificing and gutting the IT industry in US was not a problem.
So, given the incentives these days, do not remotely pin your hopes on what these CEOs are saying. It means nothing whatsoever.
Other white collar business/bullshit job (ala Graeber) work is meeting with people, “aligning expectations”, getting consensus, making slides/decks to communicate those thoughts, thinking about market positioning, etc.
Maybe tools like Cowork can help to find files, identify tickets, pull in information, write Excel formulas, etc.
What’s different about coding is no one actually cares about code as output from a business standpoint. The code is the end destination for decided business processes. I think, for that reason, that code is uniquely well adapted to LLM takeover.
But I’m not so sure about other white-collar jobs. If anything, AI tooling just makes everyone move faster. But an LLM automating a new feature release and drafting a press release and hopping on a sales call to sell the product is (IMO) further off than turning a detailed prompt into a fully functional codebase autonomously.
this software (which i am not related to or promoting) is better at investment planning and tax planning than over 90% of RIAs in the US. It will automate RIA to the point that trading software automated stock broking. This will reduce the average RIA fee from 1% per year to 0.20% or even 0.10% per year just like mutual fund fees dropped in the early 00s
more expensive silly companies will exist, but the cheap ones get the scale. SP500 index funds have over 1 trillion in the top 3 providers. cathy wood has like 6-7 billion.
BNYMellon is the custodian of $50 trillion of investment assets. robinhood has $324bn.
silly companies get the headlines though
That use case is definitely delegated to LLMs by many people. That said, I don't think it translates into linear productivity gains. Most white collar work isn't so fast-paced that if you save an hour making slides, you're going to reap some big productivity benefit. What are you going to do, make five more decks about the same thing? Respond to every email twice? Or just pat yourself on the back and browse Reddit for a while?
It doesn't help that these LLM-generated slides probably contain inaccuracies or other weirdness that someone else will need to fix down the line, so your gains are another person's loss.
But if you get deep into an enterprise, you'll find there are so many irreducible complexities (as Stephen Wolfram might coin them), that you really need a fully agentically empowered worker — meaning a human — to make progress. AI is not there yet.
If you weren’t doing much of that before, I struggled to think of how you were doing much engineering at all, save some more niche extremely technical roles where many of those questions were already answered, but even still, I should expect you’re having those kinds of discussions, just more efficiently and with other engineers.
I'd suspect the kind that's going away.
I don’t agree with it or believe it’s smart but it’s the world we live in
The vast majority of software engineers in the world. The most widespread management culture is that where a team's manager is the interface towards the rest of the organization and the engineers themselves don't do any alignment/consensus/business thinking, which is the manager's exclusive job.
I used to work like that and I loved it. My managers were decent and they allowed me to focus on my technical skills. Then, due to those technical skills I'd acquired, I somehow got hired at Google, stayed there nearly a decade but hated all the OKR crap, perf and the continuous self-promotion I was obliged to do.
* meeting with people, yes, on calls, on chats, sometimes even on phone
* “aligning expectations”, yes, because of the next point
* getting consensus, yes, inevitably or how else do we decide what to do and how to do it?
* making slides/decks to communicate that, not anymore, but this is a specific tool of the job, like programming in Java vs in Python.
* thinking about market positioning, no, but this is what only a few people in an organization have agency on.
* etc? Yes, for example don't piss off other people, help custumers using the product, identify new functionalities that could help us deliver a better product, prioritize them and then back to getting consensus.
Isn't like half of our industry just churning out JS file after JS file to yet again change how facebook looks?
huh? maybe im in the minority, but the thinking:coding has always been 80:20 spend a ton of time thinking and drawing, then write once and debug a bit, and it works
this hasnt really changed with Llm coding either, except that for the same amount of thinking, you get more code output
I interview a lot of people, and I've seen people who are astoundingly good at micro-systems, very complex regexes, etc. white struggling massively with system design. And vice versa. People have different talents.
But, in my experience, AI will vastly improve the success of the developer who's better at orchestration, architecture, and system design than the developer who's very good at tiny micro-system type of work. Yes, there is still a need for someone who can read and understand regexes... but is there anywhere near as much of a need as before? Not at all.
Now. there are very many dual threats, and most truly senior engineers are both. These people now have an even bigger leg up, because they have an understanding of system mechanics + the superpower of Claude Code/etc. They don't have to waste as much time on boilerplate and raw implementation, and yet they can check the output of their input to the AI. They are also probably better equipped to build testing harnesses, etc., that adapt well to agentic use.
I have, incidentally, held the title of "Principal Software Architect", designing distributed systems with kubernetes, and I will say this about architects: if they aren't immersed in the day to day code, they suck at their job. If you're too removed from the constraints you can't be effective at that job. I have however worked with "architects" that refused to get their hands dirty, and it was always miserable.
I'll give you an example: at one time I essentially re-implemented the behavior of a WeakMap in JS because I didn't know that the language feature existed. AI is much better at implicitly "knowing" these things because it can model the entire language and possible token-space much better than humans can. That is something I always struggled with; my long-term memory is not great.
I think that's very different from remembering how to write a basic function.
It doesn’t capture everyone’s experience when you say thinking is the smaller part of programming.
I don’t even believe a regular person is capable of producing good quality code without thinking 2x the amount they are coding
(a) thinking about, and deciding upon, what will be done, and
(b) the thinking that is required during implementation.
{type (a)} was always the majority of my time, but {type (b)} consumed a lot of effort, especially for languages or syntaxes I wasn't very familiar with. {type (b)} consumes very little of my time now.WHOAH WHOAH WHOAH WHOAH STOP. No coder I've ever met has thought that thinking was anything other than the BIGGEST allocation of time when coding. Nobody is putting their typing words-per-minute on their resume because typing has never been the problem.
I'm absolutely baffled that you think the job that requires some of the most thinking, by far, is somehow less cognitively intense than sending emails and making slide decks.
I honestly think a project managers job is actually a lot easier to automate, if you're going to go there (not that I'm hoping for anyone's job to be automated away). It's a lot easier for an engineer to learn the industry and business than it is for a project manager to learn how to keep their vibe code from spilling private keys all over the internet.
OK, to quote you: WHOAH WHOAH WHOAH WHOAH STOP!
You've made a lot of assumptions.
I'm not saying that coding is not thinking. What I'm saying is this:
There is a difference between:
(a) thinking about, and deciding upon, what will be done, and
(b) the thinking that is required during implementation.
In my experience, coding is at least 50/50 (even for the best developer) in the sense that figuring out how to structure and fix your code {type (b)} used to require very deep thinking. But then the other thinking time was spent on your system design/architecture {type (a)}, and not debugging type errors, etc.AI has already changed that split. If you have a good test harness and problem definition, you can throw Codex at a really massive task and have it do quite well at the finer details of implementation.
Other white-collar office work, as stupid as it may be, will be a lot harder to automate because it is primarily the "thinking about what will be done" {type (a)} kind of work and not the "thinking that is done during implementation" {type (b)} kind of work.
If you haven't seen what I mean by "enterprise office work" it may be hard to grasp what I'm talking about... But thinking that people are just doodling around making slide decks or writing shitty emails is the wrong mental model for the breadth of non-technical work available in a large company.
I also don't expect AI to replace software engineers any more than white-collar business people.
But what I'm saying is the work of "mere implementation" is now happening pretty quickly with AI tooling.
Most white-collar work is not "mere implementation" but rather the yak shaving and spec definition that precedes "mere implementation" — and this includes software development. For that reason, it will be harder to fully automate.
As a CEO I see it as a massive clog up of vast amounts of content that somebody will need to check. A DDoS of any text-based system.
The other day I got a document of 155 pages in Whatsapp. Thanx. Same with pull requests. Who will check all this?
The answer to that, for some, is more AI.
I had a peer explain that the PRs created by AI are now too large and difficult to understand. They were concerned that bugs would crop up after merging the code. Their solution, was to use another AI to review the code... However, this did not solve the problem of not knowing what the code does. They had a solution for that as well... ask AI to prepare a quiz and then deliver it to the engineer to check their understanding of the code.
The question was asked - does using AI mean best-practices should no longer be followed? There were some in the conversation who answered, "probably yes".
> Who will check all this?
So yeah, I think the real answer to that is... no one.
Even more troublesome is importing libraries. I have no idea which ones are AI generated and they are better and better at hiding their original authors.
Making it easier/better just means more/higher quality “worthless” work is performed. The incentives in the not-directly -productive parts of organizations are to keep busy and maintain a stream of signals of productivity. For this , AI just raises the bar. The 25% of the work that -is- important to producing economic value just gets reduced to 15%.
The workforce in large orgs that is most AI adjacent is already idling along in terms of production of direct economic value. Making them 10x more productive in nonproductive work will not impact critical metrics in a short timeframe.
It’s worth noting that these “not directly productive” activities actually can (and often do) produce value, eventually. Things like brand identity, culture, and meta-innovation, vision (search-space) are intangibles that present as cost centers but can prove invaluable in longer timescales if done right.
These are the people "shocked" when they are displaced.
Whats taught in economics textbooks doesnt always reflect reality - ha.
The manager wants a large team. The shareholder who ultimately employs the manager but does not control operations does not want that of course.
Hmm.
Figure A6 on page 45: Current and expected AI adoption by industry
Figure A11 on page 51: Realised and expected impacts of AI on employment by industry
Figure A12 on page 52: Realised and expected impacts of AI on productivity by industry
These seem to roughly line up with my expectations that the more customer facing or physical product your industry is, the lower the usage and impact of AI. (construction, retail)
A little bit surprising is "Accom & Food" being 4th highest for productivity impact in A12. I wonder how they are using it.
Could it be that employers are not seeing the difference because most employees are doing something else with the time they've saved by using AI?
There's been massive wage stagnation, benefits are crap, they play games with PTO. Most people I talk to who use AI as a part of their workflow are taking advantage of something nice that has come their way for a change.
“Autofishers” are large boats with nets that bring in fish in vast quantities that you then buy at a wholesale market, a supermarket a bit later, or they flash freeze and sell it to you over the next 6-9 months.
Yet there’s still a thriving industry selling fishing gear. Because people like to fish. And because you can rarely buy fish as fresh as what you catch yourself.
Again, it’s not a great analogy, but I dunno. I doubt AGI, if it does come, will end up working the way people think it will.
spoiler, it's not
The moment of realisation happen for a lot of normoid business people when they see claude make a DCF spreadsheet or search emails
claude is also smart because it visually shows the user as it resizes the columns, changes colours, etc. Seeing the computer do things makes the normoid SEE the AI despite it being much slower
Replace excel and office stuff with ai model entirely then people will pay attention.
iterating over work in excel and seeing it update correctly is exactly what people want. If they get it working in MSWord it will pick up even faster.
If the average office worker can get the benefit of AI by installing an add-on into the same office software they have been using since 2000 (the entire professional career of anyone under the age of 45), then they will do so. its also really easy to sell to companies because they dont have to redesign their teams or software stack, or even train people that much. the board can easily agree to budget $20 a head for claude pro
the other thing normies like is they can put in huge legacy spreadsheets and find all the errors
Microsoft365 has 400 million paid seats
Do you work extra hard to be this arrogant or does it come naturally?
For my team at least, the productivity boost is difficult to quantify objectively. Our products and services have still tons of issues that AI isn't going to solve magically.
It's pretty clear that AI is allowing to move faster for some tasks, but it's also detrimental for other things. We're going to learn how to use these tools more efficiently, but right now, I'm not convinced about the productivity gain.
What improvements have you noticed over that time?
It seems like the models coming out in the last several weeks are dramatically superior to those mid-last year. Does that match your experience?
What I'm still skeptical about is how much more productive it makes us. In my case, coding is maybe 50% of my job, and I work on complex and novel systems. The agent gives me the illusion I don't need to think anymore, but it's not the case. Agents slow me down in many cases too, I'm not learning and improving as I used to.
FWIW Gemini inside Google apps is just as bad.
LLMs are impressive and flexible tools, but people expect them to be transformative, and they're only transformative in narrow ways. The places they shine are quite low-level: transcription, translation, image recognition, search, solving clearly specified problems using well-known APIs, etc. There's value in these, but I'm not seeing the sort of universal accelerant that some people are anticipating.
Give it a year or two and let things settle down and (assuming the music is still playing at that time) you might see more dinosaurs start to wander this way.
But really, are CEO's the best people to assess productivity? What do they _actually_ use to measure it? Annual reviews? GTFO. Perhaps more importantly, it's not like anything a C-level says can ever be taken at face value when it involved their own business.
The fee-earners had KPIs tied to the sales pipeline, from leads to contracts to work completed on fixed contracts or hours billed on variable-rate contracts. It's relatively easy to measure improvements here. Though it's harder to distill the causes of that and tie it to LLMs.
The fee-burners like in IT, legal, compliance, marketing, finance, typically had KPIs tied to the department objectives. This stuff is a LOT more subjective and a lot more prone to manipulation (goodhart's law). But if you spend 60 hours a week on work in such a department, you tend to have a pretty good idea if things are speeding up or not at all. In a department I was involved in there was a lot of KYC that involved reviewing 300+ pages per case, we tracked case workload per person per day, as well as success rates (percentage of case reviews completed correctly), and could see meaningful changes one could attribute to LLM use.
Agreed though that I'm more interested in a few case studies in detail to understand how they actually measured productivity.
Steve Jobs is the only CEO of a large firm that I can re-call that always remained intimately involved.
The firmwide AI guru at my shop who sends out weekly usage metrics and release notes started mentioning cost only in the last few weeks. At first it was just about engaging with individual business heads on setting budgets / rules and slowing the cost growth rate.
A few weeks later and he is mentioning automated cost reporting, model downgrading and circuit breaking at a per-user level. The daily spend where you immediately get locked within 24 hours is pretty low.
Once the tools help the AI to get feedback on what its first attempt got right and wrong, then we will see the benefits.
And the models people use en masse - eg. free tier ChatGPT - need to get to some threshold of capability where they’re able to do really well on the tasks they don’t do well enough on today.
There’s a tipping point there where models don’t create more work after they’re used for a task, but we aren’t there yet.
you’d be surprised… the largest IP in the majority of cases is the codebase itself. once that hurdle was crossed the rest is easy decision
And the biggest irony is that the "scariest" projects we had at our university ended up being maybe 500-1000 lines of code, things really must go back to hands on programming with real time feedback from a teacher. LLM's only output what you ask and won't really suggest concepts used by professionals unless you go out of your way to ask for it, it all seems like a vicious cycle even though meaningful code blocks can range along 5 to 100 lines which. When I use LLM's I just get information burn out trying to dig through all that info or code
However, there's another factor. The J-curve for IT happened in a different era. No matter when you jumped on the bandwagon, things just kept getting faster, easier, and cheaper. Moore's law was relentless. The exponential growth phase of the J-curve for AI, if there is one, is going to be heavily damped by the enshitification phase of the winning AI companies. They are currently incurring massive debt in order to gain an edge on their competition. Whatever companies are left standing in a couple of years are going to have to raise the funds to service and pay back that debt. The investment required to compete in AI is so massive that cheaper competition may not arise, and a small number of (or single) winner could put anyone dependent on AI into a financial bind. Will growth really be exponential if this happens and the benefits aren't clearly worth it?
The best possible outcome may be for the bubble to pop, the current batch of AI companies to go bankrupt, and for AI capability to be built back better and cheaper as computation becomes cheaper.
> My own updated analysis suggests a US productivity increase of roughly 2.7 per cent for 2025. This is a near doubling from the sluggish 1.4 per cent annual average that characterised the past decade.
good for 3 clicks: https://giftarticle.ft.com/giftarticle/actions/redeem/97861f...
As tech become available to help reduce your costs and drive up your profit, the same tech also reduces your competitor's costs and perhaps lets more competitors into the market. This drives down your product prices and reduces your profit.
So you invest but see no increase in productivity, but if you don't do it - you're toast.
p.s. @dang doesn't work reliably - hn@ycombinator.com is the way to get a message delivered
Of course this doesn't take into account people who just pay to play around and learn, non professional use cases, or a few other things, but it's a rough ballpark estimate.
Assuming the above, current AI models would only increase the productivity for most workplaces by a relatively small amount, around 10-200 € per employee per month perhaps. Almost indistinguishable compared to salaries and other business expenses.
Unless I'm misunderstanding, shouldn't someone rational want to pay where (value - cost) is highest, opposed to increasing cost to the point where it equals value (which has diminishing returns)?
A $40 subscription creating $1000 worth of value would be preferred over a $200 subscription creating $1100 of value, for instance, and both preferred over a $1200 subscription creating $1200 of value.
I was more so limiting myself to the simpler heuristic where people only pay roughly what they personally think something is worth, and not significantly more/less regardless of the options. But of course, as you've pointed out, in real life the options available really do matter, and someone might decline a 200:1200 trade if there are even more lopsided options available. It does complicate the though experiment somewhat if you try to take this into account.
I am glad to see articles like this that evaluate impact, but I wish the following would get more public interest:
With LLMs we are chasing sort-of linear growth in capability at exponential cost increases for power and compute.
Were you mad when the government bailed out mis-managed banks? The mother of all government bailouts might be using the US taxpayer to fund idiot companies like Anthropic and OpenAI that are spending $1000 in costs to earn $100.
I am starting to feel like the entire industry is lazy: we need fundamental new research in energy and compute efficient AI. I do love seeing non-LLM research efforts and more being done with much smaller task-focused models, but the overall approach we are taking in the USA is f$cking crazy. I fear we are going to lose big-time on this one.
Personally I have noticed strange effects, where I previously would have reached for a software package to make something or solve an issue, its now often faster for me to write a specific program just for my use case. Just this weekend I needed a reel with a specific look to post on instagram but instead of trying to use something like after effects, i could quickly cobble together a program that was using css transforms that outputted a series of images I could tie together with ffmpeg.
About a month ago I was unhappy with the commercial ticketing systems, they were both expensive and opaque so I made my own. Obviously for a case like that you need discipline and testing when you take peoples money, so there was a lot of focus on end to end testing.
I have a few more examples like this, but to make this work you need to approach using LLMs with a certain amount of rigour. The hardest part is to prevent drift in the model. There are a certain number things you can do to make the model grounded in reality.
When the tool doesn’t have a reproducer, it’ll happily invent a story and you’ll debug the story. If you ground the root cause in for example a test, the model can get context enough to actually solve the problem.
Another issue is that you need to read and understand code quickly, but its no real difference from working with other developers. When tests are passing I usually do a PR to myself and then review as I usually would do.
A prerequisite is that you need tight specs, but those can also be generated if you are experienced enough. You need enough domain intuition to know what ‘done’ means and what to measure.
Personally I think the bottleneck will go from trying to get into a flow state to write solutions to analyze the problem space and verification.
Lots of these project have a lifespan of a week and will never ever be maintained. When you pour blood and sweat in a projet you get attached to it, when you vibe code it in an afternoon and it's not and instant hit you move on to the next one.
When you actually talk to people about what they do there are often many, many nuances, micro-events, micro-decisions and micro-actions in their work. This is why it can take days/weeks/months to completely train a new person for a job.
This level of detail is barely documented - anywhere. There is a huge amount of information buried in workflows that AI has barely had access to for training. A lot of this is more in the realm of world models, rather than LLMs.
So imagine trying to use AI to improve these workflows it knows so little about. Then imagine AI trying to reinvent them across an organization.
We find these use cases where AI provides great value - totally true - but these barely scratch the surface of what goes on.
https://www.wsj.com/video/erik-brynjolfsson-productivity-is-...
Until the handoff tax is lower than the cost of just doing it yourself, the ROI isn't going to be there for most engineering workflows.
That multiple AI agents can now churn out those lines relatively nearly instantly, and yet project velocity does not go much faster, should start to make people aware that code generation is not actually the crucial cost in time taken to deliver software and projects.
I ranted recently that small mob teams with AI agents may be my view of ideal team setup: https://blog.flurdy.com/2026/02/mob-together-when-ai-joins-t...
1) The models like us have finite context windows and intelligence, even with good engineering practices system complexity will eventually slow them down.
2) At the moment at least, the code still needs to be reviewed and signed off by us and reading someone else's code is usually harder than writing it.
I am after the automated PR agents have all passed a PR I tend to let Claude Code and Codex give me a summary, with an MCP skill to read the requirement story. I trust their ability to catch edge cases and typos more than me. I just check the general structure of the PR
> I trust their ability to catch edge cases and typos more than me.
Given the vendors EULAs etc, if poop really hits the fan with released code, then how is that likely to sound if the lawyers get involved?
> Are people still reading PRs in detail manually?
Ultimately it all depends on circumstance and appetite for risk, but yes many/most places still manually checking releases.
Also flame-wars can be autofed ad nauseam now, so there’s going to be less and less interest in engaging into them. In an act of desperation, idle trolls will turn to tasks tracked by KPIs.
- measuring productivity
- adapting to change
This article just reinforces that. Past a certain headcount, executives have little to no understanding of what IC day-to-day is like.
AI tooling doesn't fix the bureaucracy the c-suite helped to create.
My team has gained a reputation of being some sort of firefighting crew.
We are being called by PMs when projects are failing, usually engineering-data and engineering-adjacent stuff. (Mechanical/Electrical).
We automate the heck out the processes, using a mix of AI processing, RAGs, and AI assisted coding.
We rescue the projects. Finish ahead of schedule. Make fewer mistakes. We gain additional scope. We win new projects. We bring new clients.
But when higher ups ask the people we helped about productivity gains, the most generous will say stuff like "it takes as long to review as it takes to do things manually", "They really helped on {inconsequential part of the deliverable}"
If the that is the takeaway these people were taking, they would incredibly misled. Luckily for me, I have people who deal with the politics, while my team can focus on delivery.
Our reputation keeps growing, and we keep delivering faster. The heads of the departments we work with love us, the middle rank who were doing the laborious crap, maybe not so much.
CEOs are now on the downside of the hype curve.
They went from “Get me some of that AI!” after first hearing about it, to “Why are we not seeing any savings? Shut this boondoggle down!” now that we’re a few years into bubble, the business math isn’t working, and they only see burning piles of cash.
I don't have a point, just that it's an unlikely unity.
Of course, trying to automate with chatGPT 4o was stupid. Trying to automate with Sonnet 4.6 will work better. Trying to automate with the models a year from now will work all the better.
To believe we are going to stop and go back to 2019 at this stage is seriously delusional.
I wish it were true. I would love to go back to 2019 but we obviously are not. We never go backwards.
If you've already got a very effective team with clear vision/goals, this technology will almost certainly help to some degree.
If you've got a sinking ship of a business, this technology will likely drag you down faster.
You always have to work backward from the customer into the technology. AI will never change that. I've found myself waffling on advice to some clients regarding AI because whether or not they can effectively leverage it depends more on what the people in the business are willing to do than what the technology can do.
Access to capital for everyone else is dropping. And the US economy is being managed by chaos monkeys, causing all kinds of supply chain disruptions. Oligopolies in almost every market are increasingly jacking up prices above market equilibrium rates as they are emboldened by a corrupted FTC.
Despite what Peter Thiel may have led you to believe, Monopolies are not healthy for an economy in aggregate.
Of course the economy is slowing.
filling in pdf documents is effectively the job of millions of people around the world
So you lose a lot of benefits to the time sync, but since people tend to have their eye glaze over when the correction rate is low, you may still miss the 2% anyway.
This is going to put a stop to a lot of ideas that sound reasonable on paper.
"Perfect! Let's delve into the problem with the engine. Based on the symptoms you describe, the likely cause is a blown head gasket..."
The non code parts (about 90% of the work) is taking the same amount of time though.
EDIT: I am on a mobile device and don’t have a reference handy but there have been good papers on RAG scaling issues - basically the embedding space gets saturated (too many document chunks cluster in small areas of the embedding space), if my memory is correct.
I have a sweet spot for using just Emacs, no other IDEs except very occasional use of AntiGravity. For a particular fun project or researching how to use generative models in applications, I like to start low: see if a small local model with appropriate tool use or agentic library will get the job done, if not move up to using something cheap like gemini-3-flash, and only if none of these approaches work, then use an expensive model.
I was advising a friend’s company last month on their application that effectively uses LLM models, but I was blown away by their zeal to spend lots of money on tokens.
I do think we are on the verge of something tho. Once the compounding effect happens in the world of atoms (recursive robotics), it's over.
Quickly slapping "AI features" on a bunch of existing products -- like almost every SW company seems to have done in an effort to appear "on the cutting edge" -- accomplishes almost nothing.
Unfortunately I think most of the stuff they make will be shit, but they will build it very productively.
I predict a golden age for experienced developers! There will be an uncountable number of poorly designed apps with scaling issues. And many of them will be funded.
This is not good. When all that matters is how viral your app is, people no longer compete on features and quality of life.
Yes, I won’t let the door hit me in the ass on the way out…
* AI is doing real work
* Humans using AI don't seem to get more done with AI than without
There is a huge economic pressure to remove humans and just let the AI do the work without them as soon as possible.
Or even the simple utility of having a chatbot. They’re not popular because they’re useless
Which to me says it’s more likely that people under estimate corporate inertia
What is Ai missing that will make it useful to everyone?
[0]: See page 2: https://www.nber.org/system/files/working_papers/w34836/w348...
I bet many CEO PA are using AI for many tasks. It's typically a role where AI is very useful. Answering emails, moving meetings around, booking and buying a bunch of crap.
Like for instance you want to tell a coworker his work is shit, but don't know how to put it in a way that's not going hurt them or make you look like an asshole.
I know people who'd potentially spend hours on a single email like this.
Yeah, if your Fortune 500 workplace is claiming to be leveraging AI because it has a few dozen relatively tech illiterate employees using it to write their em dash/emoji riddled emails about wellness sessions and teams invites for trivia events… there’s not going to be a noticeable uptick in productivity.
The real productivity comes from tooling that no sufficiently risk adverse pubco IS department is going to let their employees use, because when all of their incentives point to saying no to installing anything ever, the idea of giving the permissions required for agentic AI to do anything useful is a non-starter.
Maybe this bothers me more than it should.
There are some real changes in day to day software development. Programmers seem to be spending a lot of time prompting LLMs these days. Some more than others. But the trend is pretty hard to deny at this point. That snowballed in just 6-7 months from mostly working in IDEs to mostly working in Agentic coding tools. Codex was barely usable before the summer (I'm biased to that since that is what I use but it wasn't that far behind Claude Code). Their cli tool got a lot more usable in autumn and by Christmas I was using it more and more. The Desktop app release and the new model releases only three weeks ago really spiked my usage. Claude Code was a bit earlier but saw a similar massive increase in utility and usability.
It is still early days. This report cannot possibly take into account these massive improvements that hav been playing out over essentially just the last few months. This time last year, Agentic coding was barely usable. You had isolated early adopters of Claude Code, Cursor, and similar tools. Compare to what we have now, these tools weren't very good.
In the business world things are delayed much more. We programmers have the advantage that many/most of our tools are highly scriptable (by design) and easy to figure out for LLMs. As soon as AI coders figured out how to patch tool calling into LLMs there was this massive leap in utility as LLMs suddenly gained feedback loops based on existing tools that it could suddenly just use.
This has not happened yet for the vast majority of business tools. There are lots of permission and security issues. Proprietary tools that are hard to integrate with. Even things like wordprocessors, spreadsheets, presentation tools, and email/calendar tools remain poorly integrated. You can really see Apple, MS, and Google struggle with this. They are all taking baby steps here but the state of the art is still "copy this blob of text in your tool". Forget about it respecting your document theme, or structure. Agentic tool usage is not widely spread outside the software engineering community yet.
The net result is that the business world still has a lot of drudgery in the form of people manually copying data around between UIs that are mostly not accessible to agentic tools yet. Also many users aren't that tool savvy to begin with. It's unreasonable to expect people like that to be impacted a lot by AI this early in the game. There's a lot of this stuff that is in scope for automating with agentic tools. Most of it is a lot less hard than the type of stuff programmers already deal with in their lives.
Most of the effects this will have on the industry will play out over the next few years. We've seen nothing yet. Especially bigger companies will do so very conservatively. They are mostly incapable of rapid change. Just look at how slow the big trillion dollar companies are themselves with eating their own dog food. And they literally invented and bootstrapped most of this stuff. The rest of the industry is worse at this.
The good news is that the main challenges at this point are non technical: organizational lag, security practices, low level API/UI plumbing to facilitate agentic tool usage, etc. None of this stuff requires further leaps in AI model quality. But doing the actual work to make this happen is not a fast process. From proof of concept to reality is a slow process. Five years would be exceptionally fast. That might actually happen given the massive impact this stuff might have.
“Oh cool, copilot is in excel! I’m going to ask it a question about the data in the spreadsheet that it’s literally appearing beside natively in-app, or for help with a formula!”
“Wait what, it’s saying it can’t see anything or read from the currently displayed worksheet? Why is it inside the application then? Why would I want an outdated version of ChatGPT with no useful context or ability to read/do anything inside all my Office applications?”
Which ones? OpenAI? Microsoft? Anthropic?
Then I started working on some basic grpc/fullstack crap that I absolutely do not care about, at all, but needs to be done and uses internal frameworks that are not well documented, and now Claude is my best friend at work.
The best part is everyone else’s AI code still sucks, because they ask it to do stupid crap and don’t apply any critical thinking skills to it, so I just tell AI to re-do it but don’t fuck up the error handling and use constants instead of hardcoding strings like a middle schooler, and now I’m a 100x developer fearlessly leading the charge to usher in the AI era as I play the new No Man’s Sky update on my other PC and wait for whatever agent to finish crap.
trying to hacksmash Claude into outputting something it simply can't just produces endless mess. or getting into a fight pointing out issues with what it's doing and it just piles on extra layer upon layer of gunk. but meanwhile if you ask it to boilerplate an entire SaaS around the hard part, it's done in about 15 seconds.
of course this says nothing about the costs of long term maintainability, and I think everyone by now recognises what that's going to look like
I think there are phases in a project’s lifecycle where it’s more appropriate, at the very beginning and very late. I do not think junior developers should be using it, because it is much much harder to learn and it kills productivity having senior developers review 3000 lines of slop. Just stuff like that needs to be figured out.
So I’m not even in the “it’s useless” camp, but it’s frankly only situationally useful outside of new greenfield stuff. Maybe that is the problem?
And Ask DeepWiki is a great shortcut for finding the right context… Granted this is open source and DW is free.
Is it the specific nature of your work?
I think in retrospect it's going to look very silly.
In the past 6 months, I've gone from Copilot to Cursor to Conductor. It's really the shift to Conductor that convinced me that I crossed into a new reality of software work. It is now possible to code at a scale dramatically higher than before.
This has not yet translated into shipping at far higher magnitude. There are still big friction points and bottlenecks. Some will need to be resolved with technology, others will need organizational solutions.
But this is crystal clear to me: there is a clear path to companies getting software value to the end customer much more rapidly.
I would compare the ongoing revolution to the advent of the Web for software delivery. When features didn't have to be scheduled for release in physical shipments, it unlocked radically different approaches to product development, most clearly illustrated in The Agile Manifesto. You could also do real-time experiments to optimize product outcomes.
I'm not here to say that this is all going to be OK. It won't be for a lot of people. Some companies are going to make tremendous mistakes and generate tremendous waste. Many of the concerns around GenAI are deadly,serious.
But I also have zero doubt that the companies that most effectively embrace the new possibilities are going to run circles around their competition.
It's a weird feeling when people argue against me in this, because I've seen too much. It's like arguing with flat-earthers. I've never personally circumnavigated Antarctica, but me being wrong would invalidate so many facts my frame of reality depends on.
To me, the question isn't about the capabilities of the technology. It's whether we actually want the future it unlocks. That's the discussion I wish we were having. Even if it's hard for me to see what choice there is. Capitalism and geopolitical competition are incredible forces to reckon with, and AI is being driven hard by both.
No. BOON. A BOON to workplace productivity.
And then the writer doubles down on the error by proving it was not a typo, ending the sentence with "...was for several years a bust."