I think I have LLM burnout
alecscollon.com
alecscollon.com
I am happy about all the little side-projects, and ideas it help my realize, and I enjoy exploring this new world, but I've noticed LLMs feed my unhealthy "don't want to take a break and waste time being idle" mindset, and I need to correct it.
W.r.t. article's main complain - I think the similar thing happened due to factory manufacturing automation. What used to be a varied skillful craft in a shop became standing in a single place of an assembly line doing the exact same thing whole day. LLM took away the more creative and variable part of the work, and left the repetitive QA rubber-stamping. Probably some of the mitigations used back then could be rediscovered today.
I confess that the above variant on the quotation is how I originally read it. And that's just about how I feel now with trying to sort through vibe-coded slop projects that are put forth by (well-meaning, probably good intentioned, not evil) people who represent them as if they're the handcrafted result of one dedicated developer.
"I did a Chat output, please fix and review it " is the kind of thing that empowers the people who used to have a minimal productivity, and now lets them to wreck things on an industrial scale.
It's not. There is no one person that has universally good taste. Also, we're not in your head, no matter how much better of a coder or whatever. We're not in your head and it's all terribly painful to navigate.
AI is not a productivity multiplier. There are diminishing results.
The ones that notice the highest increases of productivity are usually the ones that were unproductive at best and dangerously incompetent at worst.
I don't. I'm not a gambling man, and I respect my fellow engineers to not inundate them with slop.
Sure it is. Just that some of the values being multiplied are negative.
Lots of companies (nearly all, I’d wager) of any size were leaving bare-minimum a 2x software development speed increase on the table before LLMs, having nothing whatsoever to do with how fast anyone was typing or thinking up code, and everything to do with how they organized and supported development work, and with your basic ordinary corporate dysfunction.
My company, I’d say it was more like 4x or 5x they could have achieved before LLMs, by fixing processes and reducing how often management steps on their own dicks.
All the people I’m seeing with crazy-high LLM productivity at my company? They’ve been given enormous autonomy to basically go do WTF ever they want, and people are jumping to get them anything they need (and most of what they’re doing is prototyping, for that matter). So right off the bat, if they’re competent, they should see a notable multiplier on productivity even if they weren’t using LLMs. Not that those aren’t helping, too, but if you don’t change processes they’re not all that effective, because the problem wasn’t speed of code-writing (and if you can change processes, you already could have sped up development a lot before LLMs…)
Spoken like someone who is not at an org/team that has undergone layoffs and reduced hiring in the last 3 years.
You might be in the minority there - especially when it comes to those who are facing burnout.
Cost of generation has been reduced, and is highly subsidized currently.
Cost of verification has effectively not changed. I’d say as a rule of thumb: verification is the tough part.
Our brains don’t fare well under constant review pressure. https://en.wikipedia.org/wiki/Ironies_of_Automation
Then why are so many others in the thread reporting being swamped with requests to review coworkers' slop? If it's genuinely "cognition" at trivial cost, surely this review would be completely unnecessary?
This isn't the problem; the problem is that people incorrectly believe it to be true.
I think this is in part because I am one of the software engineers that always liked building products more than writing complex software. So, I am driven by the feeling of creating something. And I want to get the feature perfect and complete. But getting from 95%->100% done can take a long time with UI work for me.
So I work much longer hours now, unfortunately.
Main blocker is I am using apps like Conductor and have lots of plates spinning at once. But that's on, me and I should try and start completing the last part myself.
But it's probably a common feeling. I wonder if we'll see an increased number of people burn out in the serious, medical sense.
I'm getting so many requests to review LLM-generated documents - planning docs, docs intended for end-users, project docs, business plan docs. A team member sent me a zip file with about 30 LLM generated documents in it the other day and asked if I could review them right away. And a lot of it was just repetition and/or stuff that was just out of left field, made-up, hallucinated stuff. They're able to generate this stuff way faster than we can review. It used to be that it would take a significant part of a day for a project manager to come up with a planning doc - now they can generate one in a few minutes and send it out for review. It's just really tiring.
> do not hallucinate
They do, just less. To the degree of being usable, as long as there are guardrails and they're used responsibly. For example, if there's code being output, there should be type checking and compilation, as well as code tests that prove that it works or that it doesn't - seeing how abysmal code coverage is in most of the projects I've seem, for whatever reason people thought that they didn't really need it much. They were wrong.
This also implies you need SOTA models on max reasoning.
> make things up
Same as above. Ideally you'd give them some way to verify their claims, like web search or browsing and referencing docs, Jira tickets etc., basically improve the signal to noise ratio.
> contradict themselves
They do so way less than before, as long as the above is true.
> can review their own output into perfection
They are pretty good at reviewing things, especially if you make them do adversarial review! It will never be perfect, but can be close in quality to human output (e.g. the code they produce, when used properly and with intent, is better than the code I've seen many developers write and ship before LLMs were a thing).
This also more or less scales with how much compute you give them - three parallel review agents will turn one output artifact into something good with higher confidence than two, and definitely better than with no review. There's a cost vs quality balance and it seems that all those xhigh and max reasoning modes are still geared way too much towards cost, instead of quality. So you have to make up for that shortcoming yourself.
> regardless of task, goal or context????
Garbage in, garbage out. I won't be an asshole and say that you're holding it wrong, nor will I say that anyone should listen to the claims marketing AI (absolutely delusional takes, meant to attract investors), but we're slowly getting to a better position in regards to LLMs, year by year.
It's just a shame that the peak of inflated expectations hit while the technology still hasn't fully plateaued and reached whatever its ceiling is.
I probably also shouldn't ignore the fact that some people will not care about any of it and send AI generated slop verbatim and to an outside observer there's no way to easily tell apart the difference between the two, unless you make a technical report contain exact references to where the data is sourced from, for example (and then either verify the references yourself, or make another agent do it).
I think we will very soon move to a prove to me you've read it protocol and/or introduce speed bumps to slow things down.
Why not just review a single document quickly, find an error which invalidates the document, and send it back saying "Policy paper 1 mentions X as being on the business plan for Y, it's not on the plan, please can you fix."
Unless you can write a good-sounding reason why it's on them to review a LLM output before sending it to you, they will outsources this reviewing to you, and it's a lot of reviewing.
If it's genuinely hard to find that single bug .. perhaps the document has reached the quality required for corporate communication?
In my experience, it does take a lot time and effort to find contradictions between 10 documents. Even with good documentation, it's hard to build a mental map for that amount of information.
That ratio has changed, and verification is the hard part.
Verification is the point of all markets (and a decent part of human civ as well).
And review isn’t cost less - https://en.wikipedia.org/wiki/Ironies_of_Automation
Hopefully you mean Rate of verification/Rate of generation.
Verification/generation
He who brings the slop cannon shall be drowned by slop rain.
If this approach gets widely adopted, then I think you should probably start building an ARK.
Make it big enough to hold two of every animal species. /s
Getting an LLM to vomit out a bunch of documents and sending them straight to another colleague is absolutely unacceptable behaviour.
Which is going to win?
We also need to be motivated to stay in our jobs.
Most developers like their projects and value their work. But the chances are that it's for nothing.
Many developers know they work on bad products (gambling industry, military, surveillance, whatever) and so it's here that they focus on their technologies, tools and frameworks rather than the work they produce.
"Agentic engineering" for example.
Id be curious to see what and how Googlers are doing with their 20% time.
If you're an employee in that situation, push back if you can. If you can't, put your resume on the street. (But that may not work, these days. If it doesn't, all I can say is ride it out as best you can, and try to maintain both your job and your sanity. How? I don't know.)
That's the other nightmare of AI slop. So easy to generate endless content. Who will review?
Just today the boss request I review slides for a presentation. But it's all AI slop, generated from querying tickets and docs and who knows what. It's mostly sort of correct but also plenty misleading and incorrect. So now I have to fact check all this slop which will take hours (even with my AI assistance) and rewrite most of it.
If AI didn't exist, he would've had to do the research to generate the content and it would be 99% correct and I could just give a few notes of feedback in 5 minutes. But with the asymmetric AI workload, he can generate it in 5 minutes and I get to spend 3 hours correcting.
Maybe, depending on the boss. Some would have spent five minutes describing what they wanted, and someone else would have spent three hours creating the deck.
Problem with AI is that generation is so many orders of magnitude faster than reviewing so it's basically infinite monkeys on infinite typewriters.
You just discovered the unlock to massive AI-driven productivity increases: outsource the hard stuff to others, or just don't do it at all. Keep the easy tasks that generate a big volume of output for yourself.
Or use LLM's to generate 12+ pages of detailed reviews of those documents and return to sender.
this is just spam, people sending unsolcited data at you and expecting you to swoollow and process that data.
its just rude and unreasonable, not to mention an unconscious (hopefully) act of sabotage.
I see a different type of pressure: I'm at a company that still is requiring everyone use LLMs with token leaderboards, time-spent measurements, and impacts to performance reviews, and all that. So I find myself having to carve out some percent of my time to stop doing productive work, and "go do AI to show token use." So my workload hasn't changed (or it's gone up), but I have N% less time to work on it because I have to spend time appeasing the AI gods...
Peer Gynt Suite's "In the Hall of the Mountain King" made a prominent appearance, but so did Aqua's "Barbie Girl"
Probably like eating. Having to eat less to loose weight isn't great. Eating anything you want without worries is great. Having to eat more than you want to gain weight, not great.
Just be careful about any legal implication of doing side-projects during work with work-resources.
Same, but I really have to fight the urge to just add fun new features to things I work on any time inspiration strikes. I am an appalling 'feature factory' if I don't actively keep myself in check. The cost of just building everything is so low, but the value of those things is also incredibly low, so I'm often just bloating what I build.
There's been a lot of articles and posts about the increasing importance of 'taste' in software built with AI, and I'm finding I know need to look for strategies to find some.
I had to think of the factory scenes in Charlie Chaplin's Modern Times. The author's feeling is basically the main idea of the sketches, i.e. humans having to follow the pace of the machines instead of the other way around.
Reverse centaurs are nothing new. Ask any worker movement from the last centuries.
“I wanted a machine to do the dishes so I could concentrate on my creative work, and all I got was a machine to do my work so I’m left to wash the dishes.”
Embrace it if you’re like me and feel uncomfortable having an idle mind, embrace it! You’ll get more done and being 120% go go go is impossible over the long run so eventually your body will just say I need a break then once you recover full steam head again on the treadmill
Yes, this is to me the primary driver of the extreme AI burnout. In ~30 years in Silicon Valley and many, many startups, the pressure has never been as intense.
Before AI I'd mostly work on one thing at a time (at least within a given hour) and in the evening I wouldn't start a new 6 hour task because it's too long, so tomorrow is another day.
Now, that 6 hour task is more like 30 minutes, so there is intense pressure to just knock it off tonight. And then the next one. And one more. And while the bot is thinking, to have 4 other work streams in parallel so there is never, ever, a break in the day. The human mind is not built for 100% utilization 15 hours a day.
I find LLMs to help me manage the unrealistic workload I have, because at least now it's feasible instead of just getting more work piled on top of me with a never ending backlog (that people actually expect me to thin, not let grow). Add on top of that colleagues that would have death by commitee'd many ideas and now just have to argue against actual MVPs that work instead of ideas (or can be proven to not work and discarded without wasting time on them in some cases), and I don't even hate my job as much!
It's just that to ensure that the technology is not a net negative, I need millions upon millions of tokens every single day (tool runs, adversarial reviews, testing), but once you get that inflection point, alongside needing a good enough model, the floor for which currently I'd say GLM 5.2 on Max reasoning reaches, or use something like SOTA Anthropic/OpenAI models, it becomes a pretty good way of working. That said if you have missing pieces there (e.g. using cheap models that aren't very good), the curve of getting stuff done can go downwards and you'll just end up with a lot of slop - useless docs, bad code and an ever increasing amount of technical debt.
On average, each task that I do, needs about 15 minutes to 2 hours of planning and making the agents explore the codebase and refine the plans first.
Curiously, in my case this leads to less burnout cause I can actually pause and grab a drink, meal or go for a walk, while parallel agents do the work, once I've planned things well enough and have dispatched something that will work for 1-4 hours. I don't have to review their output immediately once they finish but can just batch things.
AI coding is addictive. Engineers are paying the price https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...
It's impossible to undo some of these linguistic wobbles. Even if you could filter out 100% of LLM input, the humans themselves are learning to say "land" at a higher frequency now.
Its like when someone points something out a in picture you never saw and now you cannot "unsee" it ever again.
I've often had to paste its output back in to ask it what it actually means. Weird.
I think the main thing is just fatigue. There's so little variety. Each model has its preferred idiolect which everyone becomes tired of due to ubiquity. That's the worst part. It's like always eating fast food.
I used to have a lot of fatigue due to it until I stopped caring.
*onanizing…
The bots (all of them) seem to show patterns of overuse of specific phrases, words, and punctuation.
Some of those are the ones you mentioned. Another that I've been seeing lately is overuse of the term "gate", wherein: As a human, I know what a gate is. A gate is a thing that can be open, or that can be closed. It might be locked or unlocked. The path beyond the gate may be passable or impassable or nonexistent. The gate is just a gate, and the presence of the gate doesn't imply whether it is open or closed.
But in bot-speak, a gate only refers to a hard block -- an impassable construct. Like a fence or a wall, or even a lava-filled moat.
But while a lava-filled moat is intended to be impassable, the bot uses "gate" -- a thing that is designed to be passed -- to describe that same kind of obstacle.
That's misuse of the term, I think, based on decades of dealing with gates in reality: Usually when I encounter a gate that is closed, I just open it and walk through.
I do have instructions that tell the bot to avoid that usage of the word and it ignores them sometimes anyway.
But "gate" is just today's problem-word that comes to mind as I write this. Yesterday, it was something different. Tomorrow, it will be something else entirely.
The overall pattern here is that of gratingly-repetitive bullshit-grade jargon that doesn't fit to begin with.
"And that's the real, no-nonsense truth!"
Another example of typical botspeak is "smoke test". Why not just say "test"? It feels like a way of downplaying the ability to detect problems.
(The project works well and I consider it to be Good Enough; I might go back and polish it more later. There was no smoke, but there could have been.)
I've been on the Internet for ~35 years. What did I miss?
But where we do have them: At a given time, the gate might be open or closed; passable, or impassable. The presence of a gate is implicit, but the status of that gate is not known without advance knowledge or direct observation. And even when it is closed (even if it defaults to always being closed), there's generally a cromulent way for a person to get that gate to open and then move beyond it. It is designed to be opened and closed.
Gatekeeping: Sure. I've run across a ton of artificial gatekeepers online in my time. I've bypassed countless scores of them. Those are easy: Just ignore them and keep moving.
These aren't examples of the hard-blocking, impassable lava moats that the bot is fond of using "gate" to describe.
Perhaps we're holding it differently.
I understand in the past HN accounts have been afraid to post code because it would reveal that they have no idea what they are doing, but now that we have the plausible deniability of "look at this stupid thing an LLM wrote!" it is curious the mindset persists. Old habits die hard, I suppose.
In reality, there's no reason for me to do so, because I'm not on trial here at all. Please keep your accusations to yourself, comrade.
My willingness to prove to you that a bot uses a word in a strange way is approximately zero.
Have a great day!
Anything written for humans should be written by humans.
Whenever I give feedback on something, the answer is just “let me tell Claude”. The person has no understanding of how everything works, and most of the code reflects that.
The other day he hardcoded in a demo mode, simply because he didn’t even know how to set up a local environment and set environment variables. I’m confused as to why Claude didn’t even knew this, but it might just be the prompting.
I limit LLM usage myself, and if I do use it, I try to use it on extremely specific tasks. It’s the only way it works for me.
I honestly don’t understand how all these companies are getting away with generating AI code. Even in a small project I quickly fall behind on my understanding of the project.
Now these people can thrive because LLM coding encourages the incurious and punishes the deep thinker.
Now your feedback is just another prompt for them, the code might be slightly better, but the person learned nothing from it.
This is an industry where American developers have successfully competed on quality with the rest of the world for years. We never were very cheap but we always were the best and worth our premium. Now that's being destroyed by a short-sighted industry.
I'd like to just do something else and work on open source. Except I know if I contribute to open source my work will just be stolen by the plagiarism bot.
It’s possible that the experienced developers you know just aren’t that good at adapting to major changes like this.
"Stroustrup points out that some senior developers are starting to retire because they're tired of constantly validating unpredictable AI code. "
https://www.newsbytesapp.com/news/science/bjarne-stroustrup-...
On the other hand my buddy is spending $10k in tokens a day on agents to build something. He's a very smart guy, former developer so it's not just AI psychosis talking.
Still trying to figure it out. Not that I have $10k to spend.
What on earth is he building?
This is what you want. You want comprehensive tests at every level, far more than is reasonable for a human to build or maintain, from unit, functional, to full end to end and beyond. Adversarial testing (both TDD-style "write tests to demonstrate this bug", and posthoc "prove this patch wrong with a new test") is the best way to keep AI on track and make those diffs you have to read clean and easy.
An even better way is to use a more strongly typed language and really lock it down, but you can use testing in any language. I feel like my background in TDD and "TATFT" has been secret sauce when working with AI
https://github.com/dprkh/eventfs
It has good test coverage, mostly unit tests but also a number of end-to-end tests. I also made the LLM build a benchmark, which you can find at the bottom of the readme. It is obviously slow, but I thought that it is good enough to work. When I tried to write a 1 GiB file, I found that it broke down, and after writing half the file, the speed went to under one megabyte per second. Implementation is 10k+ LoC, and I have no idea what is going on there.
At least with agent-run tests I care about loop speed a lot, but I care about complete coverage more, so having the odd heavy weight full stack integration test is fine, I think.
Yes tests are conceptually isolated and that helps, but I've personally seen unit tests get generated that are semantically incorrect - that is, they test the structure of the code (e.g. they can check function output types and values), but they can't know _why_ the unit tests need to be there, so the really really helpful tests never get generated. Not to mention the obvious issues with generated tests only testing is x = x, or needless redundant tests for the same thing, or them essentially testing basic features of the language.
I actually have a public (AGPL) example here: https://github.com/pgdogdev/pgdog/tree/main/integration/sql - pgdog is particularly testable since it is trying for complete transparency, so you have a perfect oracle in hand via base postgresql, but it demonstrates the concept at least.
This is also part of why I like end to end tests that use actual UI flow, so I can watch it go by in slow mode before letting it loose fully automated.
What do you mean by "be epistemically sound enough"?
You are using it as if to say "if your code is grounded in sound abstractions, you'll be fine and tests will therefore generate successfully" but preface that claim with "the code provides a baseline truth for the tests". The latter does not follow from the former, and it also does not lift the burden of responsibility away from the programmer - which is where my doubts on test generation stem from in the first place.
Additionally, what is "completely solipsistic value generation"?
You reference it like a perk in a skill tree, but to my ears "generating completely solipsistic values" seems like a way of describing AGI with a philosophical wording instead of just saying AGI.
Also you: > You have to iterate on the tests, review and validate them
Yes, "maintain" is not quite the same as "review", but the line is veeery fine. I find it really tiring to review masses of tests that an agent spews out.
Especially because I know what it has a tendency to write irrelevant/vacuous/useless tests. It's insane the amount of times I have told Codex to "write a test that reproduces the reported bug, SEE THE TEST FAIL, then implement a fix", only for it to guess an irrelevant test, not run it to see it fail, and implement a code change that has nothing to do with either the test or the actual bug.
I've been burned by this in my honeymoon period with unit testing (pretty much the reason it ended). These days, I prefer broader scope of testing, especially user-facing part. The users may be other developers or end users. I only do unit testing for tricky algorithms or math formulae.
They’re mostly a reflection of the current requirement of the project.
It is different though. Basically a lot of what I do has changed over the last 2 years. I totally get that a lot of people won't want to adapt though.
Or people don't want to be reverse centaur keeping the clankers happily running. Instead of helping to solve users/consumers problem.
It's been an experiment to see how much more performance I can squeeze from a Rust version (spoiler: it's a lot), how well the agents code in Rust (pretty great and seems idiomatic AFAICT), and if this is a good way to learn a new language (I'm learning, but the verdict on how efficient is still out).
I might be self deluding, but I do think it's been productive, even though I'm intentionally moving slow with small TDD vibe spikes followed by completely reading over everything, adding more guard rails if necessary, refining requirements and tests, sometimes ripping it out then and have the agent rewrite it more iteratively with meticulous reviews, etc. Honestly, I have the time to do this right, so I've been focused on correctness and making it enjoyable to avoid burn out... but what I find enjoyable, won't be the same thing others find enjoyable. I also have the autonomy and financial security to adopt entirely new workflows and do rewrites of my own products, which not everyone has. I would absolutely hate being forced to token max or w/e that insane BS is all about.
I save myself by skimming things like tests, templates, some UI. Anything cosmetic. But I have to read the majority of code that ends up on my back end systems.
In my personal experience, the ones most enthusiastic about LLM magic are those that can't code, but can now walk away with something functional if not quite the best code. Now that they can produce workable code, it will make everyone better. Yet, they have no idea how maintainable the slop is or if it's slop at all.
When you see a perfectly clear function or object that just isn't your style, you have to accept it and move on. Where there are concrete concerns, or it's unreadable, demand excellence, but treat it like a coworker, not an IDE.
The only time I look at code is when something isn’t right and I ask for a root cause analysis. The LLM will show me some offending code or code for reference or evidence and then I quite often say “well that’s dumb you should do it like this instead” but I never need to actually go into the files. I do sometimes look at a git status or git diff.
is the critical caveat to “that’s not how I would have done it”. Basically, choose your battles because we all have limited bandwidth. So, it’s not really a perfect binary, but a taste that you personally develop.
I’m building personal projects at a prodigious pace. In a role reversal I treat the agents like I’m one of my clients (albeit a more technical one who gives them architectural direction) and they are me. I’m using the apps and tooling they make every day. I’ve cancelled SaaS subs for tools I’ve built myself.
I watch the tool calls and realize I should be better at core command line tools so I have a study plan to catch up (just a little bit a day). I’m revisiting long standing config that I dropped in to vim and tmux way back when I started and didn’t know anything.
I guess in theory I could hold my productivity to previous levels and read more. But it doesn’t feel like that’s possible. It feels like we are in one of those sea changes where the promise is less work, but the reality is increased productivity and expectations (the Industrial Revolution feels like the right parallel to reach for). Increased expectations happen in small ways and large. The agents are so good at polishing data presentations that I always send cleaned up visually impactful reports that would have taken significant time in the past just as a matter of course now.
But, I’m tired. I’ve spent the Fable on subscription window sprinting through as much work as I can before it goes API only. (As an aside, I don’t understand how everyone is using so many tokens. I’m sleeping very little and running as much code as I can through fable and I can barely touch a 20x max plan limit.) I keep telling myself I will slow down when it comes off, now it’s extended to the 12th and my window just reset, a few more days to keep knocking out backlog items. I feel like I have to keep the robots busy overnight so when I wake up I can immediately sit down to review. I give directions to agents on my phone which feels wild to me.
OR
just keep coding by hand, thinking things in front of a whiteboard when it gets complicated, burning myself out slowly at a more humane pace as we have done for 50+ years of software engineering.
Why is this even a choice? I mean, serious question, do you people have a little bit of self-respect? I am expecting the excuse of “but my boss expects me to use AI”. It is clear most of you have not experienced true burnout, because it’s going to be a world of pain when it hits. Stop turning yourself into machines. You are not.
I’m not sure where you work, but I have to compete for every contract (and then to keep them). Hand coding everything isn’t an option anymore or won’t be for much longer. And I don’t even think it’s just freelancers who will be affected. As a solo dev I feel like I’ve been given super powers, but if I worked at a hundred head software house, I’d be looking left and right to see which 30 of us were going to be left when the dust settles.
Not to trivialize your point. I agree we aren’t machines and we should reject being thought of or treated as one. But the reality of software dev IMO has already moved. All reactions are likely over reactions, but I don’t think the swing back is going to land anywhere near where it was.
RE personal projects specifically: addiction is an interesting angle. It feels more like anxiety though, that this thing is going to be taken away and this will be a brief and glorious window when I had the means and opportunity to make all the things I wanted, but never had time for. It’s also exhausting to have a bunch of stuff that you want to do piling up for years.
Could just be a personal thing though. I had kids shortly after I started freelancing. I’m the primary caregiver for them and my partner has always worked insanity hours so just keeping our lives together is a full time job. Every billable hour had to be squeezed from a stone. Maybe I’m trying to make up for lost time.
When the rug gets pulled, what are you going to do with this giant pile of unfinished and unmaintained projects?
I would guess that anxiety and withdrawals are very similar
From the outside, this all sounds like my father who worked as a mechanic at a flour mill.
His job was mostly watching the machine do the work and then fixing the machine when it broke down. 90% of the time it ran fine.
Going from working by hand to watching the machine would seem really boring and I guess I can understand the AI hate on here from that perspective.
(which why on earth would be applicable? yet this argument is thrown around begging the question [meaning assuming its own validity, begging to be questioned, not suggesting further downstream investigation])
Namely: there was a massive shift from household manufacture in the Indian subcontinent - then home to an estimated quarter of economic activity on earth - to Britain.
Industrialization was about out competing a geopolitical rival of the UK: Mughal India. It succeeded because the Crown adopted policy sabotaging Indian production, not because the capital intensiveness was inherently better. It was merely necessary for England to even aspire to outproduce a country with 20x its own population.
Industrialization was never about the sheer efficiency of the new industrial productive system. That wouldn't have tipped the scales in a matter of one or two decades. In fact, quite the opposite: the incredibly short timescale in this massive geographic shift in economic output necessitated an approach which was costlier, and socially and environmentally damaging.
The machine won out because it was the only way to get the job done. It was a way worse experience for everyone involved. Just look at the British elite's tenure: its 100 years of zenith pale in comparison to the Mughal third of a millennium ride. And now the center of gravity for global industry is shifting back to Asia despite the extremely heavy price paid by Atlantic society and the global environment for the anomaly state which is now waning.
For client work, a similar split. Simple UI fixes, I give it a scan. A feature that will have to be supported indefinitely, I’ll spend a few days on it making sure I really understand it. It’s not really a fair comparison though because by the time I get to a serious review I’ve already done multiple iterative sessions planning and architecting with the LLM then approved each commit as it has gone in.
I’m conflicted though. For sure I don’t have as good a grasp of my own codebase as I did when I wrote every line. I think I’m ok with that though.
Prior to the last 12mos AI companies were hell bent on squeezing out the best results from mediocre models.
But... now that the top models have progressed, those same AI companies have switched their efforts into reducing the computation (cost of a producing a result) as much as possible without being too obvious.
What was an exponential slope in the quality of results over the last 36 months has now nearly flat lined.
Addendum: IMHO results have 'flat lined' not because the models aren't much more capable than a year ago, but because conserving the enormous processing cost (of an over subscribed user base) supersedes the goal of following the user's explicit instructions (e.g. especially if that means more processing cost) to generate the best results.
Before, they could stay in thinking mode for more than 7 minutes. For example, "find a source for this claim" would search, analyze, and self-adjust the query. Nowadays, even if I push for it, I cannot make these tools work for more than 30 seconds before they give generic answers, even in "Pro" mode.
How empirical are your comparisons of new and old outputs?
Hell, the Opus 4.5 moment was only last November, and that was when agentic coding and most coding CLI tools became truly first class options. That's a wild paradigm shift. Hell, GPT-5 wasn't even out (that's August of last year). Most people were using 4o. Their current offerings are wildly better for coding than 4o was.
But that's also let me use "agent" stuff longer, I guess? The better you were at knowing what you wanted and how to ask for it, the less of an inflection point that you got from Opus 4.5 or GPT 5.
Some of the highest-time-saved-for-max-ROI agentic problems I've solved to date were in September and October of last year with Claude or Cursor.
Just because we work with computers doesn't mean we don't take, er, social-damage. Or perhaps parasocial damage, in this case.
You'll frustrate yourself by not using this tool without above. You'll definintely frustrate yourself expecting the tool to be "genuinely sorry"!
I got into programming because the problems of programming were interesting to me. But if the problems go from "figure out why this calculator is off by one in France" to "Get this LLM to stop spamming cutsey emojis", then maybe it's time for a career change.
My latest is, I'm really into fizzy/soda water and wanted my own continuous carbonator. My entire build from water source to tap with an ESP32 controlled pump, pressure, water level, cooling fans.
There were so many areas I made mistakes in my shopping cart and it found it - like Home Brewer likes 8mm lines but water filter systems like 9.5mm. Really optimized the versions from a simple on/off pump w/ float switch to effectively a full on PLC system. So many iterations gained by chatting with "someone more experienced". Once I get the parts I can build and have the software side running in less than an hour.
It doesn't make money, but man I really enjoy it.
Fun fact, if you see in the course catalog that’s listed as Remote/Online, check the scheduled times; if there are some, they’re either mandatory meeting times or in-person testing days; if there’s none, then it’s a fully asynchronous class you can THEORETICALLY complete in parallel to your job, whenever you like.
You could dip your toe in the water very slowly the first term and set calendar reminders for the drop & withdraw deadlines. At worst you don’t like it, at middling you realize you can’t multitask school and work, at best you pass the course. One step closer: wax on, wax off.
We’ve had 3 production incidents this week that slipped past CI because there’s a whole team that is just shoving out PRs without understanding what’s going out.
It's not surprising that if you have a hundred separate, isolated contexts working on the same business, that don't cross-talk and have no ability to subconsciously receive and collate, prioritize the thousands of signals we get from our work environment, that you end up shipping lots of incomplete or incompatible work.
The experience is much closer to working with an external API that you don't have control over and which simply doesn't do what the documentation says. Those have always been the most frustrating parts of programming, but at least previously you could reverse engineer the actual implementation to work around bugs. You can't even do that now because the "boundary" randomly change every day.
Sorry, some of us have a joy for programming where the how is just as important, if not more so, than the what and the why. No matter how much people proclaim that the how doesn't matter to them, it isn't going to suddenly make it true for others.
Why should I regret that? Why should I care about your purity tests?
This is dishonest. Your quip was to imply that I'm reducing what is otherwise a fun activity to an automation, on the basis of a purity test.
> Some people merely enjoy programming more than engineering.
And I haven't said otherwise, so I'm not sure what point youre trying to make. My initial goal was to provide my own viewpoint on how to enjoy the process while taking advantage of modern tools, not to tell people they shouldn't enjoy programming more than engineering.
This isn't a response I expect from people who are here for a productive discussion. I'm sorry that you are sick of hearing this, but I'm not responsible for making sure you only read what you find worthy of your own personal brand of respect. Instead of attacking someone for simply offering their point of view, in what appears to be a quasi-gatekeeping effort, maybe you should look inward and discover what is making you this upset toward a complete stranger.
__I cannot take away the joy you have for programming simply by stating what drives me.__
Look, I don't have a problem with your personal motivation. I just hate seeing it suggested that people should abandon their passion because someone else doesn't share it. There's absolutely nothing wrong with "I enjoy making something useful" just as there's nothing wrong with "I enjoy making something with my hands or figuring out how to make it". My problem is with "your enjoyment of that is invalid because I don't enjoy that, so learn to enjoy this".
0: Not really a statement of your drive, is it? More of a directive or suggestion.
then so is your opinion on the matter. this goes both ways. I can just as easily dismiss your commentary as "drive-by".
> Not really a statement of your drive, is it? More of a directive or suggestion.
Yeah, so was my initial comment. Merely a suggestion concerning view points, and how a shift in that can bring back some amount of joy.
You don't have the moral high ground here.
As a programmer, you also had to work at that higher abstraction level anyway.
It's a myth that you are "moving" up a level; you were always at that level, just not exclusively.
do you mean my enjoyment from building things? I'm genuinely confused by this response.
Those who like having a finished thing. Product people. These people love LLMs.
Those who love the process of building a thing, working through a problem, learning something new. Finishing a project is generally not required. For them LLMs are soul sucking hell.
I'm not surprised - your GGP comment indicated that you are more interested in the destination than the journey (you enjoy the output more than the process of crafting that output).
Nothing wrong with that - lots of programmers are interested in the final deliverable and don't really care about how the sausage is made, but you're reading a comment from someone who makes the sausage.
That's tantamount to "I don't care about the coding itself, or the underlying quality of my systems".
Your comment is incredibly aggressive, hostile, and rude, of out nowhere. You should review the guidelines [1]. It's deeply disingenuous to either communicate really poorly or to backtrack from your position unannounced and then get aggressive about it when people engage with you.
This is "incredibly aggressive, hostile, and rude, of out nowhere". Please leave me alone, I have nothing else to say to people like you.
For anyone else, imagine a winemaker that cares about their end product more than they care about growing grapes.. does that mean this winemaker doesn't care about the quality of their grapes? Probably not, but apparently to this individual the english language requires that we believe the winemaker to be disingenuous if they clarify their position on grape quality.
this AI bubble will pop. when it does you'll be hot stuff all over again.
Of course if you're supposed to achieve so much output that it's not possible to do anything but vibe it, fair enough.
Like when I'm trying to get it to create an image, and the first pass is beautiful, but ten different request to modify it, with different phrasing and even example images, produce the same image ten times. Or when you tell it not to use a cheap hack in AGENTS.md about six different ways and in your prompt, and it still does it again, and again.
It's like arguing with an idiot. And THAT gives me burnout.
Also: I've never once seen an emoji in LLM output. What are people talking about?
It seems to be heavily dependent on the task.
It takes like 5 seconds to add a quick CLAUDE.md / AGENTS.md with a quick style guide, or even just “EMOJI ARE FORBIDDEN” and I find it makes llm output significantly more tolerable. That and a quick style guide with some banned words and phrases.
That won’t help with the false assumptions though, gotta use old fashioned careful reading and critical thinking the catch those.
</irony>
Me: "The gap, stated plainly:" stop using the type of language.
Claude: "You're right. That's one of the constructions your preferences told me to drop, and I used it anyway."
But when I ask it to do data analysis or modeling, the emoji are all over the place, yes.
(And judging by what I've seen on GitHub over the last year or so, I would never in a million years consider asking an LLM to write a project README or documentation unsupervised.)
The productivity drive and the sheer feature set you can generate in record time makes it easy to forget proper sdlc hygiene.
… but I am almost certain I’d never have developed those in the first place if I hadn’t spent 25ish years programming on a bunch of different platforms and setting up servers and networks and all that, without LLMs.
I dunno how you make another “me”, now, while before lots and lots of programmers naturally ended up as someone with skills and knowledge like mine, and those skills seem super useful when writing code with LLMs.
What helped was a sleep and work system, oriented around being offline that was inspired by nature and from my earlier days in working in tech while car camping across the national parks.
Basically: the sun wins in terms of how all energy on the earth is structured, and expressed. All manners of cycles of organisms and living systems are in relation to its rise and fall, and even its particular color spectrum phases (whether thats night oriented or day). I call this our real circadian rhythm; it's used to being signaled by the light of the sun and maybe fire for millions of years and it isn't until recent centuries when we started tricking our biology with LEDs and lights. So the solution is simple. Orient yourself around the light of the sun and make sure it's the first and last major light source you see; blue limiting is the most important part BEFORE sunrise and AFTER CUT OFF ALL BLUE LIGHT. On my Mac I use a red light filter (using it now, it's 11:07pm ET and the sun went down about 2.5 hours ago). It's really hard to stay alert and chatting with an LLM when the only light sources are red and you keep them dim at that. Our ancestors would rest when the sun's at its peak (~1:05 pm today) and that's a good time to divide my own day productively as well. With intentional breaks diving the middle of the day with sunlight anchoring it, my nervous system is more relaxed, and by the evening time, it's also ready to transition out of anything blue-light assisted and most intellectual work and problem solving falls into this bucket. It's really hard to explain but it really works so simply. To enjoy the process a little more I made this fun sun clock, check it out at https://sunsignal.app
Cured my lifelong “night owl” “trait” in a couple days. Shockingly effective.
Turned out to be hard to keep up and still, like, exist with other people, and you’d probably need to relax it a little in Winter unless your job lets you work reduced hours to kinda “hibernate” (otherwise when would you do anything that’s not work but requires light or electronics?) but it sure worked.
But burning non-petroleum candles a couple hours each evening (and I found it took two tapers to have enough to comfortably read by) is pretty expensive. Plus burning anything isn't great for air quality, even if it smells OK. One person doing it in a full-sized house, alight, maybe. Three or more people in a house carrying around pairs of candles every night? You might start to notice the air quality and your shelves or walls getting a bit sooty.
If I wanted to do it longer term I'd probably build a few shaded (it actually hurts to look directly at even something as dim as a single candle flame, let alone a slightly-brighter bulb, when your eyes are adjusted to low light; I didn't come up with a solution for this with my candles, in the ~3 weeks I was using them) nightlight-level-output battery powered lamps. Most of the lamps or lanterns you can buy have built-in LEDs these days (so, limited ability to modify them) and their dimmest setting is way brighter than two actual candles, plus you'd want a candle-warmth color temp and lots of them are cool sunlight-temp. So I think building a few would be the best option.
Standard modern nighttime lighting started seeming insanely bright after a few days of doing the candle thing. It really takes very little light to navigate your house and do many tasks safely and comfortably, if you let your eyes adjust and don't have some kind of serious vision problems, and if you've got a few people awake in your house doing things after dark it's easy to have hundreds of times as much illumination active as anyone really needs. I also kinda liked carrying my light with me, rather than flipping switches from room to room.
>Standard modern nighttime lighting started seeming insanely bright after a few days of doing the candle thing. It really takes very little light to navigate your house and do many tasks safely and comfortably
Yes agreed. It's so nice once you get used to this, and I think this is how our ancestors lived for millions of years followed by candles, incandescent and then horrifyingly today's blue lights.
Did you notice blue in specific feel incredibly unnerving and painful to look at at night even in low light output after adjusting to night time settings?
On the “pre” side, the specification of the problem becomes much more important. On the “post” side: QA and verification that the change has its desired effect, and no ill effects, also becomes much more important.
Sure, these are the next things to be automated, and people will try, but it’s easier on the backend (testing/verification can be automated) than the front end (the spec will be human-written as long as someone cares = forever for brands that matter), there will always be a need for humans on the specification side.
Burning out to grind out tons of code - for which you get paid no extra above your salary - is not a net win for coders. It's only a net win for the employers and it's turning coders into serfs. People are going to get wise and realize the electricians have a way better deal right now. This is no longer a nice way to earn money for a 20 to 40 year career.
My mind still can't function well without having knowledge about everything.
That said, your reaction is totally human. I personally get sick of how the LLM writes prose with always the same tricks and formulas (even if you prompt it not to). Humans need variants and novelty, that's why fashion exists. We get fed up with repetition and after seeing too much green shoes, seeing a red one is so relieving :) (quick note: I don't like fashion - I'd advocate diversity and personal styles, not fashion)
But that's also the way you work with AI that might be part of the problem. Personally, I don't review all the code the AI generates. I look at it, and I review only the code that matters. And with time on a given project I review less and less because I trust more the architecture and ability of the AI to follow it. In my settings, the AI gets confined to the existing architecture (that we define together at the beginning of the project), and has to ask for authorization to create new things (that's when I review the more). Hoping this could help to avoid burnout myself...
https://en.wikipedia.org/wiki/Ironies_of_Automation
https://ckrybus.com/static/papers/Bainbridge_1983_Automatica...
“Look at the first letter of the prompt, and use the style of the matching literary stylist enumerated here:
For English:
A — W. H. Auden
B — Bill Bryson
C — Italo Calvino
D — Joan Didion
E — T. S. Eliot
F — William Faulkner
G — Gabriel García Márquez
…”
I certainly don’t see emojis any more.
There's some YOLO approach to it, but now Codex has self-approving as well as Claude Code (auto mode). I implemented the same feature by my own on Pi with models through OpenRouter and found results very stable thus I have (as always) limited confidence it can fly.
So (disclaimer: I'm Jujutsu advocate :)) I do "jj new", tell it what to do and then let it run, and check in back later.
If there are things I'm not comfortable (like creating PRs or pushing to repo) I ask it to create Ruby scripts instead named like "__pr.rb" (double underscore files are in my global gitignore). So I can leave it working and then inspect back and edit manually before I run "ruby __pr.rb".
The only thing that's not yet there is tying multiple tmux Claude/Codex session together, but I'm thinking about creating a small Rust app that communicates with Tmux for a preview (or a Ruby script that communicates with my LogSeq directly and manages nodes there :))
As a long-time engineering manager, PM and, eventually, product owner my response is, "Congrats! You've just been promoted to management." :-)
As a new manager, your first challenge will be successfully delivering commercial results using only a team of 'differently abled' new grad interns. Don't complain, new managers don't get to pick their first team! To be honest, these guys are more like alien brains raised in a vat with no direct senses. They've only ever experienced a data feed of the internet and, oh yeah, they get near-total amnesia a few times a day (but maybe you can teach them to write notes for themselves). They also have ADHD and are somewhere on the spectrum. But don't worry because what they lack in common sense, experience and intuition is offset by having a sort-of photographic memory and a willingness to grind on a problem 24/7. You should be fine. Good luck, we're all counting you...
Including myself as well. It's how you grow into the role.
At this point you should stop risking to burn your central cognitive capacity. Be advised.
Note: myself a passionate Claude code user with multi agent parallel approach to dev. Plus 30 yrs of various oldskool dev experience. My blood and sugar and all tests are all within norm, and I bike and swim regularly. I’m not a major drinker and try to avoid alcohol in general. I count 25 trees outside my window and I’m not on any amphétamine pill such as Adderall.
Anyone else working on something like this or know of any projects attempting it?
Directionally if what you're doing is straightforward it's an amazing experience to be able to slap in an epic planning document and wake up the next day to it being "done", with a big asterisk that done-ness is directly proportional to how good of a spec and how good of a model you were using.
That being said, these days if you use Fable, slap in an epic planning document, and ask it to run a workflow (be sure to specify that subagents should use, say, Sonnet, or wave goodbye to your wallet), it's almost as good as gastown/gascity but far more predictable.
I've taken a bit longer than I wanted but it will be open sourced soon.
It's a durable orchestration engine that takes in specs/requirements and coordinates agents externally (meaning the engine drives the loop, not an agent) until the work is fully implemented/verified and reviewed.
It's meant to be used with any harness as basically the last step. You plan your work with whatever LLM you use and then hand off implementation to the engine (through an MCP server or other surfaces)
It can use your OpenAI/Anthropic subscriptions or any other provider and you can mix and match models across implementation and review in any way you want with fan out for parallel reviewers and more.
The goal is to produce high quality unsupervised code that matches your requirements and is reviewed throughout the implementation rather than at the end only, so that mistakes don't compound.
https://engine.build if you want to get notified when it releases.
A good coworker will admit not knowing something, or if unsure give their best guess but discuss its limitations and why they might be wrong.
Question: Has anyone experimented with using voice to directly prompt an LLM, without doing speech-to-text? If an LLM can pick up on the skeptical nuances in a person's response, it might be prompted not to be overconfident in its subsequent output.
Even projects that used to be challenging enough to impress people with your skills can now be built in 10 minutes with AI just by describing what you want. It's an incredible shift, but it also changes how I think about the craft and what it means to be a good engineer.
https://github.com/JuliusBrussee/caveman
It's for getting it to output shorter answers, but also could help with your burnout.
It's ok if you don't get it, but other people do, so can't generalize with your lack of comprehension.
But I also know that trying to optimize every second of my workday is a recipe for stress and eventually burnout. Nobody benefits from that.
So the key is to find the right balance between productivity and longevity.
If you feel stressed and overwhelmed then you are not in balance.
This is me. I hate the internet (and search) so much these days that I embrace anything that allows me to not Google a thing.
Review AI code line by line is like watch movies frame by frame, and is impossible, very difficult, terribly boring, or abandoned sooner or later.
Getting sent IM responses that are copy pasted LLM nonsense. Getting a massive PR to review that was generated overnight and the author didn't read it first.
I do not understand these complaints. Yes, those are the defaults and they're annoying, although the general public seems to like them. But you are not stuck with these. You can just tell the LLM how it should interact with you. If you're using any sort of harness beyond the chat window in a web browser, you can codify these instructions in a rules.md file or similar and have it automatically included in any new chat. It's not any harder than changing the default wallpaper or color scheme on your desktop operating system.
In reverse order, you can just tell the LLM to never use emojis. I don't like emphatic staccato fragments either, so I tell it to eschew the language of marketing and hype and stick to a factual and plain language, or to employ an academic tone. I explicitly instruct mine to ask clarifying questions whenever context is ambiguous and to push back on false assumptions or common misconceptions (by me). Hallucinationsa re the biggest problem of those you mention; it's not easy to totally eliminate them (for the same reason it's not easy to instruct people to not fall for scams or disinformation), but you can considerably reduce them by setting standards for citations.
I have ideas about reducing hallucinations over work material (ie a codebase) but am omitting them here as they are not fully thought out or tested.
Ouch. I love my local AI setup with Qwen, but that is a mismatch right there. That model is not the right match for that project. It's like trying to develop a major software solution by just throwing in hundreds of fresh junior programmers and have them spew out random code bits, while what you needed is a good PM, a great architect and a handfull of senior engineers. Might as well pack it in for a year until your model has grown into the ability for those rolls. There is a reason why Opus 4.6 and now Fable dramatically changed the SWE capabilities, and IMHO Qwen is not there yet.
Do does managnig agents.
At work, I ended up doing other chores, getting a lot more involved in projects I wouldn't even care to touch. Turns out it's kinda fun being the source of truth at work. I now have a clear sketch of what the company has done and what we can improve on.
Being able to fill in the gaps at a company that doesn't do much feels like a company within a company. Sure, its not Silicon valley, but it's still fun! And job security is guaranteed if you do a bit more than just play the ticket factory