5,713 karma · joined August 11, 2008
jon at the domain above
Founder at pyn.ai and cultureamp.com
As a counterpoint -- it's rare I've seen a new UX issue fixed with a PR.
If you valued quality you could feature-flag some different improvements, get feedback and refine. AI is great at this. We have CS directly submitting Pull Requests now... and they're not junk, 95% of the the time they want things fixed/correct. And it's stuff that usually would sit on the backlog forever. The quality has gone up.
Your experience is representative I'm sure - but I do think there is a way to get this right and those that do will see a big upside.
> Were in for the golden age of cyberattacks, let me tell you.
Agree in full there.
Then throw away the ones you don’t like.
It also prevents reinforcement of your incoming pov.
I’ve found this has made me way way better at steering.
In coding I’ll do what I call a Battleship Prompt - simply just prompt 3 or more time with the same core prompt but strong framing (eg I need this done quickly versus come up with the most comprehensive solution). That’s really helped me learn and dial in how to get the right output.
No.
CLAUDE.md is just prompt text. Compaction rewrites prompt text.
If it matters, enforce it in other ways.
Once you get into higher power (laptops and up), switching and distribution get harder, so the advantages fade.
For bigger appliances (fridge, etc), AC is fine + practical.
Large organizations are making major decisions on the basis of it. Startups new and old will live and die by the shift that it's creating (is SaaS dead? Well investors will make it so). Mass engineering layoffs could be inevitable.
Sure. I vibe coded a thing is getting pretty tired. The rest? If anything we're not talking about it enough.
There are features you can skip safely behind feature flags or staged releases. As you push in you fine with the right tooling it can be a lot.
If you break it down often quite a bit can be deployed safely with minimal human intervention (depends naturally on the domain, but for a lot of systems).
I’m aiming to revamp the while process - I wrote a little on it here : https://jonathannen.com/building-towards-100-prs-a-day/
I realized recently that I've subconsciously routed-around merge conflicts as much as possible. My process has just subtly altered to make them less likely. To the point of which seeing a 3-way merge feels jarring. It's really only taking on AI tools that bought this to my attention.
FWIW I've struggled to get AI tools to handle merge conflicts well (especially rebase) for the same underlying reason.
In terms of tradeoffs, if you’re coming from the single event loop model, they’re pretty consistent with the rest of JS. Isolation-first, explicit sharing, fewer footguns. So I think the tradeoffs are the right tradeoffs.
FWIW, traditional threads have their own tradeoffs (especially around IO). In JS that’s mostly a non-issue, so the "I need 1000s of threads" case just doesn’t come up very often.
I'd assume that this will be a bit like JSON schemas - the decoding will eventually get smart enough to validate in the output in line with more complex rules.
Agree on the "behind it's back" too. I might make a change that in the case of "--fix" give the LLM the diff on the spot.
The other advantage I've not been able to quantify yet - I've been able to *remove* stuff from CLAUDE.md/etc in favor of lint rules. e.g. prefer ?? over || -- all the way through to "use our logging framework" -- as a lot of the nitpicks in my instructions were bits like this. This keeps the instructions to higher level architectural stuff.
If you’re exploring an idea or iterating, the roles can help break it down and understand your own requirements. Personally I do that “away” from the code though.
Definitely would entertain -- I do agree with your framing. I just think the article undersells the impact of fast+cheap codegen.
Lowering the cost of implementation will (has) expose new bottlenecks elsewhere. But imho many of those bottlenecks probably weren’t worth serious investment in solving before. The codegen change will shift that.
But there’s a more important difference: I can’t spin up 20 decent human programmers from my terminal.
The argument that "code was never the bottleneck" is genuinely appealing, but it hasn’t matched my experience at all. I’m getting through dramatically more work now. This is true for my colleagues too.
My non-technical niece recently built a pretty solid niche app with AI tools. That would have been inconceivable a few years ago.
The post captures something real about LLMs: the interface makes the interaction feel like a social exchange even when you know perfectly well it isn’t. Despite knowing better we attribute intention/emotion/feeling to the LLM. I felt that the most in her (somewhat bleak) sign off at the end.
1. You can make the script very specific for the skill and permission appropriately.
2. You can have the output of the script make clear to the LLM what to do. Lint fails? "Lint rules have failed. This is an important for reasons blah blah and you should do X before proceeding". Otherwise the Agent is too focused on smashing out the overall task and might opt route around the error. Note you can use this for successful cases too.
3. The output and token usage can be very specific what the agent needs. Saves context. My github comments script really just gives the comments + the necessary metadata, not much else.
The downsides of MCP all focus on (3), but the 1+2 can be really important too.
Anyway. Somewhat ironically, I use a wired set of headphones for this. It's not just the speakers that are better. I often get people remarking how much better the audio is on their end too... i.e. the cheap inline microphone.
The Visual Basic comparison is more salient. I've seen multiple rounds of "the end of programmers", including RAD tools, offshoring, various bubble-bursts, and now AI. Just because we've heard it before though, doesn't mean it's not true now. AI really is quite a transformative technology. But I do agree these tools have resulted in us having more software, and thus more software problems to manage.
The Alignment/Drift points are also interesting, but I think that this appeals to SWE's belief that that taste/discernment is stopping this happening in pre-AI times.
I buy into the meta-point which is that the engineering role has shifted. Opening the floodgates on code will just reveal bottlenecks elsewhere (especially as AI's ability in coding is three steps ahead and accelerating). Rebuilding that delivery pipeline is the engineering challenge.
So for me this is a pretty huge change as the ceiling on a single prompt just jumped considerably. I'm replaying some of my less effective prompts today to see the impact.
I think that's ended up with a bit of a mess.
1. You might be speeding up something that is inherently not productive (the "faster horses" trope). I see companies using AI to generate performance reviews. Same company using AI to summarize all the new performance stuff they're getting. All that's happening is amplified busywork (there is real work in there, but questionable if it's improved).
2. Some things are zero sum. If you're not using AI for marketing you might fall behind. So you adopt these tools, but attention/etc are limited. There is no net gain, just competition.
3. You might speed one part up (typing code), but then other parts of your pipeline quickly become constraints. It might be a long time before we're able to adapt the end-to-end process. This is amplified by coding tools being three strides ahead.
4. Then there are actual productivity improvements. One of these PRs could have been "translate this to German". That could be one PR but a whole step-change for the business.
So much of what is happening falls in buckets 1+2+3. I don't think we've really got into the meat of 4 yet.
My reaction is more to the broader tone of some of these discussions. In my experience engineering cultures can become quite dogmatic or obstructive, and that can block improvements just as much as the opposite problem.
At our definitely-non-Meta-scale, we’ve been experimenting with letting more of the team get their own PRs up with LLM help. Overall it’s been pretty transformative. Interestingly, people tend to work on QoL and polish improvements that many SWE workflows often don’t prioritise or have time for.
There are outliers of course, but we learn, revert, and move on. If the outcome somewhere like Meta is PMs building nonsense, that feels more like a deeper systemic issue than something inherent to opening up the codebase.
I love writing software. I love that others are now getting to share this.
I think the issues here are valid. Equally there is lots of hard engineering work to reduce these issues. That's where I'm putting more energy.
My scale is decidedly non-Meta, but we're investing to make the whole team able to get their own PRs up. It's not been without it's bumps, but on the whole I think it's been transformative for everyone.