It does not pool memory or split one inference request across machines. Adding machines increases parallel throughput but won't let you run bigger models.
556 karma · joined June 18, 2019
It does not pool memory or split one inference request across machines. Adding machines increases parallel throughput but won't let you run bigger models.
> models often determine their answers based on implicit biases tied to question templates, then construct reasoning chains to justify their predetermined conclusions > its reasoning was correct right until the final step (Yes/No answer)
This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.
> Fairphones lack the updates and hardware-based security features expected by GrapheneOS.
I've done that before and the code was a mess. It works at the beginning but APIs do much more than piping data from the database. When you start dealing with ACL, external calls, code reuse, etc. It's just nice to have all the tools available to you from something like Python or Go.
I would much rather have someone on my team who ships less but whose work I can trust than someone much faster whose changes leave me wondering what problems we’re going to discover later.
And when production breaks (and it will), I need the person who made the change to actually understand it well enough to help fix it, instead of showing up with no idea what is going on.
You can obviously be an asshole about how you do it but I don’t think pushing back makes someone toxic.
You need to be flexible and compromise when the business trade-off makes sense. But you also need a backbone. If you think something is going to cause real problems, bringing it up is part of the job.
If the limiting factor is the rate at which experienced engineers can understand and validate changes, you have three options: generate less, find a genuinely better way to validate or accept lower quality.
We had tests, CI, code review, QA, architectural reviews, etc. None of those disappeared. But they were designed for a world where producing such a large amount of change was impossible.
Tests don’t solve that. Tests can tell you that the behaviours you thought to test still work. How many times have you had a completely green CI with 100% coverage and still shipped a bug?
But do the people still understand what the system is doing and why?
You can have code that perfectly follows every standard and passes every test while gradually building a system nobody has a mental model of.
> an engineer only has to take a cursory look at the code
If that means you’ve automated away checking syntax and implementation, great. But we already had that before. If it means nobody needs to understand the change anymore, then this is exactly the risk I'm talking about in the article.
On a large system, the customer being happy today isn’t enough. You need other engineers to be able to understand the system.
Have you ever been on call and been woken up in the middle of the night to fix a production incident in a system you didn’t write?
If everything you build is small, isolated and easy to replace (basically fire-and-forget), then yeah... who cares? Ship the ugly thing, get paid and move on.
If you’re going to be working on something for the next 5+ years, you should definitely spend some time thinking about what you’re doing.
If you genuinely understand the resulting system, then that's fine. That’s not the behaviour I’m criticising.
This was already happening before. But now, a lot more companies that previously might have taken many years to reach an unmaintainable state can now get there in just a few months.
If you don’t understand the system, don’t know what the code is doing and you’re mostly prompting an LLM to make the decisions and implementation for you, what exactly is your contribution?
The ability to operate the tool isn’t much of a moat if everyone else has access to the same tool.
I don’t review assembly produced by a compiler because the compiler isn’t deciding what my system should do. It’s translating a program whose semantics were already specified. More importantly, that translation is deterministic.
A compiler takes a human-specified program and translates it into another representation while preserving its semantics.
If in five years I can give an agent a complete specification and reliably verify the resulting machine code against it, then sure, reviewing code may become obsolete and I’d happily stop doing it.
Also, keeping changes small isn’t just about making individual lines readable. It limits blast radius, makes behaviour easier to reason about, isolates mistakes, makes changes easier to revert and so many other things. None of those properties suddenly become obsolete because code generation got faster.
You’re deliberately constraining the output so changes stay small, human-reviewable and reversible. That’s good engineering culture.
It isn’t all or nothing. I use AI heavily but that doesn’t mean I have to pretend there aren’t serious problems with how it’s being used.
Whether people will buy AI-generated software is a different question. The problem I’m describing already exists inside businesses with paying customers.
I’m talking about established teams working on products that already have users and make money, where AI lets individual engineers introduce changes faster than the rest of the team can properly understand and review them.
I don’t have any issue with intentional debt when you understand the trade-off and have a clear payoff plan.
See "AI Won't Take Your Job, Someone Using AI Will" > There's no lucrative middle ground where 'AI whispering' is a high-value skill.
https://blog.florianherrengt.com/vibe-coder-career-path.html
You still have to understand the change yourself if you’re going to take responsibility for approving it.
That work is fundamentally much slower than generating the code, even with the help of AI.
We’ve made producing a large change extremely cheap and fast. We haven’t found an equivalent shortcut for building a correct mental model of what that change does, how it interacts with the rest of the system and whether the decisions behind it are actually sound.
Maybe one day we'll find one. As of today, I don’t think we have.
AI massively increases the speed and scale at which bad engineers can do damage while fixing it is still slow, difficult work.
We still teach arithmetic and algebra despite computers being vastly better at calculation. We teach spelling, grammar and essay writing. We even teach history and geography while everyone permanently has a device in their pocket that can look up almost any fact in seconds.
How else would you be developing the mental models required to understand, question and verify anything?
Although, I don’t think immigration or country of origin is particularly relevant here. I never worked in the US but there are plenty of "assembly line" developers here in London too.
If your job is essentially taking a ticket and turning it into code without contributing much else, then yes, I think that category of developer is going to be decimated.
That applies equally to everyone.
You are still an engineer but you’ve delegated your technical judgement to an LLM. You just stopped doing the most important part of your engineering job.
> Fixing it would require such a colossal amount of work that it would be impossible to even start justifying it to anyone in management. > And what are you even thinking about? It would end up in the exact same state again in just a few months anyway.
The management side of this probably deserves a whole post of its own. There’s only so far I can take each tangent before the original post becomes too long for anyone to finish.