Nice way of putting it.
7,362 karma · joined February 24, 2009
Currently working on static analysis at GitLab, formerly quantum computing at Pasqal, safety tools at Element, performance at Mozilla, Rust contributor. Enjoys tech and product challenges, performance programming, safety guarantees, systems programming, programming language design, distributed programming, static analysis, compilers, formal methods, back-ends, databases, ...
Also, science vulgarization, storytelling, improv acting, ...
D.O.Teller+hn@gmail.com https://www.linkedin.com/in/davidteller/ https://github.com/Yoric https://yoric.github.io meet.hn/city/fr-Lyon
Nice way of putting it.
In our latest evaluation, one of the criteria was whether you trust AI, with "trust" being defined as letting the AI write all the code, without reviewing it. And of course, not trusting AI meaning that you were behind the curve.
Sadly, at my current company, this doesn't seem to be acceptable practice.
Not really? In my experience, to get anything precise done, you have to fight the LLM every step of the way. And then when you come back after a few days, you realize that it has overwritten the carefully crafted code or data structures.
I love these parts. But in my experience, LLMs break much of that.
Planning/architecting? Great. So far, I haven't found any agent that actually follows the plans set out, though. They get something wrong, and it snowballs from here.
Directing a team, steering and giving advice? Absolutely. Some of my greatest achievements involve mentoring. But human teams learn from their mistakes, grow up and contribute insights. Agents don't.
Reviewing code? Well, maybe not so much fun, but it's usually a good way to understand what's going on, and to share experience. Except with agents, you need to spend most of your brainpower seeing through the misleading comments and documentation and choices and sycophancy, and the agent never learns from its mistakes.
Designing interfaces and APIs? Absolutely. Yet every time I look at code modified by my agent, I see that the contracts (internal or public) have been broken by the latest edits.
In my experience, LLMs can be very useful, for refactorings and as learning and review assistants, and sometimes as replacement for missing documentation. But codegen is the worst way to use them.
1. Human gives high-level design. 2. Agent generates wrong code with misleading comments. 3. In further iterations, agent get mislead by said code and comments, ends up generating insane workarounds.
Sadly, I know (major) companies that insist it's the process that needs to be solved, because it improves velocity (for some definition of velocity that involves dropping pretty much all quality gates).
I consider myself a professional. I lean towards understanding problems, making sure that they're fixed once and for all, and being able to foresee issues that the users will encounter and that product cannot see. And while I enjoy coding, code is secondary to all this. Yet I'm in Hell.
We're losing our ability (and time) to understand problems, we prefer churning out patches fast than actually fixing them, we delegate meaningful choices to an AI that has no ability to predict issues, we remove the opportunity to review and we end up leaving testing to the end user. As far as I can judge, we have stopped building products and are now shipping glorified demos.
Also, I can't speak for other products, but on the type of features I work on, AI is blatantly incompetent, but our C-suite just refuses to believe it, assuring us that it's just us having difficulties with the transition.
And in our twice-yearly self-improvement reports, we have to write how we can use AI further. To reach grade 4/4, you need to "trust" the AI.
That... is definitely not my experience.
I don't remember actually reading the article, so I have no idea how serious that was, but I'd be surprised if this wasn't being studied actively by all major countries already. And that may be one of the core reasons for the China vs US AI competition.
Yeah, we did that, too.
But every time I end up, for some reason, digging up deep in the code, I realize that it's not nearly sufficient in our case.
Reviews? They were the first casualty.
I'm currently working on porting a mid-sized project to a new architecture, new programming language and of course adding new features.
Getting a new feature implemented is quite easy. You spend a few hours brainstorming specs with the agent, then ask it to implement it. This gives you extremely frequent code drops that add a new brick, add a new feature, etc. All of this with 100% code coverage (we also have mutation testing, strongly-typed code, standard and custom linters, etc.)
Then you look at the code. Code that has passed review, generally. You realize that the database schema has been broken silently, and that the agent has rewritten the tests or the golden fixtures to match. You realize that it has made assumptions that contradict the specifications and the product is going to break once it's in the hand of users. You realize that the 100% code coverage is essentially a convenient lie, because the code and tests have been written to make passing easy. You realize that none of the security golden rules have been followed, and that has managed to happen because the agent has somehow deactivated linting.
Why did it pass reviews? Well, because of deadlines. And because there is simply so much code (and so much unparsable/misleading documentation) that it's simply impossible to review all of this. And because things move so fast that nobody understands the CI pipeline anymore, and the explanations of the agent are convincing enough that surely, it knows better than you?
On the upside, bugfixing becomes so fast! Just add a new test, wait a few dozen minutes, and a new Merge Request appears. With equally convincing/misleading explanations, and something else broken.
After ~4 months, we had a bare bones deliverable, which we're now steadily expanding. If we had had to write the product manually, I suspect that it would have taken us at least one year, possibly two. So, that's the productivity increase. The productivity decrease is that what we have is not a product but a glorified demo, something that will work very nicely on the happy path, but on any other path, all bets are off.
Not really?
About 20 years ago, I was working on Firefox and we had millions of tests on CI. I was working on a host of other open source apps and they all had tests (most of them had no CI, of course).
We're mid-way through a similar process at work. Rewriting a legacy app in a new language, with new architecture and new features.
And it's a mess.
We're at 10x loc (admittedly, the new programming language is more verbose than the old one), comments make no sense. Yes, we have ~100% coverage, but most of the tests are meaningless. The agent keeps removing our tests to replace them with tests that are easier to pass, breaking code invariants, removing all the engineered data structures and replacing them with stringly-typed code, etc.
And of course, given the number of LoC (and the fact that the agent rewrites so much code all the time), it's physically impossible that all of them were reviewed by a human being.
AI made it possible, insofar as upper management would never have greenlit the project without AI, but I can't escape the feeling that we're building on quicksands.
That would be lovely.
It's sad that we're being forced to vibe code, though, because my day-to-day experience of that is that the agent does not respect the data structures I feed it, nor the architecture, nor the algorithms.
> Pure vibe coding is dull and unsustainable with current technology for all but the simplest systems; AI-assisted coding, on the other hand, rekindled my passion for computers.
Agreed.
Let's say it is. What about the rest of my paragraph?
> And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?
At this stage, I'm the one who doesn't know what to tell you. It took me years to grow from "research intern" into a competent researcher (and parallel years to turn into a competent developer). The research interns I've worked with were... vaguely useful, at best?
A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM or agent, I have the ability to learn, so I eventually got better.
Which turned out to be true.
I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?