HNHacker News
TopNewBestAskShowJobs

Yoric

7,362 karma · joined February 24, 2009

Hi, I'm David. If you have ever cursed at Promise or async/await in JavaScript, I'm sorry, I promise it's not (only) my fault! The same if you ever cursed at try! or ? in Rust!

Currently working on static analysis at GitLab, formerly quantum computing at Pasqal, safety tools at Element, performance at Mozilla, Rust contributor. Enjoys tech and product challenges, performance programming, safety guarantees, systems programming, programming language design, distributed programming, static analysis, compilers, formal methods, back-ends, databases, ...

Also, science vulgarization, storytelling, improv acting, ...

D.O.Teller+hn@gmail.com https://www.linkedin.com/in/davidteller/ https://github.com/Yoric https://yoric.github.io meet.hn/city/fr-Lyon

submissionscomments
Yoric··on The Normalization of Inexplicable Failures
> So I would say we are not normalizing failures (yet) but rather normalizing legacy.

Nice way of putting it.

Yoric··on The internet discovers TLA+. Now what?
Nah, jump straight to pi-calculus.
Yoric··on How to keep enjoying programming in a world of LLMs
Unrealistic deadlines, productivity being measured in PR counts and tokens spent, and generally a 100% focus towards having a software factory produce all code.

In our latest evaluation, one of the criteria was whether you trust AI, with "trust" being defined as letting the AI write all the code, without reviewing it. And of course, not trusting AI meaning that you were behind the curve.

Yoric··on Japan moves to tighten rules for foreigners
I'm confused. Aren't you actually describing racism?
Yoric··on How to keep enjoying programming in a world of LLMs
Oh, if I had the ability to do these things, agents would be much nicer to use.

Sadly, at my current company, this doesn't seem to be acceptable practice.

Yoric··on How to keep enjoying programming in a world of LLMs
> But you can still do that, much easier, with plain human words.

Not really? In my experience, to get anything precise done, you have to fight the LLM every step of the way. And then when you come back after a few days, you realize that it has overwritten the carefully crafted code or data structures.

Yoric··on How to keep enjoying programming in a world of LLMs
> And who likes planning, architecting, directing a team, steering and giving advice, reviewing code, designing interfaces and APIs. Coding became more mentally exciting.

I love these parts. But in my experience, LLMs break much of that.

Planning/architecting? Great. So far, I haven't found any agent that actually follows the plans set out, though. They get something wrong, and it snowballs from here.

Directing a team, steering and giving advice? Absolutely. Some of my greatest achievements involve mentoring. But human teams learn from their mistakes, grow up and contribute insights. Agents don't.

Reviewing code? Well, maybe not so much fun, but it's usually a good way to understand what's going on, and to share experience. Except with agents, you need to spend most of your brainpower seeing through the misleading comments and documentation and choices and sycophancy, and the agent never learns from its mistakes.

Designing interfaces and APIs? Absolutely. Yet every time I look at code modified by my agent, I see that the contracts (internal or public) have been broken by the latest edits.

In my experience, LLMs can be very useful, for refactorings and as learning and review assistants, and sometimes as replacement for missing documentation. But codegen is the worst way to use them.

Yoric··on Jev Based Code Review
As far as I can tell from just looking at the documentation, yeah, it looks like a "better than nothing" and "definitely better than LLMs for a few tasks" tool, and quite possibly something that I'd build into a demo, but I'd be really careful deploying it in any kind of product where exactness matters.
Yoric··on Jev Based Code Review
It's funny, because in my experience, the process often looks like:

1. Human gives high-level design. 2. Agent generates wrong code with misleading comments. 3. In further iterations, agent get mislead by said code and comments, ends up generating insane workarounds.

Yoric··on Jev Based Code Review
It is.

Sadly, I know (major) companies that insist it's the process that needs to be solved, because it improves velocity (for some definition of velocity that involves dropping pretty much all quality gates).

Yoric··on I am done with this shit
Not only that, but even when you get the LLM to give you the right answer, if you ask it to turn that answer into actual code committed to your repo, it usually destroys something else in the process.
Yoric··on I am done with this shit
Yes, I've concluded a few years ago that VC money has been really bad for software quality.
Yoric··on I am done with this shit
I think that tinkerers is the wrong term.

I consider myself a professional. I lean towards understanding problems, making sure that they're fixed once and for all, and being able to foresee issues that the users will encounter and that product cannot see. And while I enjoy coding, code is secondary to all this. Yet I'm in Hell.

We're losing our ability (and time) to understand problems, we prefer churning out patches fast than actually fixing them, we delegate meaningful choices to an AI that has no ability to predict issues, we remove the opportunity to review and we end up leaving testing to the end user. As far as I can judge, we have stopped building products and are now shipping glorified demos.

Also, I can't speak for other products, but on the type of features I work on, AI is blatantly incompetent, but our C-suite just refuses to believe it, assuring us that it's just us having difficulties with the transition.

Yoric··on I am done with this shit
I don't know about other companies, but my company was to a large extent turned into a bodyshop after we decided to go all in into AI. We laid off / reorganized away the team, then amped up the deadlines and promises, leaving the handful of survivors to cover up for so many features and tickets that we have no time to actually work on quality.

And in our twice-yearly self-improvement reports, we have to write how we can use AI further. To reach grade 4/4, you need to "trust" the AI.

Yoric··on Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences
> Now even though Claude seems smarter than any person I interact with on a daily basis, it is all suspect.

That... is definitely not my experience.

Yoric··on Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences
I used Google when it appeared, and I'm pretty sure that I never heard these sentences uttered out. If anything, everybody was wowed by PageRank.
Yoric··on Study: Young users (9 to 18Y) ditch Google for AI, with unknown consequences
I seem to recall headlines about Russia infecting AIs with propaganda, one or two years ago.

I don't remember actually reading the article, so I have no idea how serious that was, but I'd be surprised if this wasn't being studied actively by all major countries already. And that may be one of the core reasons for the China vs US AI competition.

Yoric··on Don't be the out of touch Kung Fu master
> Code review was difficult at the start due to the volume of code. One insight we gained midway through though is that the value humans bring to code review is judgment and business context. So we created a code-atlas skill that creates an artifact for PR reviewers. That artifact highlights the most important parts of the code.

Yeah, we did that, too.

But every time I end up, for some reason, digging up deep in the code, I realize that it's not nearly sufficient in our case.

Yoric··on Don't be the out of touch Kung Fu master
Benchmarks look good.

Reviews? They were the first casualty.

Yoric··on Don't be the out of touch Kung Fu master
In my experience, they are fairly good at catching stuff in human-written code, but they tend to gloss over agent-written code, just as human beings being lulled to complacency by superficially-looking competent code.
Yoric··on Don't be the out of touch Kung Fu master
Not GP, but I can.

I'm currently working on porting a mid-sized project to a new architecture, new programming language and of course adding new features.

Getting a new feature implemented is quite easy. You spend a few hours brainstorming specs with the agent, then ask it to implement it. This gives you extremely frequent code drops that add a new brick, add a new feature, etc. All of this with 100% code coverage (we also have mutation testing, strongly-typed code, standard and custom linters, etc.)

Then you look at the code. Code that has passed review, generally. You realize that the database schema has been broken silently, and that the agent has rewritten the tests or the golden fixtures to match. You realize that it has made assumptions that contradict the specifications and the product is going to break once it's in the hand of users. You realize that the 100% code coverage is essentially a convenient lie, because the code and tests have been written to make passing easy. You realize that none of the security golden rules have been followed, and that has managed to happen because the agent has somehow deactivated linting.

Why did it pass reviews? Well, because of deadlines. And because there is simply so much code (and so much unparsable/misleading documentation) that it's simply impossible to review all of this. And because things move so fast that nobody understands the CI pipeline anymore, and the explanations of the agent are convincing enough that surely, it knows better than you?

On the upside, bugfixing becomes so fast! Just add a new test, wait a few dozen minutes, and a new Merge Request appears. With equally convincing/misleading explanations, and something else broken.

After ~4 months, we had a bare bones deliverable, which we're now steadily expanding. If we had had to write the product manually, I suspect that it would have taken us at least one year, possibly two. So, that's the productivity increase. The productivity decrease is that what we have is not a product but a glorified demo, something that will work very nicely on the happy path, but on any other path, all bets are off.

Yoric··on Don't be the out of touch Kung Fu master
> just 20 years ago, all the most popular software shipped with NO tests. are you getting it?

Not really?

About 20 years ago, I was working on Firefox and we had millions of tests on CI. I was working on a host of other open source apps and they all had tests (most of them had no CI, of course).

Yoric··on Don't be the out of touch Kung Fu master
That is interesting, but... is it actually improving structure?

We're mid-way through a similar process at work. Rewriting a legacy app in a new language, with new architecture and new features.

And it's a mess.

We're at 10x loc (admittedly, the new programming language is more verbose than the old one), comments make no sense. Yes, we have ~100% coverage, but most of the tests are meaningless. The agent keeps removing our tests to replace them with tests that are easier to pass, breaking code invariants, removing all the engineered data structures and replacing them with stringly-typed code, etc.

And of course, given the number of LoC (and the fact that the agent rewrites so much code all the time), it's physically impossible that all of them were reviewed by a human being.

AI made it possible, insofar as upper management would never have greenlit the project without AI, but I can't escape the feeling that we're building on quicksands.

Yoric··on Don't be the out of touch Kung Fu master
> We are going from the era of manual, line-by-line mental model transcription to one where software engineers can focus on data structures, software architecture and algorithms.

That would be lovely.

It's sad that we're being forced to vibe code, though, because my day-to-day experience of that is that the agent does not respect the data structures I feed it, nor the architecture, nor the algorithms.

> Pure vibe coding is dull and unsustainable with current technology for all but the simplest systems; AI-assisted coding, on the other hand, rekindled my passion for computers.

Agreed.

Yoric··on Research acceleration: The view inside OpenAI
Who, me?
Yoric··on Research acceleration: The view inside OpenAI
> "Destroying the ecosystem" is just FUD.

Let's say it is. What about the rest of my paragraph?

> And if you don't find "average AI research intern" impressive, I'm not sure what to tell you. Have the goalposts moved so far that open ended problem solving at "average CS student fresh out of the uni" levels is suddenly trivial?

At this stage, I'm the one who doesn't know what to tell you. It took me years to grow from "research intern" into a competent researcher (and parallel years to turn into a competent developer). The research interns I've worked with were... vaguely useful, at best?

Yoric··on Research acceleration: The view inside OpenAI
Am I the only one who's a bit disappointed that we're spending trillions, destroying the ecosystem, drowning democracies and learning in slop, preparing a big financial crash, all of this to achieve an "average AI research intern"?

A long time ago, I used to be a (AI-adjacent) research intern, and frankly, I wouldn't trust any non-trivial task to that younger me. Fortunately, by opposition to an already trained LLM or agent, I have the ability to learn, so I eventually got better.

Yoric··on Meta Security Researcher's AI Agent Accidentally Deleted Her Emails
...when it's not entirely ignoring your guardrails, AGENTS.md and skills.
Yoric··on Continuous Diffusion Language Models (CDLM's)
As I understood it at the time, the rationale was that GPT2 made it too easy to produce content at scale and could easily be used for propaganda, phishing or other kinds of cons.

Which turned out to be true.

Yoric··on Continuous Diffusion Language Models (CDLM's)
Asking out of curiosity, because I have limited experience in that domain.

I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?

Page 1 of 34Next →