HNHacker News
TopNewBestAskShowJobs

dylanratcliffe

2 karma · joined November 2, 2021

submissionscomments
dylanratcliffe··on Ask HN: How do you guys stop coding agents from acting out of line?
It's often easier to catch and scold them than to prevent it. We have rule for example that agents (or anyone) can't touch autogenerated code. They still routinely do, but there is a CI job that re-generates everything and fails if there is a diff. Which the agents then see and fix. We have also use hooks for the same thing since they run locally and the feedback cycles are shorter.
dylanratcliffe··on Ask HN: What happens to code review process when using LLMs?
We want both to generate new features quicker than can be traditionally code reviewed, and still want to personally understand the architecture and code that is being generated.

This is possible, but not without being willing to change the level of abstraction at which you're working. We have 5 devs and changed from doing code review to plan review. You still peer-review what people are doing, but one layer of abstraction further back. Then when the PR gets raised there is a loop that just compares the implementation plan to the actual code and makes sure it was followed.

You still get to understand how things are getting solved i.e. "We're going to do this database migration in this way" but without reading the SQL. This is the plugin we wrote for it: https://github.com/until-dev/plugins

dylanratcliffe··on Ask HN: What eval harness holds up in practice, and what is still missing?
We've been using Gauntlet and enjoying it
dylanratcliffe··on Ask HN: How do you code review?
We did some measurement at the beginning of the year and found that our PRs were getting bigger and review time was coming down even without us doing anything, which is not a good sign.

We ended up moving peer review to the implementation plan rather than the PR, then having a loop that validates the code against the plan when the PR is raised. That way the agent gets a CI failure if it deviates from the plan, which it then fixes or acknowledges. Anything with no differences gets merged without human review, differences get approved by the original person who peer reviewed the plan.

I puled some stats the other day for a presentation I'm working on about what we did:

Matched plan on first pass: 25% (169/677) Had differences: 75% (508/677)

Differences per PR: Median 2 Mean 2.60 P90 6 Max 20

1,757 findings:

- Missing (skipped planned work): 44.6% of findings, 51.7% of PRs - Changed (done differently): 42.7% of findings, 51.3% of PRs - Beyond (extra, still in scope): 9.1% of findings, 19.8% of PRs - Scope (unplanned feature): 3.6% of findings, 6.6% of PRs

dylanratcliffe··on Is this the end of human code review?
From my reading of this it seems that the bug they were fixing was that there was a chat system which wasn't originally designed to be able to recover from a reload as it didn't track the state of both clients and didn't have any way for clients to catch up. The fix cost $2,430, took 3 days and touched ~50k lines, 189 files (7.1% of the codebase by LoC)

I totally agree that with the right systems we can do away with code review, and it's a solidly impressive achievement by the model/harness. But three days an nearly $2,500 to fix something that is a very foreseeable requirement if you'd done a little bit of planning doesn't seem like a terribly impressive outcome.

I think the more interesting tradeoff here is that the system they had to fix feels like it was "vibed" rather than "engineered" (or maybe the lack of state tracking was a deliberate tradeoff, but it doesn't really read that way) it sounds like they're at the "find out" stage TBH. But maybe spending $2,500 to fix things that would have been easily solved by some forethought is a tradeoff worth making if you're shipping things at ludicrous speed. Maybe you just ship 10 features and only fix the ones that get traction and it works...?