HNHacker News
TopNewBestAskShowJobs

dhorthy

1,188 karma · joined March 2, 2018

building @ humanlayer.com
submissionscomments
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
awesome - i have updated the post with a link to this thread!
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yeah my best articulation of taste is something i got from Jake Nations[1] while he was still at netflix -

"you know a bad pattern when you see it because at some point you were up at 2am debugging it"

taste is the hard-earned intuition about every anti-pattern and landmine that has blown up in your face since you started doing software

1 - https://www.youtube.com/watch?v=eIoohUmYpGI

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
we out here trying
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
it is absolutely wild to me that this keeps floating to top comment thread when the guy very clearly did not read even the first half
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
> I spent a week or so and like a billion+ tokens trying to refactor and save it. It just wasn't worth it.

this is exactly what happened to github.com/humanlayer/humanlayer - it was overslopped and we reset from scratch to build it right - spent 2 weeks in VS CODE - not even cursor, plumbing the core patterns from scratch. codebase is part of the prompt, yada yada

> I wish people would be pragmatic about this

that might be the tl;dr for the whole post haha

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
one thing I probably didn't mention is we do the program design having already done an in-depth codebase research, with current patterns and architecture surfaced - that actually seeds every step of the flow including even the product part -

but yes if you're working in very small slices then I think it's very feasible to skip program design and just review the code as you go, and resteer live. I do this all the time for tasks that are too big for a oneshot but on the smaller side overall.

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
and build it incrementally!

You don't have to build the entire software factory at once. You don't have to mastermind the whole future system, instead you're actually stacking and layering these small, isolated problems.

I think that's a really good approach to start getting value tomorrow or this week without saying, "I'm going to revolutionize how we ship."

It's just:

1. Find places where you can use agents.

2. Figure out where the right leverage points are for humans and where the right leverage points are for agents.

3. Just start building those things and plugging them into each other.

One day you'll wake up, and 80% of all of your stuff is automated.

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
I have had the phrase "programming is building a theory" spinning in my head for days, especially since watching this pragmatic engineer pod with Kent Beck https://www.youtube.com/watch?v=ddHQQtjIOpw
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
> Historically I know that the majority maintenance problems occur from slow continuous evolution of a system that it initially was never designed for. And the only way to address this was continuous system design.

yes exactly - this is what I'm advocating for - that you can't skip the system design, and that actually good system design goes down to the typedefs and object graph at the code level, not just mermaid charts and db schemas and service contracts.

I will highlight what a few others have said along the lines of "a sufficiently detailed spec IS code" - that is, to make the spec guaranteed to produce the code you want, the spec will look a lot like code (and will be roughly the same effort to review as the code itself anyway, saving you no time)

what I'm proposing is "how can you maximized the odds that the code WILL be good or close enough to good that its easy to get there, with the LEAST amount of human effort/attention" - how can you move fast without skipping what matters

https://haskellforall.com/2026/03/a-sufficiently-detailed-sp...

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
appreciate that context! I definitely did not mean to come out and say "its definitely not working" or anything, but would love to hear from y'all a retrospective on the ~5-6 month anniversary - what was right, what did we get wrong, etc
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
this is 100% right. you have to guard the codebase patterns with your life. because the codebase is part of the prompt.
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
> In order for coding with LLMs to go well, there has to be more rigor, more discipline, more good engineering hard-assedness. To reiterate, the teams seeing the best results with AI were already high-discipline and high-hygiene.

hard agree. But i don't think this is sufficient. Even formal verification has its limitations.

> AI works on data. The better the data, the better the likelihood of a desirable outcome. Code is data. If you have bad code, no matter how awesome the model you let loose on it, you can't get as good a result as if you had good code to start with. This principle has been well known in AI/ML circles since the 20th century.

hard agree. but also RL data is shaped differently than SFT data that has driven the majority of AI/ML innovations since ~2000, and its where there's so much room for innovation still. e.g. ImageNet was all just hand-labeled answer pairs.

> it's not a skill issue, it's an effort/laziness/rigor issue

I'm sorry but this feels like a semantic argument - the point of "skill issue" is "you didn't put in the effort or learn the techniques"

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yes exactly
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
i can't tell if this is a compliment or not
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yeah i like this and others in the thread mentioned that understanding RL and RLHF and the shape of the data is really important (at least the fundamentals, I'm sure there's quite complex industrialization of RL inside labs as Nathan Lambert says)
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
blake smith has a really good post on this - that mental alignment among the team is the primary purpose of code review - https://blakesmith.me/2015/02/09/code-review-essentials-for-...
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yeah that was another thing i hoped would pour through here - that deterministic systems are much better for evaluating quality (test, linters, cyclomatic complexity, etc) - but that we don't have such a system for code maintainability, at least not one that's widely accepted or adopted
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
fair point, this is the thing I struggled most to extract out while writing it - if you can propose an RL environment that penalizes a model for bad design, then I'm all ears - right now there's no fast oracle/verifier for this (as stated in the post)

My current evolving take on "how would you build such a thing" is you need to tee up a roadmap of 20 features and feed them to a model one at a time, so it can't design up front for what's coming.

That way if it builds the first 10 features and the codebase goes to slop, it get's penalized when it can't build features 11-20, or when those features take wayyy more tokens/time/cycles than a model that maintains a clean codebase can do.

This is how most real software is built by most teams - incrementally, getting feedback from users along the way, and steering goals in response.

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yeah I 100% agree - and I think the most popular coding agent workflows / skill kits are designed to pull those insights and intuition out of humans in a way that optimizes for the developer's experience building the plans or building the code, e.g.

- claude code plan mode - mattpocock/skills - obra/superpowers - research/plan/implement

etc etc

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
normative specifications can help, but the thesis here is that specs that define behavior of the product or even architecture are helpful but there's MORE that can be done and even though "program design" feels too in the weeds it's still essential if you care about maintainability
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
interesting - i'd say my main goal is to put the current "agentic software factory" hype in the historical context of "we've actually been rube-goldberging software deploys for a while now"
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yeah this is along the lines of what some friends of mine call "core vs. pragmatic modules" or even s/modules/codebase zones/

the idea that if you have a solid core and decoupled modules, you can have "zones" in your codebase where you allow the model to run wild and do a little slop, because you know the blast radius is contained

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
my perhaps controversial take is that opus 4.1 was smarter than 4.5 for complex engineering work, but 4.5 was faster and "squishier" - it responded better to simpler prompts, it read between the lines of user input better, and that this was really important to converting new users quickly
dhorthy··on Show HN: Hibernate and restore Claude Code sessions across reboots
this is cool tech and i like the use case but I'm always afraid of building against these internal JSONL apis that claude code uses since they change all the time / not a formal product interface
dhorthy··on Mercedes-Benz commits to bringing back physical buttons
functional programming taught us this decades ago. State is the root of all evil.

If the outcome of my interaction with the interface (e.g. tap a place on the screen) is a function of not just where i tap but the last 2-6 places i recently tapped (menus etc) suddenly you've added massive complexity and mental overhead.

can't wait to get back to a button that does the same thing every time every time i press it [1]

tesla screens, carplay, mercedes screens, its been getting worse for a while

1) I know in reality most are sliders or an on/off toggle but the point stands

dhorthy··on CodingFont: A game to help you pick a coding font
diabolical
dhorthy··on Anatomy of the .claude/ folder
I think one of the main examples that i saw in a swyx article a while back is that using the sort of ALL CAPS and *IMPORTANT* language that works decently with claude will actually detune the codex models and make them perform worse. I will see if I can find the post
dhorthy··on Get Shit Done: A meta-prompting, context engineering and spec-driven dev system
it is very hard for me to take seriously any system that is not proven for shipping production code in complex codebases that have been around for a while.

I've been down the "don't read the code" path and I can say it leads nowhere good.

I am perhaps talking my own book here, but I'd like to see more tools that brag about "shipped N real features to production" or "solved Y problem in large-10-year-old-codebase"

I'm not saying that coding agents can't do these things and such tools don't exist, I'm just afraid that counting 100k+ LOC that the author didn't read kind of fuels the "this is all hype-slop" argument rather than helping people discover the ways that coding agents can solve real and valuable problems.

dhorthy··on Verified Spec-Driven Development (VSDD)
software engineering is still software engineering.

just because you don't type out the characters doesn't mean you're not designing systems and thinking critically and leveraging your experience.

also: do we think this is written by ai? do we care anymore?

dhorthy··on A Brief History of Ralph
there is the theoretical "how the world should be" and there is the practical "what's working today" - decry the latter and wait around for the former at your peril
← PreviousPage 2 of 8Next →