HNHacker News
TopNewBestAskShowJobs

dhorthy

1,188 karma · joined March 2, 2018

building @ humanlayer.com
submissionscomments
dhorthy··on AI;DR (AI; Didn't Read)
> AI-written emails where the commenter would rather see the prompt

this a thousand times this

LLMs are information transformers. if you're trying to take some rough idea and blow it out into something that "feels" substantive, you're just combining your high-signal idea / prompt with a bunch of meaningless bloat and noise from the model weights.

dhorthy··on Show HN: /show-me: agent skill for compact visual representations
they seem to be able to do the call-stack diff stuff pretty well still
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
so wait is the finding that most of those skills reduce pass rates against SCB? wild
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
why would that worry you?
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
agree, i think the implication is that low quality code is harder to change in the future
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
somebody get this man a curl-pipe-bash stat
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
oh i really like the idea of flipping around the order of checkpoints and comparing results. Could be an interesting way to increase/decrease difficulty even

i will look into how easy it would be to zip up some subset of the results without leaking anything...probably doable

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
Yeah I would hold that models don’t know how to simplify because most rl/benchmarks doesn’t penalize complexity
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
Yeah the main reason I skipped fable was because we have a ZDR with anthropic and I didn’t feel like spinning up another account to circumvent that. Next run will have fable and sol
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
i laughed at the pelican bit its good

yes the labs will always prioritize the vibeslop dopamine casino as far as I can tell - making the models useful and addictive for unsophisticated users, sometimes at the expense or at the very least at the ignorance of the needs of power users

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
I agree this is an option, and the next thing on my radar is to try with a more realistic "factory-shaped" harness where you have feedback from linters and other models after each coding episode that refines the architecture.

For readability specifically, I've found it hard to get the models to do this with prompting. If you've talked to opus/fable for a long time on prose writing you probably felt this too

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
> - what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is

this is a nicely succinct way to put this - a multi-dimensional space where no single metric is really useful

state space of the system is interesting too. I would guess that for any production software with dependencies like databases/third parties that might be too hard to measure, but if you can silo off parts of your system into bounded state machines, it may be a value metric on some module behind a clean interface.

I think the kubernetes control loop model is a great instance of this, a handful of scoped components that own a control loop across a well-defined state machine, that can operate / recover in the face of most network partitions or downtime - the promise of CRDTs but rather more a pragmatic approach to it

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
yes sol is still my daily driver for most coding tasks

I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)

but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked off.

> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 https://arxiv.org/html/2603.24755v1 > Get all the context, fetch any mentioned repos, and then propose a plan to me.

> i have an anthropic API key in .... > Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?

> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.

> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.

> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable

dhorthy··on Benchmarking Opus 5 on SlopCodeBench
i hope that is because you hate slop and not because you write it
dhorthy··on Benchmarking Opus 5 on SlopCodeBench
yeah this was just a start - the fastest cheapest thing we could try for a brand new model.

I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark

I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model

dhorthy··on If coding has been solved, why does software keep getting worse?
the one good thing about the current ios version is that it is the best it will ever be from this point forward. all future versions will be worse
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
I agree this rocks I do this almost daily
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
yes well said
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
don't forget the andon cord
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
i humbly disagree - horizontal means touching one plane of the stack across, vertical means cutting down through it and touching multiple layers https://en.wikipedia.org/wiki/Vertical_slice
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
i did a write up on fable while it was out - it can do big refactors, but it does not know what to change without human steering. For that, you need humans to know what to ask for.

lights off is still out for me

https://x.com/dexhorthy/status/2064747631885398231

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
its the optimizing utilization instead of overall throughput all over again. eli goldratt talked about this in the 1970s[1]. we still haven't learned

1 - https://en.wikipedia.org/wiki/The_Goal_(novel)

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
exactly. a great pr is a joy to review. we've found some success in agents generating static HTML walkthroughs that order the diffs in something other than GitHub's default alphanumeric ordering, but it can only go so far
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
i guess to clarify my contention:

generic prompting like "review this code" or "make the architecture better" will raise the floor but cannot come close to human-quality code without humans understanding what code exists and what to ask for.

Obviously cursor and other labs have a stake in this going one direction, but I don't buy it, and I don't think you should either.

In fact, "Rebuild sqlite from spec" has all the problems with every other benchmark that I cited - model knows the whole problem up front and never has to iterate on the pile of slop it created cheating its way to a solution.

In any case, politely, I think we're mostly arguing vibes here and I'm not sure its going to get anywhere.

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
(op here btw) - the question isn't "can models make code better" - its "left fully unattended, will they turn your codebase to slop over time"

from the footnotes (sorry if this got a little buried)

> yes of course you can get gpt-5.5 xhigh to do BRILLIANT refactors. But you had to tell it to do that. And to tell it to do that you had to understand your codebase well enough to know it needed doing. We're here talking about why lights-off wont work

i suppose this would also be a good place to ask if you fall in the "high stakes production systems" camp

> If you love vibe coding, please, go on vibing. I still vibe code lots of things, I just also maintain lots of production software (and through HumanLayer, help 1000s of other engineers do the same), so the rest of this is aimed at folks solving hard problems in complex codebases.

dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
the thesis of the post is that this is not true. fable can solve more problems but for complex systems it is not sufficiently diligent that you can turn the lights off.
dhorthy··on Why Software Factories Fail (or: harness engineering is not enough)
that's probably part of it but every enterprise in the world (even teams as small as 10-20 engineers) are paying per token, not with subscriptions. Claude code did make it into those orgs because people played with it at home with subscriptions first, but the bulk of that revenue is almost certainly coming from people paying per token, not the ones getting the ~90% subsidy on subscriptions
Page 1 of 8Next →