HNHacker News
TopNewBestAskShowJobs

dhorthy

1,188 karma · joined March 2, 2018

building @ humanlayer.com
submissionscomments
dhorthy··on Show HN: Vibe Kanban – Kanban board to manage your AI coding agents
i feel so strongly that this will rapidly become true over the next 6 months. if you don't believe me check out Sean Grove's talk from mid June - https://www.youtube.com/watch?v=8rABwKRsec4
dhorthy··on Spec Engineering and the New Code – Sgrove from OpenAI [video]
my personal take following work of AmpCode folks, Claude Code team, and our internal process evolution - spec development will replace 99% of code development within 3 months
dhorthy··on What would a Kubernetes 2.0 look like
Overall I like this but confused about this part

“Yaml doesn’t enforce types but HCL does”

Is the same schema-based validation that is 1) possible client-side with HCL and 2) enforced server-side by k8s not also trivial to enforce client side in an ide?

dhorthy··on Show HN: SecureBuild – Zero-CVE Images That Pay OSS Projects
this looks cool - your homepage video should open with what it is though!
dhorthy··on Hyper Typing
Great post and def needs to be said. Won’t name names but there are some VERY popular ts library that are guilty of this. Once you have type constructor constructors I tend to bail and just cast things as described here. But then you lose guardrails, and confidence that you’re consuming the library properly
dhorthy··on Hyper Typing
This guy gets it
dhorthy··on Show HN: AG-UI Protocol – Bring Agents into Frontend Applications
I had kinda wondered about this for a while - I called it "MWP - model workload protocol" - a client-agnostic (sms, whatsapp, in-app, react components) way to display what an agent is doing

- working - thinking - calling tools - hit an error - needs human input - needs human approval

etc etc etc

dhorthy··on Tilt: dev environment as code
Been using tilt as a make alternative for years. Great tooling, even as just file watch + pythonic syntax for running tests, etc.

Obvs the real magic is the live syncing patches into remote containers though

dhorthy··on Come for the Tool, Stay for the Network
this is cool. curious what other examples folks have seen. Probably lots of content creation/management tools in general - figma, github, etc.

Any examples of this in lower level devtools? I think of things like K8s as "the real value was always in the standardization" but maybe it really did just start as "hey this is a cool way to run software"

dhorthy··on How to think about agent frameworks
this is probably good advice, but seems a little off to use "customer" in an article that feels like its trying to appeal to everyone, not just customers of langchain

> If you do use a framework, ensure you understand the underlying code. Incorrect assumptions about what's under the hood are a common source of customer error.

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
*probably :)
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
control is good, and determinism is good - while the primary goal is to convince people "don't give up too much control" - there is a secondary which is: THESE are the places where it makes sense to give up some control
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
yeah. the issue is when you're baked into a tool stack/framework where you cant go customize in that 1% of cases. A lot of tools try to get the right abstractions where you can "customize everything you would want to" but they miss the mark in some cases
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
i'd be interested to check it out

here's a take, I adapted this from someone on the notebookLM team on swyx's podcast

> the only way to build really impressive experiences in AI, is to find something right at the edge of the model's capability, and to get it right consistently.

So in order to build something very good / better than the rest, you will always benefit from being able to bring in every optimization you can.

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
interesting - I think I have to side with the Boundary (YC W23) folks on this one - if you want bleeding edge performance, you need to be able to open the box and hack on the insides.

I don't agree fully with this article https://www.chrismdp.com/beyond-prompting/ but the comparison of punchards -> assembly -> c -> higher langs is quite useful here

I just don't know when we'll get the right abstraction - i don't think langchain or dspy are the "C programming language" of AI yet (they could get there!).

For now I'll stick to my "close to the metal" workbench where I can inspect tokens, reorder special tokens like system/user/JSON, and dynamically keep up with the idiosyncrasies of new models without being locked up waiting for library support.

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
this guide is great, i liked the "chat interfaces are dumb" take - totally agree. AI-based UIs have a very long way to go
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
excellent work on BAML and love it as a building block for agents
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
Typed outputs from an LLM is a game changer!
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
Yeah definitely. I think the pattern I see people using most is “start with slow, expensive, but low dev effort, and then refine overtime as you fine speed/quality/cost bottlenecks worth investing in”
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
cool thing about open source is you can work on whatever you want, and it’s the best way to meet people who do similar work for their day job as well
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
I link in a few places to https://github.com/got-agents/agents where I have a few of these real agents
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
this guy gets it
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
funny you say

> I have learned 80% the hard way

because the other working title for this was "Agents the Hard Way" (in the spirit of https://github.com/kelseyhightower/kubernetes-the-hard-way)

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
This is a great bit of feedback - what kinda of use cases do you think would make sense?

Definitely wanna evolve this in the open with the community

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
They kept telling me automation would do my chores so we could spend more time on writing and art. I write less and still have to do my own laundry
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
THIS so much. People are like "why human supervision when we can have agent supervsion" and always respond

> look if you don't trust the LLM to make the thing right in the first place, how are you gonna PROBABLY THE SAME LLM to fix it?

yes I know multiple passes improves performance, but it doesn't guarantee anything. for a lot of tool you might wanna call, 90% or even 99% accuracy isn't enough

dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
i like the terminal UI and otel integrations - what tasks are you using this for today?
dhorthy··on 12-factor Agents: Patterns of reliable LLM applications
i have seen a ton of good ones, and they all have ups and downs. I think rather than focusing on frameworks though, I'm trying to dig into what goes into them, and what's the tradeoff if you try to build most of it yourself instead

but since you asked, to name a few

- ts: mastra, gensx, vercel ai, many others! - python: crew, langgraph, many others!

dhorthy··on Google will let companies run Gemini models in their own data centers
i'll add that on-prem is getting 10-100x easier than it was 10-20 years ago (still very hard), and "i want to run this in my own datacenter" is becoming accessible to much smaller companies than just F500 enterprises
dhorthy··on Reverse engineering OpenAI code execution to make it run C and JavaScript
I think most code sandboxes like e2b etc use Jupyter kernels because they come with nice built in stuff for rendering matplotlib charts, pandas dataframes, etc
← PreviousPage 5 of 8Next →