2,307 karma · joined August 15, 2012
Fixed.
-- Proverbs 16:16
It sets up your repo to ensure agents use a workflow which breaks your user requests down into separate beads, works on them serially, runs a judge agent after every bead is complete to apply code quality rules, and also strict static checks of your code. It's really helpful in extracting long, high-quality turns from the agent. It's what we used to build Offload[1].
0: https://github.com/imbue-ai/rust-bucket : A rusty bucket to carry your slop ;)
Papal Encyclicals[0] are solely authored by the Pope, even if there has been secular scholarship involved in the writing. It is never "presented" by anyone else, and to frame it as presented primarily by Christopher Olah "alongside" the pope is to betray an ignorance of what's officially going on.
Not sure how we arrived at the present title, "Anthropic co-founder to present AI encyclical alongside Pope Leo XIV", but it makes as much sense as "Iceberg nearly completes mainden voyage across Atlantic, with famous ship as passenger."
Of course, your response admits, "second to Rust", which I am guessing is an unspoken question in the grandparent's mind.
For your conceptual-model.md, do you find natural language is sufficient? Do you use pseudo code, Entity Relationship modeling, or anything like that?
And how do you go about round-tripping it from the source code, keeping it up to say with any changes, and yet at the right later of abstraction?
Remote digital fist bump, and prayers for you.
This is the reason why AI-assisted programming has not turned out to be the silver bullet we have been hoping for, at least yet. Muddled prompting by humans gets you the Homer Simpson car you wished for, that will eventually collapse under its own weight.
I've been thinking a lot about Programming as Theory Building [0] as the missing piece in AI-assisted engineering. Perhaps there are approaches which naturally focus on the essence while ignoring the accidents, but I'm still looking for them. Right now the state of the art I see ignores both accident and essence alike, and degrades the ability to make progress.
Please inform me if there are any approaches you know that work! And lest this sound pessimistic, far from it. This state of affairs is actually intoxicatingly motivating. Feels like we have found silver, and just need to start learning to mould bullets.
[0] Another classic required reading of the industry https://pages.cs.wisc.edu/~remzi/Naur.pdf
As a downside, the compile time is somewhat offset once you're using agents (and especially parallel agents) anyway. Since all of your edits cost a round-trip API call to a third party server, you can accept a slightly slower compile step.
The purpose of a sandbox should be understood to be limited to isolating changes to the inner state of the sandbox: filesystem, git, installed binaries like compilers, interpreters, checkers, running processes, etc.
In short anything that gets rebuilt when you rebuild the sandbox.
Harness to API control is an orthogonal surface, that may be reasoned about independently. You may initiate and control it from within the sandbox, but equally (and perhaps more) valid would be to do it from the outside.
Why would doing that lose control over the interface? Could you not secure the harnesses means to create outgoing connections and validate it that way?
I would argue that control from outside gives you MORE control as you could trust guardrails you've built outside the sandbox more than anything that's running in the same space where the agent has permission to execute arbitrary bash commands.
Latchkey does support some form of permissions management too.
I'm not following why this would this be the case? The purpose of calling the API is to get data or effect a state transition on some remote service, but I don't follow why the originating machine matters.
Or is your objection about auth?
I've heard many claims that because LLMs are tuned to specific harnesses, we should expect worse performance with novel architectures. That seems to make people reluctant to try to put effort into inventing them.
I'm really intrigued by your point on read-memory vs a dedicated read interface, because it is a real insight about success rates in harness design.
How did you come to the conclusion you did? Could you speak a little to the evaluations you ran, or the data or anecdotes you collected to validate that decision?
I'm also curious about the overall framing of the question, which I'll challenge with, does the agent have to have a where?
An agent could be modeled by a set of states and transitions. I don't think that there's anything inherently necessary about the current "one process claude" approach for harnesses, other than convenience. Why hasn't a fully distributed harness, built on functions and tables, gained more mindshare?
0: https://slatestarcodex.com/2018/10/30/sort-by-controversial/
You can use "API-style" pricing on these providers which is more transparent to costs. It's very likely to end up more than 200 a month, but the question is, are you going to see more than that in value?
For me, the answer is yes.
Hertz just thinking about it.
This is enough of a command reference that with it, agents are able to work with jj pretty well.
You mention doubling up on items for daytime/nighttime--are there any items where you half-up, since you only need one for both babies?