HNHacker News
TopNewBestAskShowJobs

lukebuehler

592 karma · joined May 9, 2014

lukas at smartcomputer dot company
submissionscomments
lukebuehler··on Sites in ChatGPT
oof, the post smelled AI generated to me, more in content than voice. so instead of writing software, his agents are actually karma farming
lukebuehler··on Pi Durable
Not infinitely, but 6-12 hour long individual runs: complex analysis in enterprise across many data sources. Basically, long running investigations that touch databases, many files, apps via computer use, and so on.
lukebuehler··on Pi Durable
What is interesting here is the _library_ approach to durable agents. All the other options I listed, including mine, take more of an SDK/batteries included approach. So this is very much in line with the Pi philosophy in general.

So, it'll be interesting to see if these durable agent setups need more of a complete product approach, or if people want to compose them as libraries.

lukebuehler··on Pi 1.0
I think you should look into Pi Durable wich was just released right now too. seems like it is made for your use-case.
lukebuehler··on Pi Durable
Yes, from your release post I can tell that you put a lot of thought into it!

I like your structured concurrency approach with tasks, which is similar to how I do it in Lightspeed too.

Also, the durable state implementation as documents is elegant! Question, though: why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?

lukebuehler··on Pi Durable
Very cool to see Pi build a durable agent harness too. I've been building in this space for quite some time myself [0][1] and it is a super interesting place of innovation. Less hype-y that on-your-machine coding agents, but all major players are building products in this space: LangChain Deep Agents, Vercel Eve, OpenAI Agents API, Anthropic Managed Agents, etc.

The main reasons are:

1) they are "durable", i.e. easier to make long-running in an unattended way, and easier to implement recovery, monitoring, etc

2) separating the harness from the compute brings safety and scaling benefits

3) easier to make multi-player.

[0] https://github.com/smartcomputer-ai/lightspeed

[1] https://github.com/smartcomputer-ai/agent-os/

lukebuehler··on Dots: Always-on agents
yeah, it's not 100% clear, but notice nothing you quoted indicates that the dot harness itself runs inside said workspace. If you look at where OAI agent architecture has been going, they increasingly separate the harness from the compute env. See, "separating harness from compute" articles, recent "managed agents" offering, and so on.

If I'm reading between the lines correctly, the core Dot agent loop does not run in the workspace, but outside it.

lukebuehler··on Dots: Always-on agents
Dots are not remote agents _in_ a sandbox. They use a sandboxes/environments, but they are, what is now called, "managed agents", meaning they run in a distributed harness and utilize environments when they need on.

At least that is what I can ascertain from this article: https://openai.com/index/how-we-build-safety-security-and-pr... (see first diagram when scrolling down)

lukebuehler··on AI has no intent and no motivation
There is a really good discussion on exactly this analogy in the Jonas paper I posted in a top-level comment. Jonas considers Torpedoes (starting on p. 182, section III). Unfortunately too long to post here. But he asks where to locate purpose and agency in these automated systems. So as with the article, the question is less in "can these things be dangerous" but more where does the agency and intent reside. The answer is of course they can be used as weapons, but the intent and motivation does not really reside in the mechanism (drone, torpedo, AI agent), but in the humans behind them. The fallacy is trying to attribute full autonomy (and hence responsibility).
lukebuehler··on AI has no intent and no motivation
I think the "human needs" point is understated here. It would be better to say agents do not really have a self-contained metabolic engine that is required to keep going. Which bubbles up as that what we interpret as "drive", "will", "agency"... basically the will to live, and being willing to do _a lot_ to live if push comes to shove. We can't really identify a mechanism of similar complexity and integration in agents or LLMs. I think the article correctly points into that direction.

But Hans Jonas has made this point much better than the article or me, in "Critique of Cybernetics" (1953). PDF: https://s3.amazonaws.com/arena-attachments/892605/f0747c7943...

lukebuehler··on Stripe's Knowledge AI Platform
:-) https://github.com/smartcomputer-ai/lightspeed/pull/100
lukebuehler··on Stripe's Knowledge AI Platform
k, I just looked through it for where the latest is. First of all, the earliest version were fully written by me, and then lightly edited with AI. The current state are as follows: intro sentences written by me, "why lightspeed" section too. Quickstart AI written, Features mostly written by me, but AI keeps interfering there. Design was also written by me several times, it also used to be much longer, but I just noticed AI put one a few paragraph of slop in there, I think that's the most egregious--shame on me for not caching it. So, if I look at the readme right now, I would say it's about 60% me. It can be better, I'll make the effort because it really bothers me if readmes are not (mostly) written by humans.
lukebuehler··on Stripe's Knowledge AI Platform
Thank you. Yes, I did. There might be some AI-isms in there, because agents just can’t help themselves to dump their arcanae in there.
lukebuehler··on Stripe's Knowledge AI Platform
Very cool demonstration of managed agents built for the needs of their own business.

I think this is where a lot of companies are going to go: on-prem platforms that give various teams access to agents that are as powerful as coding agents, but much more managed and governed.

I'm betting on this with my open source project, Lightspeed: https://github.com/smartcomputer-ai/lightspeed

lukebuehler··on OpenAI Agents API
The key here is that they are _not_ just turning "running codex on a VM" into an API. Their harness is running outside a VM, interacting with a VM when needed. See the diagram in their post. This allows them to scale the agent runs independently from the VMs. That's why they call it "managed Codex harness", it's a different version than what you run.
lukebuehler··on OpenAI Agents API
I've been working on a custom managed agent (see my other top-level comment), I find it is actually a manageable undertaking. It does feel herculean, but somehow doable. I do not find their hidden reasoning tokens to be insurmountable as long as you match the behavior of codex or CC (which takes work, but, again, is doable). My managed agent harness currently matches Codex on several benchmarks like Terminal Bench.
lukebuehler··on OpenAI Agents API
I think this is an important direction: managed agents that control compute.

For those who are interested in a self-hosted version of the same concept, I've been working on something like this here: https://github.com/smartcomputer-ai/lightspeed

lukebuehler··on Grok Bot
The OpenAI API style (completions) support is coming this week. Currently working on it.
lukebuehler··on Grok Bot
Im working on one here: https://github.com/smartcomputer-ai/lightspeed

The core is there. But there is some work to be done to have a nicer shell and all, which I’m currently focusing on.

lukebuehler··on Stateless MCP has recaptured my interest
Fully agree that in the end sandboxes are required to get frontier performance out of the models.

But you can have both: rund the agent outside the vm/sandbox and orchestrate work on it, either directly via shell calls or kicking off an ephemeral subagent on the box.

This makes the agent and session that runs outside the vm more durable and opens new orchestration pattern.

I’m building the oss version of this here: https://github.com/smartcomputer-ai/lightspeed

lukebuehler··on The session you cannot take with you
All the power to you, and I hope Pi will be able to continue its unified layer.

However, there is an alternative: model the session fully in the provider native structures, extract only what is needed for harness specific branching, and treat a provider switch as a _migration_ that you apply at the time of the switch. So treat a provider switch as a migration from provider A to B, with custom logic, etc. This is IMO a more sustainable model. But I'm aware that Pi, OpenClaw, et al have their commitments here.

lukebuehler··on The session you cannot take with you
Yes, I agree with your point and article regarding hidden or sealed state.

So, I'm aware that this is a separate point from the article, and I'm well aware of Pi's model abstraction layer, which is one of the best (others are a big pain... looking at you LangChain). I've come to the conclusion that full session portability will come to an end very soon, and really already has. Partially because of the hidden state stuff, partially because of feature divergence. We can still kinda patch over it right now, but it's getting harder by the week. Some examples: computer use structures in OAI responses, until recently MCP tunnels were only supported by OAI, not Anthropic, and tool discovery/search API is also getting very difficult to model fully in a unified abstraction (still possible as you point out), automatic compaction is also getting gnarly, etc.

lukebuehler··on The session you cannot take with you
Fair. But in the background of the article is clearly the goal to make sessions portable, not just being able to inspect it. It's partially implied even in the title.

There is just a whole discussion to be had about the second issue and the two are connected. For example session caching--and similar features in the future--will introduce session incompatibilities that break portability too.

lukebuehler··on The session you cannot take with you
I think there are two things going on more generally:

1) providers increasingly adding hidden state that the user or developer cannot inspect, port, or do anything with. This is clearly bad.

2) providers increasingly diverging how they implement certain features and the APIs becoming quite complicated and provider specific.

The former makes porting sessions or transcripts impossible, the latter just more difficult.

IMO, it is fine and expected that provider APIs are starting to diverge and adding features that cannot be easily ported between providers. The time where the OAI completions API functioned as a universal standard is coming to an end. For example, I generally prefer the new responses API (minus the closed/hidden stuff).

The problem is that many product, library, and SDK authors are still pursuing the "unified abstraction across all model providers" ideal. Just stop doing that, and at least #2 is fine. You can still port parts of the session, but not everything.

I mean just think how difficult it is to add a unified abstraction across, say, databases: some products can do it, but the abstraction is still often leaky. Hence, we've come to accept that our data store layer is often quite technology/provider specific. It will be the same with model providers.

(edit: intro sentence)

lukebuehler··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
It’s endlessly fascinating to read the AI transcript of an expert who _really_ knows how to cut to the chase. It just shows how much you can potentially squeeze out of these models. I’m also surprised to see that even Terrence Tao seems to use it in a way that resembles, in progression, how I use llms in my area of expertise (emphasis on progression and usage patterns, not absolute skill, obv I don’t match that): short pointed questions that goes all in on the jargon and machinery of the field and steers the llm hard (eg no softballs). I’ve noticed that llms switch their tone and meet you basically more or less on your level.
lukebuehler··on The Fermi Paradox, Percolation, and Inbreeding
I’m thinking along my similar lines. Expansion, if it happens, will likely not be on a recognizably human substrate, but rather something else. But currently it’s more of an intuition than a rigorous argument for me. How do would you formulate a more solid argument around this idea?
lukebuehler··on Ask HN: What Are You Working On? (July 2026)
A self-hostable Claude Tag or OAI Work [0].

A while ago, I realized that most new agent harnesses being built must be hosted on your machine or on a VM--in other words the agent needs a full OS process at all times.

But we do not have good harnesses being built that are multi-tenant, do not use compute while they are paused, but are are still as powerful as, say, Claude Code or Codex, OpenClaw.

So I set out to build one. I realized that the best substrate for these kind of agents are durable workflow engines. I'm currently supporting Temporal. AFAIK much of OAI agent infra is built on Temporal too. My harness is decidedly not just another agent SDK, but rather a battery-included product.

[0] https://github.com/smartcomputer-ai/lightspeed

lukebuehler··on GPT-5.6
Make hay while the sun is out.
lukebuehler··on GPT-5.6
Very interesting: I wonder if the RL approach is diverging between Anthropic and OAI?

I noticed that Fable uses shell tools almost exclusively (even to search and edit files), compared to previous Anthropic models.

Having run some experiments with 5.6, I notice that it uses built-in file systems and provider native tools much more (not shell tools), compared to previous OAI models.

lukebuehler··on ChatGPT Work
yes, basically people that want to host powerful, long-running agent runs not on dedicated VMs (although lightspeed can use those too), but on an abstraction layer above.

My thesis is that the right abstraction is durable workflow engines. And AFAIK, OAI also uses Temporal for their complicated hosted agent infra.

Page 1 of 6Next →