HNHacker News
TopNewBestAskShowJobs

afshinmeh

4,379 karma · joined January 31, 2013

afshin.io
submissionscomments
afshinmeh··on Please Do Not Vibe Fuck Up This Software
Genuinely wondering though: is the problem that the patch was vibe coded, or is that no one reviewed the changes?
afshinmeh··on Show HN: VT Code – open-source terminal coding agent in Rust
what does "LLM-native code understanding" mean in this context?
afshinmeh··on SQLite Is a Library of Congress Recommended Storage Format
I love SQLite and thanks for sharing it but there should be a "(2018)" at the end in the title:

> As of this writing (2018-05-29) the only other recommended storage formats for datasets are XML, JSON, and CSV.

afshinmeh··on Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
That's exactly how I tried to address that problem with https://github.com/afshinm/zerobox -- you control what network access (e.g. `--deny-net *.amazonaws.com`) your agent has and you also get snapshotting out of the box.

That said, using LakeFS is probably a better long term solution and I like this approach.

afshinmeh··on DAG Workflow Engine
I'm mainly looking at Rust based projects and haven't been able to find something to use out of the box, without hacky RPC/Shell execs. Curious if you have any suggestions?
afshinmeh··on DAG Workflow Engine
Sort of. My thinking is that the input to define the workflow should be anything you prefer to use (TS, Go, YAML, etc.) and the orchestrator's job is to model that and execute the job, given your deployment model.
afshinmeh··on DAG Workflow Engine
Yeah, that makes sense. I looked at a few workflow orchestrators and I'm building something that I will release soon, but my thinking is that the "workflow engine" should be an abstraction that takes the input and executes the steps. "What" you use to define that workflow is probably the SDK layer though, but I can certainly see the value in using type safe code to define as opposed to a YAML file.

I'm mainly focusing on the portability aspect of it (e.g. use TS/Python/etc. to define the workflow/steps or just simple a simple YAML file).

afshinmeh··on DAG Workflow Engine
Curious, what format would you prefer to use to represent a workflow instead of YAML?
afshinmeh··on DAG Workflow Engine
https://github.com/vivekg13186/Daisy-DAG/blob/main/backend/s...
afshinmeh··on The agent harness belongs outside the sandbox
I wonder though, what about cases where you have multiple agents or LLM backends and the credentials is shared between all of them?
afshinmeh··on The agent harness belongs outside the sandbox
Agreed and it's a pattern that OpenAI suggested a few days ago, too [1]. I also built a cross platform process level sandboxing that uses parts of OpenAI Codex for the same purpose [2]

[1] https://openai.com/index/the-next-evolution-of-the-agents-sd...

[2] https://github.com/afshinm/zerobox

afshinmeh··on Flue is a TypeScript framework for building the next generation of agents
Vibe coding aside [1], it's very interesting software projects these days don't really care about adding a single test [2].

[1]: https://github.com/withastro/flue/blob/8fdf8e0e9df5bd33c3120...

[2]: https://github.com/search?q=repo%3Awithastro%2Fflue+test+pat...

afshinmeh··on Tendril – a self-extending agent that builds and registers its own tools
> It solves the problem the originating user asked it to

Interesting. And is there a mechanism to go back and "fix" the tools after they are published? What happens if the tool decided to use the "id" attribute to click on buttons and now you have a new website that follows a different pattern to find the right target?

I agree that "correctness" of a tool could have different meaning depending on the context of the problem though (e.g. would you consider OOM a correctness bug even if it addresses the user's ask?)

afshinmeh··on Tendril – a self-extending agent that builds and registers its own tools
> how do we avoid burning tokens solving the same problems over again

Letting the LLM write half baked tools is the recipe for burning more tokens.

> There's a wiki the LLM searches before solving a problem, that links saved programs for past actions to their content entry.

What's the criteria for marking an LLM written tool as useful/correct before publishing it?

afshinmeh··on The Prompt API
https://github.com/mozilla/standards-positions/issues/1067
afshinmeh··on An AI agent deleted our production database. The agent's confession is below
It's actually interesting to me that the author is surprised the agent could make an API call and one of those API calls could be deleting the production database.

It's a sad story but at the same time it's clearly showing that people don't know how agents work, they just want to "use it".

afshinmeh··on ClawRun – Deploy and manage AI agents in seconds
I have been struggling with the same issue but help me understand this:

> The lack of predictable output/outcomes

How does that actually show up in practice for you? Asking because "lack of predictable output" could mean different things depending on the context.

afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I'm adding snapshotting as well https://github.com/afshinm/zerobox/pull/21

Then you can run:

```

zerobox --snapshot -- sh -c 'echo "abc" > a'

```

and also `zerobox snapshot list/diff/restore`

afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I agree. What would be the ideal DX from your point of view?
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
Thanks for sharing this. I really like the idea
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I've been testing this on Docker today, including the credential injection, env vars, net calls control. I will add more docs but one interesting use case would be to have something like `zerobox --profile nanoclaw -- nanoclaw`, or something similar.

I'd like to hear your thoughts.

afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I will run the same benchmark test on wasm sandboxes just to be able to compare it with Zerobox. I will share the results tomorrow.
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I added some basic --debug support earlier today, but I will work on proper JSONL/Otel integration soon.
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I think there is still a valid case for sandbox logs/otel. strace would give you the syscalls/traces but not _why_ a particular call was blocked in side the sandbox (e.g. the decision making bit).
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
That's true and the expected behaviour but I see your point. The example there is not great, I should've used `sk_s123...` to show that you are passing the env var to the sandbox as opposed to setting it on the host, then proxying it. I will update it.
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
Here is the video, running Claude with Zerobox, you can see the latency, etc. https://www.youtube.com/watch?v=xzsGsSsx0OI
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I know. I will add more docs soon though, that should make it easier to navigate the code and understand what's going on.
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
That was my thinking, too. The only other option would be reimplement it in Rust (never researched what exists though).
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
I'd love to hear your thoughts! I've been primarily testing this with Bun + Vercel AI SDK for tool call sandboxing.
afshinmeh··on Show HN: Zerobox – Sandbox any command with file, network, credential controls
It does but because I'm inheriting the seatbelt settings from Codex, I'm not resetting it in Zerobox (I thought it's a safer option). Let me look into this, there should be a way to take Codex' profile and safely combine/modify it.
Page 1 of 12Next →