HNHacker News
TopNewBestAskShowJobs

benswerd

726 karma · joined June 6, 2021

Building Freestyle (YC S24)

github.com/freestyle-sh github.com/worldhealthorganization/app github.com/theswerd

swerdlow[dot]dev

submissionscomments
benswerd··on Brainless: Shadcn components that look like Claude Code, Codex and Grok
Generally, when using these components I ended up wanting to customize a lot. I switched around the options, coloring, the words in the loading, I mix and matched the components from different CLIs, etc.

I think these are more useful as baselines than as final destinations, and I expect production users to customize them far more than options in components.

I also separately don't really believe in traditional components anymore, code is cheap. The value in these components is that I took the time to pixel match a bunch of the CLIs, not the specific interface used to integrate them.

benswerd··on Launch HN: Agnost AI (YC S26) – Extract user feedback from agent conversations
Without using agnost, what are some basic SQL queries I can run on my data to find outliers I'd otherwise be missing?

How far can I get with just keywords, common phrases, boring traditional analysis?

Depending on what I measure there, when is the right time for me to consider upgrading to something like Agnost/what is a specific example of what it will find that traditional/rigid analytics approaches will miss?

benswerd··on Builders Fallacy
Been feeling this more recently, building got 10x easier but choose what to build kinda got harder.
benswerd··on How much do sandboxes cost?
a pricing calculator showing the prices of different sandboxes with dials to compare them in practice.
benswerd··on Show HN: Smol machines – subsecond coldstart, portable virtual machines
Live migrations and the tech powering it was the hardest thing I ever built. Its something that I think will come naturally to projects like smolVM as more of the hypervisors build it in, but its a deeply challenging task to do in userspace.

My team spent 4 months on our implementation of vm memory that let us do it and its still our biggest time suck. We also were able to make assumptions like RDMA that are not available.

All that to say — as someone not working on smolVMs — I am confident smolVMs and most other OSS sandbox implementations will get live migration via hypervisor upgrades in the next 12 months.

Until then there are enterprise-y providers like that have it and great OSS options that already solve this like cloud hypervisor.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Its actually almost O(1) with respect to fork count. We have some O(N) behaviors but I expect to be able to remove those in the next 6 months and get to full horizontal fork O(1) any VM any fork count.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
We read your blogs when building all of this!
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Freestyle isn't designed for an individual engineer working on their Github repos. Its designed for platforms building coding agents that want to take the place of Github all together. Those platforms need some source of truth alongside the VMs, just like how you don't store all of your important documents on your personal computer. That is why we offer git.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
I'm not sure what you saw as slow, I'd love to improve it. Do you mean the dashboard?

We're built as an API for platforms to build on rather than tool for individual developers. Oriented at platform orchestrating tens of thousands at VMs rather than individuals using CLI. We also have a CLI but its primarily a debugging and testing tool.

Resuming a freestyle VM with claude code in it will just work. You can do that via SSH.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
So first I don't, I think startup times are fundamentally really important. 5s is different than 1s is different than 500ms is different than 200ms and users notice.

I don't think people run real world benchmarks on what that coldstart really means though, like time to first response from a NextJS is a very important benchmark for Freestyle and we've spent a lot of time on it. While Daytona sandboxes boot faster than Freestyle ones our first response is an order of magnitude ahead of theirs.

I think another important one is concurrency: In worst case scenarios how many VMs can you get from a provider in a 5 second period is important.

I also think not enough time is spent on "Does it actually work on this VM", stuff like postgres, redis, ntftables, complex linux binaries that are hard to run need to work on these sandboxes because AI is going to need them and I don't think there has really been a feature-bench system yet.

Networking/snapshotting/persistence characteristics all also need to come into this.

I

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
This is a good note. We've never been great at explaining what we're doing and plan to do a lot more work on making it accessible/make sense.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
I don't believe so. while it is technically easy to fork claude code running in these VMs, its not technically difficult to fork a conversation loop outside of the VM as well.

What matters is that its all forked atomically, which can be done with resources outside of the VM as well.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
If git isn't for you we'd still love to support you. We believe to build the sandboxes for coding agents you also need to provide git repos for them so we do that as well. You can easily say give me this vm with these 3 repos and these permissions with us.

But that said, the sandbox stands on its own without it.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
This assumes you can retain the same state after an operation.

> "I wonder if this is slow because we have 100k database rows" > DELETE FROM TABLE; > "Woah its way faster now" > But was is the 100k rows or was it a specific row

Thats a great place where drilling bugs and recreating exact issues can be really problem, and testing the issues themselves can be destructive to the environment leading to the need for snapshots and fork.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
That + its not always simple to replicate state. A QA agent in the future could run for hours to trigger an edge case that if all actions to get there were theoretically taken again it wouldn't happen.

That can happen via race conditions, edge states, external service bugs.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
There is no partial state really possible. We can run out of space on a Node and just say no. But the nature of memory forking is if you don't literally do it 100% right it crashes immediately (I know cuz it took me a while too get it right).
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
TBH I wouldn't recommend using it for this. I'm a big believer in agent chat running outside of the VM, where you can get much better control over the chat loop. I would treat the VM as a tool the agent is using rather than the agent's environment. Like the agent is a human using a machine and watching it, rather than trying to watch it from inside the machine. Then there are great existing observability tools, my fav is langfuse.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
We're actually median under 500ms — ~320ms median — I just didn't want to piss of hacker news with over estimatation.

We have another set of optimizations that we believe can take us to ~200ms in the next few months but beyond that we're pretty much completely stuck.

Realistically other sandboxes will be able to get there before us because we've chosen to support so much of Linux/if you don't run an operating system or don't support custom snapshots that is much easier.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Ah I see. This is very interesting but not what we're focused on right now. I will keep this in mind for future prioritization.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Billed for wall time. whichever plan you are on you get in credits, so hobby plan gets $50 of credits and beyond that billed on per CPU wall time.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
We do not allow long term persistence for the free tier.

This is purely a defense mechanism, I don't want to guarantee storing the data of an entire VM forever for non paying users. We have persistence options for them like Sticky persistence but it doesn't come with the reliability of long term persistence storage.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
CI Builders/QA Agents can do this very well. User session starts, bring VM up with content + dependencies, when session is done throw it away. Keeps it clean, debuggable, fast and cheap.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Freestyle has really built with this in mind. We propose a primary architecture built around declarative configuration of the vm with a git repo as external source of truth.

If the VM crashes/you have another idea/you want to try something else it should be reconstructable from outside of the VM.

However, I think this is potentially unrealistic. While it is the ideal architecture, I hear more and more every day people who just want to have the VMs run for months at a time.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Yessir, we haven't mastered it yet but we've compiled the kernel with enough flags for stuff like nftables and KVM to make it possible.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
It is impossible.

Our tech is not decades old so there is a chance we've missed something but our layer management is atomic so I'd be shocked if you'd be able to corrupt state across forks/snapshots.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Generally kernel level attacks and neighbor performance impacts on the security side.

On the functional side without a kernel per guest you can't allow kernel access for stuff like eBPF, networking, nested virtualization and lots of important features.

Here is a good blog from docker explaining how even the best container is not as safe as a MicroVM https://www.docker.com/blog/containers-are-not-vms/

theoretically you can get to fairly complete security via containers + a gVisor setup but at the expense of a ton of syscall performance and disabling lots of features (which is a 100% valid approach for many usecases).

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
So forking across multiple nodes in that speed is not possible — we run extremely beefy nodes in order to avoid moving VMs across nodes as much as possible.

We are researching systems of hot moving VMs across VMs but it would have very different performance characteristics.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
So thats what we did. We've made forking a whole gas town performant in 100s of milliseconds. Try it — you can definitely see it working on free tier.

In respect to large and powerful RAM + Size is important but I was more-so referring to full Linux power. The ability to run nested virtualization, ebpf, fuse, and the powerful features of a normal Linux machine instead of a container.

benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Great for simple things, but git worktrees don't work when you have to fork processes like postgres/complex apps.
benswerd··on Launch HN: Freestyle – Sandboxes for Coding Agents
Proxmox forking in a few seconds is a miracle!

These are likely only a better value for you at large scale/if you start wanting to run hundreds.

← PreviousPage 2 of 5Next →