HNHacker News
TopNewBestAskShowJobs

BenceRed

50 karma · joined February 26, 2025

bence [at] hoplite.sh
submissionscomments
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Re: the dev box, it works very well for individuals and small sized teams, but starts to become an operational burden past a certain size. Our ideal customer is one who has a ton of engineers and wants great multiplayer/observability, as the case for Hoplite becomes a lot clearer -- "Run through project setup once, then onboard all engineers with one email (and they can bring their entire local setup with one CLI command)".

On the point of microVMs, agreed that it's a very difficult problem to solve. Luckily sandbox providers are continuously improving their APIs to make this slightly easier, but I wouldn't be surprised if we need to migrate over to AWS Lambda MicroVMs and roll a lot of the orchestration logic ourselves. Our goal is to get our P95 project setup time (i.e. connect -> fully running in the sandbox) to around 5 minutes, most of which we imagine being dependency installation. This is one clear point of differentiation where if we nail it, we'd be leagues above the rest of the competition.

The legacy players such as Devin, Cursor, and Factory are certainly well entrenched in their market position, but this space has the unique advantage of completely reworking how it operates every 6 months. These existing tools need to balance keeping up with new user demand for features, while also maintaining the old legacy workflows for their existing customers. We're lucky in that we can now build a product that we believe resembles how the majority of development will work ~1 year from now, meaning we have a lot more flexibility in how we can move forward.

And fundamentally, outside of large enterprise features like on-prem deployment, the key differentiators are 1) UX, 2) cost, and 3) harness performance. We can certainly win at the first, are at parity with the second, and likely struggle at the third (need to do benchmarks/evals -- if those go poorly then we'll transition from the custom harness to using the first party Claude Code/Codex. So quite fixable.)

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
We're going to be investing pretty heavily in evals/benchmarks over the next couple of weeks, so that should give us a much better understanding of how our custom harness stacks up to the official ones.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
We use 'phalanx' internally!
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
In general, we've found that over the past couple of years agent harnesses have gotten much less restrictive, allowing the agent freedom to choose its own way of doing things. It seems like the project agnostic vs. specific harnesses will follow that trend. The pattern will work well for the current gen models, but eventually Fable 7 will be able to intuit how it should approach a specific project very well, at which point the challenge is making sure it has the tools to do so.

And I've made some changes to the coupon, does it work now?

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
We've observed this pattern as well, and have counteracted it by keeping the tasks very finely scoped. For example, "The test(api) check has failed. Fix it, then immediately commit and push." -- or -- "The following comments have been added by reviewers. Resolve each one, then immediately commit and push."

This specificity helps Sol stay on track (most of the time). It doesn't work as well when the comment questions a complex piece of the architecture though.

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
I think for use cases like that, we'd offer on-prem deployments (similar to Factory), potentially coupled with a FDE. Still need to do a lot more research into the enterprise space.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Agreed, we're working with a designer and are going to be fixing this very soon. Our main focus has been on making sure the product itself looks and feels very good to use -- probably not the best approach from a marketing POV.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Agreed that at the moment it's a very difficult problem, but one we're looking to solve! I think it becomes a no-brainer for most people if we're able to give each agent a replica of their production stack.

What does your current setup look like? And are you using an open source solution like OpenInspect for your in-house version, or building it from the ground up?

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Yes, agents have a persistent Chromium session they use via the agent-browser CLI. Typical workflow would involve starting the preview, seeding data, then the agent going through the old and new UX flows for a before + after view. We've also got some optimisations around saving the aforementioned flows in a QA library, so that they can be replayed without needing an agent to run through it all again.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
exe.dev works quite well for giving an agent a computer and managing it remotely, but seems to require a fair bit more configuration to achieve parity with what we offer out of the box. Namely automations, PR autofix, visual QA, and general UI/UX polish.

I think it comes down to whether configuration or ease of use is valued more, and Hoplite favours the latter a bit more. (They shouldn't really be mutually exclusive, but we have a long way to go before we're happy claiming that we match/beat self-hosting in that area)

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
This is still something we're working on making seamless. The current approach is to install our MCP server and ask the local agent to start up a new thread on Hoplite when you want to transition to the cloud, but it doesn't carry over file system changes. (Unless you first push the contents to a remote branch, at which point the Hoplite agent can pull it down.)
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
That setup is pretty much what we're trying to offer with Hoplite!

Using us means losing freedom and control with regards to infrastructure, however we think that's a tradeoff people would want to make in exchange for easier onboarding and a more polished experience.

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Modal just released some features that would allow users to bring custom Docker images, and I'm working on getting your exact use case supported! Aiming to get it out by the end of the week.

Noted the pricing feedback! We're still figuring out exactly what works best so it's still very much so subject to change.

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
This one: https://www.daytona.io. Their platform was OSS for a long time but they decided to go closed source recently.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
You can do either. If you don't want to migrate over fully, I'd recommend setting up an automation to fix Sentry/PostHog issues as they come in. You can get a good feel for the platform and how it fits into your workflows that way.

We also have an MCP server that you can use to delegate tasks (e.g. research, debugging, SRE work) to Hoplite via your existing local setup.

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Modal has a lot of small niceties that made them easy to implement, such as filesystem snapshots and programmatic build images. But I did see that AWS recently launched Lambda MicroVMs, and since we're an AWS house we may transition to using them.

My main issue with Modal is that their autoscaling is not as good as Daytona's. You have to stop the machine, resize, then start it, which takes ~3s and terminates all running processes. Daytona supports scaling up (but not down) without stopping the VM.

Also would recommend checking out ColeMurray/background-agents if you're planning to self host. Very good alternative! And the team behind it are great

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Agreed. Per-thread VMs are quite similar to how local agents use worktrees to avoid cross contamination, but with the added benefit of being able to easily scale up/down compute requirements on demand.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Agreed 100%. We've been working with a designer on a complete redesign of our landing page to avoid that vibey-smell.
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
At the moment we're using Daytona as a redundant fallback in case Modal experiences an outage, but they have very stringent limits on how many resources we can consume concurrently. We're evaluating adding a second provider to help ease this so would love a chat! Feel free to email bence [at] hoplite.sh
BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Thank you! Regarding models, as you said we're not locked into a specific provider, and are able to offer open weight models like Kimi K3 and GLM 5.2

Our pricing is higher than other providers because we do not upcharge on token or sandbox costs. We believe that people should be running as many agents as they possibly can handle, and an upcharge would create a monetary incentive for us to say that, when it's a genuine belief we hold.

We also offer features out of the box that would usually be behind enterprise gating (e.g. sandbox baking).

BenceRed··on Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents
Yeah, luckily they're in quite a different domain to us -- hopefully shouldn't have too much trouble winning the SEO battle