HNHacker News
TopNewBestAskShowJobs

frabonacci

513 karma · joined October 10, 2023

Founder Cua AI (YC P25)

https://github.com/trycua/cua

submissionscomments
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thanks, really appreciate it!

The LLM interacts with the VM through a structured virtual computer interface (cua-computer and cua-agent). It’s a high-level abstraction that lets the agent act (e.g., “open Terminal”, “type a command”, “focus an app”) and observe (e.g., current window, file system, OCR of the screen, active processes) in a way that feels a lot more like using a real computer than parsing raw data.

So under the hood, yes, screen+metadata are used (especially with the Omni loop and visual grounding), but what the model sees is a clean interface designed for agentic workflows - closer to how a human would think about using a computer.

If you're curious, the agent loops (OpenAI, Anthropic, Omni, UI-Tars) offer different ways of reasoning and grounding actions, depending on whether you're using cloud or local models.

https://github.com/trycua/cua/tree/main/libs/agent#agent-loo...

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
We’re still figuring things out in public, but a few key differences:

- Open-source from the start. Cua’s built under an MIT license with the goal of making Computer-Use agents easy and accessible to build. Cua's Lume CLI was our first step - we needed fast, reproducible VMs with near-native performance to even make this possible.

- Native macOS support. As far as we know, we’re the only ones offering macOS VMs out of the box, built specifically for Computer-Use workflows. And you can control them with a PyAutoGUI-compatible SDK (cua-computer) - so things like click, type, scroll just work, without needing to deal with any inter-process communication.

- Not just the computer/sandbox, but the agent too. We’re also shipping an Agent SDK (cua-agent) that helps you build and run these workflows without having to stitch everything together yourself. It works out of the box with OpenAI and Anthropic models, UI-Tars, and basically any VLM if you’re using the OmniParser agent loop.

- Not limited to Linux. The hosted version we’re working on won’t be Linux-only - we’re going to support macOS and Windows too.

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Yes - pig.dev is a great product! You should definitely check it out.

Also, let us know on Discord once you’ve tried out c/ua locally on macOS: https://discord.com/invite/mVnXXpdE85

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you - we appreciate it!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you - we appreciate it!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you for your support!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thanks! Great question - those are definitely relevant, but they depend a lot on the deployment model. Since CUAs often run locally or in controlled environments (e.g. a user’s own VM or cluster), we can sidestep a lot of traditional SOC2/HIPAA concerns around centralized data handling. That said, if you're running agents across org boundaries or processing sensitive data via cloud APIs, then yeah - those frameworks absolutely come into play.

We're designing with that in mind: think fine-grained permissioning, auditability, and minimizing surface area. But it’s still early, and a lot of it depends on how teams end up using CUAs in practice.

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Good question! Specifically around computer-use agents (CUAs), I haven't seen much exploration yet - and I think it’s an area worth exploring for vertical products. For example, how do you securely handshake between a CUA agent and an API-based agent without exposing credentials? If everything stays within a local cluster, it's manageable, but once you start scaling out, authn/authz becomes a real headache.

I'm also working on a blog post that touches on this - particularly in the context of giving agents long-term and episodic memory. Should be out next week!

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you - we appreciate your support!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
If you're running Cua from VS Code or Cursor, have you checked out this issue? https://github.com/trycua/cua/issues/61

Feel free to ping me on Discord (I'm francesco there) - happy to hop on a quick call to help debug: https://discord.com/invite/mVnXXpdE85

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thanks for trying out c/ua! We still recommend pairing the Omni loop configuration with a more capable VLM, such as Qwen2.5-VL 32B, or using a cloud LLM provider like Sonnet 3.7 or OpenAI GPT-4.1. While we believe that in the coming months we'll see better-performing quantized models that require less memory for local inference, truth is we're not quite there yet.

Stay tuned - we're also releasing support for UI-Tars-1.5 7B this week! It offers excellent speed and accuracy, and best of all, it doesn't require bounding box detection (Omni) since it's a pixel-native model.

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Sure - just followed you back!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
I love AgentDesk’s take on Kubernetes - it’s something we had considered as well, but it didn’t make much sense for macOS since you can only spin up two macOS VMs at a time due to Apple’s licensing restrictions.

Feel free to join our Discord so we can chat more: https://discord.com/invite/mVnXXpdE85

frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Yes, we’re currently running pilots with select customers for a hosted service of Cua supporting macOS and Windows cloud instances. Feel free to reach out with your use case at founders@trycua.com
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thanks — we really appreciate your support!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Windows host support is on our roadmap - we're currently exploring virtualization options with KVM/QEMU. Please join the discussion on our Discord: https://discord.com/invite/mVnXXpdE85
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you so much - we truly appreciate your support!
frabonacci··on Launch HN: Cua (YC X25) – Open-Source Docker Container for Computer-Use Agents
Thank you - your support means a lot to us!
frabonacci··on Show HN: Lume – OS lightweight CLI for MacOS and Linux VMs on Apple Silicon
Yes, lume relies on Apple's Virtualization framework and can run BSD on a Mac with Apple silicon: https://wiki.freebsd.org/AppleSilicon

I'll definitely document the option in the README, thanks!

On Lima:

- Lima focuses on Linux VMs and doesn't support managing macOS VMs.

- It is more of a container-oriented way of spinning up Linux VMs. We're still debating whether to go in that direction with lume, but the good news is that it would mostly require tweaking the hooks to expose lume’s core to adopt the containerd standard. I’d love to hear your thoughts - would you find it useful to have a Docker-like interface for starting macOS workloads? Similarly to: https://github.com/qemus/qemu-docker

- Still many dependencies on QEMU, which doesn't play well with Apple silicon - while we opted to support only M chips (80-90% of the market cap today) by relying on the latest Apple Virtualization.Framework bits.

On Tart:

- We share some similarities when it comes to tart's command-line interface. We extend it and make it more accessible to different frameworks and languages with our local server option (lume serve). We have also an interface for python today: https://github.com/trycua/pylume

- Going forward, we'd like to focus more on creating tools around developers, creating tools for automation and extending the available images in our ghcr registry. Stay tuned for more updates next week!

- Lastly, lume is licensed under MIT, so you’re free to use it for commercial purposes. Tart, on the other hand, currently uses a Fair Source license.

← PreviousPage 4 of 4