HNHacker News
TopNewBestAskShowJobs

frabonacci

513 karma · joined October 10, 2023

Founder Cua AI (YC P25)

https://github.com/trycua/cua

submissionscomments
frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Thanks! On graphics - currently it's paravirtualized via Apple's Virtualization Framework, so basic 2D acceleration but no GPU passthrough. Fine for desktop use, web browsing, coding, productivity apps. Wouldn't recommend it for anything GPU-intensive though.

Good news is there are hints of GPU passthrough coming (_VZPCIDeviceConfiguration symbol appeared in Tahoe's Virtualization framework), so that might land in a future macOS release. We're keeping an eye on it.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Nice, thanks for sharing! It'd be interesting to integrate MIST into lume's ipsw command - right now Apple's native features in Apple Vz only provides download links for the latest supported version of the host, so grabbing older versions requires workarounds like this.
frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Both use Apple's Virtualization Framework, so core VM performance is similar. Main differences are around agent-first design (HTTP API, MCP server), unattended setup via VNC + OCR, and registry support for VM images.

We've also built a broader ecosystem on top - the Cua computer and agent framework for building computer-use agents: https://cua.ai/docs

We went through the comparison with Tart, Lima etc here: https://github.com/trycua/cua/issues/10

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Yeah, Apple intentionally provides no unattended setup. Plus any process trying to control the UI programmatically needs explicit accessibility permissions, which defeats the purpose.

So we just click through like a human would via VNC. Version-specific but works with their security model rather than against it.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
We haven't observed any networking degradation with Lume on Tahoe so far - things have been working smoothly in our testing. Give it a try and let us know if you run into any issues!
frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Correct, Containerization APIs are Linux VMs specific.

There's a kernel-level check in the Hypervisor framework that enforces the 2 VM limit, and bypassing it violates Apple's EULA.

Nice technical deep-dive on the how here: https://khronokernel.com/macos/2023/08/08/AS-VM.html

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
LoganDark is right. I've personally never tried, and don't think it'd be easy for any macOS predating Apple Virtualization Framework. For that you'd need something like UTM since they're relying on QEMU - these configs might help: https://github.com/adespoton/utmconfigs
frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
re: unattended setup.

You're both right - Apple's official zero-touch setup requires MDM + DEP, which needs Apple Business Manager (and yes, a DUNS number).

But for VMs specifically, DEP doesn't work anyway - VMs don't have real serial numbers that can be enrolled in Device Enrollment Program.

VNC-based setup automation is the only practical option - it's what the ecosystem has converged on for macOS VMs. Lume connects to the VM's VNC server and programmatically tabs, clicks, types through Setup Assistant.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
re: https://github.com/dockur/macos

A closer comparison here is Lumier, which provides a "Docker-like" interface to spin up VMs with a noVNC server: https://cua.ai/docs/lume/guide/advanced/lumier/docker

The key difference: dockur/macos uses QEMU+KVM, which only works on Linux hosts. It can't run on macOS hardware since Apple doesn't expose KVM. See: https://github.com/dockur/macos/issues/256

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Docker on Mac runs Linux containers inside a Linux VM - you can't run macOS in Docker. So if you need Claude / Codex / OpenCode to interact with:

- macOS GUI apps (Xcode, Numbers, Safari, etc.) - macOS desktop automation (screenshots, mouse/keyboard input, accessibility APIs) - macOS CI/CD (building iOS/macOS apps, running XCTest)

...you need an actual macOS VM, which is what Lume provides.

frabonacci··on The Olivetti Company
Growing up in Italy in the 90s, Olivetti was already fading but still everywhere. My grandmother had a Lettera that I swear will outlive us all.

Reading these comments is interesting—for most of you it's nostalgia for nice hardware. In Italy it hits different. We grew up hearing about Olivetti as this national wound. Adriano dies in 1960, Tchou in a car crash a year later, electronics division sold to GE. It gets brought up whenever people complain about "cervelli in fuga" (brain drain)—look, we once had this company that attracted top talent and led the world, and we let it slip away.

I've been living abroad for 10 years now and the irony isn't lost on me. The machines were great. But in Italy what stings is the what-could-have-been.

frabonacci··on CLI's completion should know what options you've typed
+1 to this. I’ve seen the same thing - once completion respects earlier flags, defaults matter less and the CLI becomes self-discoverable. Fish gets part of the way there, but having it modeled at the parser level feels like the missing piece
frabonacci··on Why Windows XP is the ultimate AI benchmark
Yes, in a simulated environment you can do this today using plain JS and connecting to a real VPN, while driving the desktop UI. No infra provisioning needed.

If you need a real Windows OS + corporate VPN, we also support binding agents to actual Windows sandboxes. This example shows automating a Windows app behind a VPN: https://cua.ai/docs/example-usecases/windows-app-behind-vpn

you'll need to define a new task in the cua-bench registry first though - just sign up on the website for early access!

frabonacci··on Why Windows XP is the ultimate AI benchmark
We spent the last few months trying to understand why computer-use agents (Claude Computer-Use, OpenAI CUA, Gemini 2.5 Computer-Use) fail so inconsistently.

The pattern we kept seeing: same agent, same task, different OS theme = notably different results.

Claude Sonnet 4 scores 31.9% on OSWorld and Windows Agent Arena (2 of the most relevant benchmarks for computer-use agents) — but with massive variance. An agent trained on Windows 11 light mode fails on dark mode. Works on macOS Ventura, breaks on Monterey. Works on Win11, collapses on Vista.

The root cause: training data lacks visual diversity. Current benchmarks (OSWorld, Windows Agent Arena) rely on static VM snapshots with fixed configurations. They don't capture the reality of diverse OS themes, window layouts, resolution differences, or desktop clutter.

We built cua-bench — HTML-based simulated environments that render across 10+ OS themes (macOS, Win11, WinXP, Win98, Vista, iOS, Android). Define a task once, generate thousands of visual variations.

This enables: - Oracle trajectory generation via a Playwright-like API (verified ground truth for training) - Trajectory replotting: record 1 demo → re-render across 10 OS themes = 10 training trajectories

The technical report covers our approach to trajectory generation, Android/iOS environments, cross-platform HTML snapshots, and a comparison with existing benchmarks.

We’re currently working with research labs on training data generation and benchmarks, but we’d really value input from the HN community: - What tasks or OS environments should be standardized to actually stress computer-use agents? - Legacy OSes? Weird resolutions? Broken themes? Cluttered desktops? Modal hell?

Curious what people here think are the real failure modes we should be benchmarking.

frabonacci··on AI and the ironies of automation – Part 2
The author's conclusion feels even more relevant today: AI automation doesn’t really remove human difficulty—it just moves it around, often making it harder to notice and more risky. And even after a human steps in, there’s usually a lot of follow-up and adjustment work left to do. Thanks for surfacing these uncomfortable but relevant insights
frabonacci··on iPhone Typos? It's Not Just You – The iOS Keyboard Is Broken [video]
Same here. I even blamed it on switching between Italian and Spanish all the time and thought my brain was short-circuiting. But when you see the right key light up and a different letter shows up, something’s clearly off. Also: with battery saver on it’s basically unusable - the lag makes typing way worse. The video was oddly comforting. Turns out I’m not losing it.
frabonacci··on When hackathon judging is a public benchmark: my report from Hack the North
Yeah, a lot of these corporate hackathons are basically just lead gen in disguise. "Use our SaaS product, maybe we’ll give you a t-shirt." They're more about getting conversions than actually teaching anything useful to the students.
frabonacci··on Ollama Web Search
This is a nice first step - web search makes sense, and it’s easy to imagine other tools being added next: filesystem, browser, maybe even full desktop control. Could turn Ollama into more than just a model runner. Curious if they’ll open up a broader tool API for third-party stuff too
frabonacci··on Docker Hub Is Down
Duplicate https://news.ycombinator.com/item?id=45366942
frabonacci··on Docker Hub is down (again)
Also this comes just a couple of days after a similar incident affected all of Spain
frabonacci··on Abundant Intelligence
> Our vision is simple: we want to create a factory that can produce a gigawatt of new AI infrastructure every week.

The moat will be how efficiently you convert electricity into useful behavior. Whoever industrializes evaluation and feedback loops wins the next decade.

frabonacci··on Microsoft Favors Anthropic over OpenAI for Visual Studio Code
Microsoft has a dozen vertical Copilots to build, so picking the model with the best capability today makes sense. If Claude Code is stronger for dev productivity, using it in VS Code just raises the bar for everything else they ship
frabonacci··on React is winning by default and slowing innovation
Really thoughtful piece. It reminds me of how Angular once dominated by default, until its complexity and inertia gave space for React. The same dynamic could be repeating now - React’s network effects create stability, but also risk suffocating innovation
frabonacci··on Dissecting the Apple M1 GPU, the end
From "draw a triangle" to upstream Vulkan on M1. Practically, this makes the Venus/virtio path viable for guests on Apple Silicon (no passthrough in VZ), which is what many people actually need.
frabonacci··on Claude for Chrome
I thought we had pivoted away from bundling browser-use features in Chromium extensions. Why take a step back instead of bundling their own browser?
frabonacci··on A visual introduction to big O notation
Good intro! I first learned Big O from Cracking the Coding Interview since many universities in Europe notoriously skip complexity notations even in basic programming classes. This definitely explains it in a much simpler way.
frabonacci··on Everything I know about good API design
The reminder to "never break userspace" is gold and often overlooked.. ahem Spotify, Reddit and Twitter come to mind.
frabonacci··on Show HN: Lumier – Run macOS VMs in a Docker
that's also true :)
frabonacci··on Show HN: Lumier – Run macOS VMs in a Docker
I like it - but there seems to be already another YC company with the name: https://lmnr.ai
frabonacci··on Show HN: Lumier – Run macOS VMs in a Docker
Yes, it should be possible: https://developer.apple.com/forums/thread/709453

Lumier does not expose the capability but the underlying Lume CLI does.

lume run <VM_NAME> --recovery-mode true: https://github.com/trycua/cua/tree/main/libs/lume#usage

← PreviousPage 2 of 4Next →