HNHacker News
TopNewBestAskShowJobs

frabonacci

513 karma · joined October 10, 2023

Founder Cua AI (YC P25)

https://github.com/trycua/cua

submissionscomments
frabonacci··on Making a Pokémon fan trailer in 2 hours with Opus 5.5
full conversation: https://x.com/francedot/status/2103909301605761276
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
We've also seen similar improvements on a M5 max. no M1 pro or M3 pro results yet though - would love to see someone try those
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
my guess is apple chose a conservative profile for compatibility across different chips and guest releases
frabonacci··on Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
agreed on the title. added more context below on the exact scope and why this is really a VM capability-reporting issue: https://news.ycombinator.com/item?id=49260087
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
apple silicon is what made us start Lume in the first place last year. the hardware is so good (M1 is now 6 years old!) that people keep pushing through the gaps in the platform. just yesterday we ran a fully offline computer-use agent with Cua Driver and Muse Glimmer, all locally on Apple Silicon, an now the same kind of agent can run isolated inside a macOS VM and use Apple’s GPU path too
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
the better question is why a throwaway account is doing capitalization forensics
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
yeah the naming is confusing. Apple family 9 isnt M9, it's a Metal GPU feature family. Apple maps family 7 to M1, family 8 to M2, family 9 to M3/M4, and family 10 to M5
frabonacci··on Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
RunAnywhere or Conifer?
frabonacci··on Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
> this won't speed up llama.cpp for everyone, just for users running it in this particular kind of Virtualization.framework VM.

correct. these figures apply to llama.cpp inside the macOS guest configuration we tested. Lume is the VM frontend we used, while Apple's Virtualization.framework provides the virtual GPU. bare-metal llama.cpp is unaffected.

> The fix here works around a problem where the VM was causing llama.cpp to select the wrong kernels.

mostly, with one nuance: llama.cpp is selecting the correct kernels for the capability answers it receives. the stock guest reports an older Apple GPU family and a 32 KB threadgroup memory limit, so llama.cpp chooses slower kernels. Our process-scoped layer reports the tested Apple 9 and 64 KB values while allowing llama.cpp to select newer paths that the paravirtual GPU successfully execute

the layer itself though works at the Metal API boundary, independently of llama.cpp. other Metal compute and graphics apps now may select newer paths from the same capability answers, although this is still preliminary and each app needs separate testing. for example, MLX-LM stayed flat in our tests

historically related limitations have been coming up across Apple Silicon VM frontends for a while e.g. Tart tracked MPS/GPU support back in 2023: - https://github.com/openai/tart/issues/501 - https://github.com/openai/tart/issues/1032

UTM also has related cases where apps detect the Apple paravirtual Metal device but falls back to software rendering: https://github.com/utmapp/UTM/issues/7671

frabonacci··on Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
yeah fair point. it's always tricky to get the whole idea across within HN's title limit. tldr: we ran the same workload in the same Lume macOS VM on the same Apple Silicon host, first with stock Metal capability reporting and then with our process-scoped dynamic library. The 11.08x figure is prompt processing, while 16.36x is token generation. the mechanism technically extends to graphics workloads too but these figures are specifically from llama.cpp
frabonacci··on MacBook Neo Deep Dive: Benchmarks, Wafer Economics, and the 8GB Gamble
i bought this for my girlfriend as an entrypoint laptop considering she is coming from Windows - and overall satisfied. the battery though could be improved especially considering for a couple of hundred bucks more we could have gotten a used macBook air
frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
A few examples i'm excited about:

- Closing the coding feedback loop by having agents verify their own changes in a real app

- Automating repetitive workflows across apps that don't have good APIs

- Agents recording product demos of them using software. One compelling use case here: https://x.com/trycua/status/2047383207612645426

- Creating CLI and APIs for apps by reverse implementing their GUI, e.g. see: https://github.com/HKUDS/CLI-Anything

frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
Nothing prevents using it as a general automation library.

If you want to use it directly as an automation framework, you can take a Swift dependency on 'CuaDriverCore': https://cua.ai/docs/cua-driver/guide/getting-started/swift-i...

frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
Thanks! We haven't gone deep on Windows yet because we're still focused on polishing the macOS release. We want to go deeper on the Mac experience before going broader across platforms, and there are still a lot of features we want to ship and use cases we want to share.
frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
really appreciate it. macOS has powerful primitives already, but they weren’t designed as one coherent agent API so you end up stitching together and hitting roadblocks. If Apple doesn't make this more first-class, Linux/Android-style environments may move faster because they’re easier to instrument. I think the OpenAI/Jony Ive AI hardware rumors are yet another signal that people may start building agent-native CUA devices instead of retrofitting agents onto existing desktops
frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
Thanks for trying out Lume! We definitely haven't given up on the idea of sandboxing GUI agents in local macOS VMs. Cua Driver is aimed at a different use case though, letting coding agents and general agents use the Mac you're already on, asynchronously and in the background. That also makes the economics better since multiple agents can share the same machine instead of each needing its own VM
frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
We don't have a specific testing framework yet. cua-driver is closer to an automation interface than a test runner. that said, you could definitely build one on top of it. For reference these are some of our integration tests: https://github.com/trycua/cua/tree/main/libs/cua-driver/Test...

One useful trick is to cua-driver 'launch_app' instead of the default 'open' or other osascript, since it can start the app without raising/focusing it, and the tests don't disturb your active desktop while they run

frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
Thanks for starting that thread, I definitely drew some inspiration from it. But ultimately the secret sauce for the background click came from discovering yabai's window_manager_focus_window_without_raise https://github.com/asmvik/yabai/blob/f17ef88116b0d988b834bb2...
frabonacci··on Show HN: Drive any macOS app in the background without stealing the cursor
Fair criticism. We took a similar approach to established dev tools like Homebrew, with an anonymous, opt-out telemetry to understand install issues, crashes, and high-level usage. For cua-driver specifically, telemetry is limited to command/tool-level events and basic environment metadata. We don’t send screenshots, recordings, app contents, prompts, typed text, file paths, or tool arguments. That said, we should make the opt-out path clearer
frabonacci··on GitHub's fake star economy
what's even more alarming is how exploitable GitHub Trending itself is these days. you can get the star count to fork ratio right and you land on the front page, which then pulls in real organic stars
frabonacci··on Show HN: Holos – QEMU/KVM with a compose-style YAML, GPUs and health checks
it reminds me of https://github.com/dockur/windows with its compose-style YAML over QEMU/KVM. The difference i'm seeing is scope: dockur ships curated OS images (Windows/macOS), while holos looks more like a generic single-host VM runner. Is that a fair read? also curious any plans to support running unattended processes for OS installs?
frabonacci··on Show HN: Cua-Bench – a benchmark for AI agents in GUI environments
Thanks - trajectory export was key for us since most teams want both eval and training data.

On non-determinism: we actually handle this in two ways. For our simulated environments (HTML/JS apps like the Slack/CRM clones), we control the full render state so there's no variance from animations or loading states. For native OS environments, we use explicit state verification before scoring - the reward function waits for expected elements rather than racing against UI timing. Still not perfect, but it filters out most flaky failures.

Windows Arena specifically - we're focusing on common productivity flows (file management, browser tasks, Office workflows) rather than the edge cases you mentioned. UAC prompts and driver dialogs are exactly the hard mode scenarios that break most agents today. We're not claiming to solve those yet, but that's part of why we're open-sourcing this - want to build out more adversarial tasks with the community.

frabonacci··on Show HN: Cua-Bench – a benchmark for AI agents in GUI environments
Fair point - we just open-sourced this so benchmark results are coming. We're already working with labs on evals, focusing on tasks that are more realistic than OSWorld/Windows Agent Arena and curated with actual workers. If you want to run your agent on it we'd love to include your results.
frabonacci··on Show HN: Cua-Bench – a benchmark for AI agents in GUI environments
Hey visarga - I'm the founder of Cua, we might have met at the CUA ICML workshop? The OS-agnostic VNC approach of your benchmark is smart and would make integration easy. We're open to collaborating - want to shoot me an email at f@trycua.com?
frabonacci··on Porting 100k lines from TypeScript to Rust using Claude Code in a month
The author's differential testing (2.3M random battles) is great as final validation, but the real lesson here is that modular testing should happen during the port, not after.

1. Port tests first - they become your contract 2. Run unit tests per module before moving on - catches issues like the "two different move structures" early 3. Integration tests at boundaries before proceeding 4. E2e/differential testing as final validation

When you can't read the target language, your test suite is your only reliable feedback. The debugging time spent on integration issues would've been caught earlier with progressive testing.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Fair point - both use VNC for unattended setup. The difference is implementation: Tart does it via a Packer plugin (Go), we built it natively in Swift with a customizable YAML schema that's less error-prone. User-facing difference is --unattended flag vs Packer workflow.
frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Thanks! On API call visibility - Lume's MCP interface doesn't expose outbound network traffic directly. It's focused on VM lifecycle (create, run, stop) and command execution, not network inspection.

For agent observability, we handle this at the Cua framework level rather than the VM level:

- Agent actions and tool calls are logged via our tracing integration (Laminar, OpenTelemetry) - You can see the full decision trace - what the agent saw, what it decided, what tools it invoked - For the "what HTTP requests actually went out" question, proxying is still the right approach. You could configure the VM's network to route through a transparent proxy, or set up mitmproxy inside the VM. We haven't built that into Lume itself since network inspection feels orthogonal to VM management.

That said, it's an interesting idea - exposing a proxy config option in Lume that automatically routes VM traffic through a capture layer. Would that be useful for your workflow?

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
MDM platforms can skip Setup Assistant, but they require the device to be pre-enrolled in Apple Business Manager before first boot - VMs can't be enrolled in ABM, so those hooks aren't available.

defaults write only works after you have shell access, which means Setup Assistant is already done.

There are tools that modify marker files like .AppleSetupDone via Recovery Mode, but that's mainly for bypassing MDM enrollment on physical Macs - you'd still need to create a valid user account with proper Directory Services entries, keychain, etc.

The VNC + OCR approach is less elegant but works reliably without needing to reverse-engineer macOS internals or rely on undocumented behaviors that might break between versions.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Good catches, thanks! Just updated the page:

Fixed the registry description—you're right, GHCR is an OCI registry. Both tools use OCI-compatible registries, we just default to GHCR/GCS.

Added licensing to the "when to choose" sections.

frabonacci··on Show HN: Lume 0.2 – Build and Run macOS VMs with unattended setup
Thanks for the feedback! You're right that a proper comparison page beats hunting through GitHub issues.

We just put one together (with some help from Claude Code, naturally): https://cua.ai/docs/lume/guide/getting-started/comparison

Page 1 of 4Next →