HNHacker News
TopNewBestAskShowJobs

heresalexandria

25 karma · joined June 8, 2026

submissionscomments
heresalexandria··on Pawvis: Control your Mac via camera and train gestures (local FOSS)
I made Pawvis because I wanted to use my computer the way we've been promised for decades that we would in the future. Pawvis reads your hand through the webcam and lets your control your mouse, manage windows, and bind custom actions to any gesture (even user-trained).

Dip a finger to click, fold middle and ring finger in like you're slinging web pages up and down to scroll, dip your pinky in to right click - these are the initial primitives but only the start, and all tracking is performed on-device with code you can read/clone/build yourself from the open source repo.

While I've seen other concepts for hand/gesture tracking through video and voice control none had hit the mark of where I felt that technology could go from a usability and extensibility standpoint.

A good input mechanism needs to be frictionless, it needs to be intuitive, and it needs to be able to adapt to the unique user using it when possible.

Pawvis provides the an ideal blend of utility, technology, and whimsy - and I hope you all enjoy using it as much as I have! Let me know if you have any thoughts or questions, and feel free to clone it (it's open-source MIT licensed) or open any PRs/issues you think might take it to the next level.

heresalexandria··on Ask HN: What are you working on? (August 2026)
I'm working on Pawvis, visual hand-tracking mouse & voice control that's open-source and fully local on your Mac (with optional handoff to Codex or Claude Code for complex computer use tasks).

We live in the future, and I wanted to use my computer the way it's been depicted in movies for decades. It turned out more capable than I'd expected and frankly really fun to use.

Homepage: https://pawvis.app

GitHub: http://github.com/alexandriax/pawvis

Demo: https://www.youtube.com/watch?v=1mdwqP0bwUk

heresalexandria··on Vision based touch-less macOS control via hands (FOSS)
Built a tool to let you control your Mac cursor including click, right click, and scroll with your hands a la Minority Report.

Beta voice commands also involve control via voice, including invoking computer use via Claude Code or Codex CLI.

Fully open-source & free.

heresalexandria··on Seedance 2.5
Been using this and it’s really a phenomenal model, big step up in quality and capability.

This feels like the serious inflection point for high quality full length feature film productions using this tech.

heresalexandria··on A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
That's fair and I agree with this framing.
heresalexandria··on A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
I'm not suggesting that fine-tuned models don't have their place, all I'm saying is that the constant drumbeat of "cheap model X beats more expensive model Z" completely misses that the more expensive model is capable of doing more things at a higher level.

If the appropriate qualifiers were added to say "cheap model X does better at test Y than expensive model Z when we fine tune X to take Y test of existing knowledge" then it would be a more accurate statement, but naturally less impressive.

heresalexandria··on Show HN: Yap – OSS on-device voice dictation for macOS with no model to download
Ahh that makes sense, I was under the impression that OS dictation used exclusively on device models for dictation now but you're right - they do still send this data for transcription.

That definitely changes my take and I'm interested in trying this!

My initial assumption was that this was something I could invoke via CLI or within scripts to transcribe audio from files or user input using on-device models, which is something I would presently probably use Whisper for.

heresalexandria··on A $500 RL fine-tune of a 9B open model beat frontier models on catalog review
This continuous cycle of fine-tuned open models beating frontier on (often vaguely labeled/defined) benchmarks doesn't provide an accurate comparison to the expanding generalized capabilities of the SoTA, which makes them effectively meaningless.

If we were to take these at face value, why is it that the frontier labs' models are making legitimate new discoveries (e.g. Erdős and Jacobian conjectures) and these models are not?

To me, a better signal of capability would be similarly performing novel work at the same or better level, which they presently are not. I say this as someone who very much looks forward to open models being more capable, but to deny the gap is misguided hopeful hype.

heresalexandria··on Show HN: Yap – OSS on-device voice dictation for macOS with no model to download
I don't think get it unless I'm missing something - how is this different from just using the built in Mac OS dictation feature?

At first I thought it was going to be a CLI/package to interface with that API which sounded interesting, but I already use a hotkey to dictate text on my Mac via the OS.

heresalexandria··on NYC Roam: 3D world with real transit and building data
An explorable Manhattan built from real map & transit data. Ride any subway, bus, or bike along its true routes & stops, or take helicopter mode and fly above the city to explore.

Walk up to any building's address plaque for its story Wikipedia info & link to historical photos from Old NYC.

heresalexandria··on Airport Simulator
This is super fun! Could be neat to add keyboard controls for auto routing (i.e. select plane number n and auto route it to strip x or have it fly a go-around to wait). Would also be sick if you added models of real world airports to play.

That said I love this as is and will definitely be playing & sharing it.

heresalexandria··on Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
My impression was that they made this editor with Fable, and its JSON project structure would only serve well for manipulation by lesser models.
heresalexandria··on Show HN: FableCut – A browser video editor AI agents can drive (zero deps)
Cool concept, will try it out! I've had decent results with computer use operating conventional editing tools, but being able to directly edit JSON project files is a solid optimization and opens up a lot of opportunities with things like modular templating.
heresalexandria··on Please stop the AI confidence theater
For sure, I appreciate your comment - this is a tough crowd, but it's their loss.
heresalexandria··on Please stop the AI confidence theater
That's fair, I do agree that you don't need a harness or ultra-high thinking mode for many problems. Many folks evaluate without those things on a task that would benefit from them leading to the sort of attitudes in this article and its comments section, which is where my comment was coming from.

If you're just saying different tools are best suited for different problems, apologies - that's my take as well.

heresalexandria··on Please stop the AI confidence theater
Never said I was bad at math, but I am aware of the fact that computers can do math better and faster than me - and with our powers combined...
heresalexandria··on Please stop the AI confidence theater
This was pretty cool, knocking out a problem that the best minds in maths couldn't for 80 years: https://openai.com/index/model-disproves-discrete-geometry-c...

Also this is a remarkable (and realistic) evaluation of where these systems are for general work which speaks to both the room to grow as well as the pace: https://www.remotelabor.ai/

For some practical examples of what the leading consumer grade AI can do, Ethan Mollick consistently has great writeups with demos: https://www.oneusefulthing.org/p/what-it-feels-like-to-work-...

heresalexandria··on Please stop the AI confidence theater
Offline models are becoming increasingly more capable - merely a few years ago it would've been unthinkable to run the LLM I have on my phone even on my MacBook Pro.

Are you suggesting that losing electricity in the modern age (entirely absent AI) doesn't upend one's world?

You seem to be saying "we should avoid this thing because we'll become dependent on it," but we're highly dependent on all manners of technology for all sorts of things and would seem to be better for it.

heresalexandria··on Please stop the AI confidence theater
It shouldn't be a surprise that the baseline for "best" shifts as better tech comes out, but that doesn't make dated models any less capable than they were when they came out.

Skeptics continue to move the goalposts on what constitutes this mattering, but the fact that frontier systems are making novel maths & sciences discoveries and I can run an LLM on my phone for simple tasks that would've been unthinkable a few years ago are testaments to the directionality of the tech.

heresalexandria··on Please stop the AI confidence theater
That's exactly what the people in my orbit and whom I'm watching are doing, and some of their outputs are fueling the excitement.

If you aren't seeing remarkable things being done with this tech, I'd argue you aren't looking hard enough. I understand there's a lot of noise obscuring the signal, but that's always the case with a "big thing."

heresalexandria··on Please stop the AI confidence theater
Qwen is a lightweight locally hosted model that's many months behind the SoTA available from the big three - while the crowd here (myself included) is excited for locally hosted models to catch up to the usable baseline, regardless of what benchmarks you based that selection on they aren't there yet.
heresalexandria··on Please stop the AI confidence theater
This sounds like you may be using subpar models and/or tools - have you had this experience using Codex with GPT-5.5 on at least "high" reasoning or on Claude Code using Opus 4.8 (both with ability to browse web and sufficient context for your project)?
heresalexandria··on Please stop the AI confidence theater
Totally agree that the lack of a common base of evaluation is terrible for the discussion, and benchmaxing only contributes to this.

The only way to get a sense for these systems is to use them on things you know well, and everyone knows different things at different levels.

People also tend to underestimate how fast this is moving and base their take on dated and subpar systems for a variety of reasons, a key one being that the firehose is too big for any one person to have a proper focus on all of it.

heresalexandria··on Please stop the AI confidence theater
Did you try providing it documentation for the respective formats (via browsing/tool use or input to the prompt)? And were you using a modern thinking model from Anthropic or OpenAI?

The crucial breakdown here sounds like either lack of proper context/harness or insufficiently capable model (there's a huge gulf between GPT-5.5/Opus 4.8/Fable class models and anything not from the big three) or both.

heresalexandria··on Please stop the AI confidence theater
Something a lot of folks struggling with these systems don't get is that the instruction and management of them is often quite important - just because they're capable doesn't mean they're mind readers.

Most of the skepticism I encounter on this front is due to lack of proper direction, process involving planning and review before execution, and appropriate attention given to evaluation and feedback loops.

If you asked the smartest person in the world to YOLO a task with the sort of instruction the average denier uses to evaluate an LLM, you'd likely find they wouldn't get back what they were expecting either - and if you're evaluating on subpar models/tools, you shouldn't be surprised to get subpar results.

heresalexandria··on Please stop the AI confidence theater
They're literally doing novel research. The smartest mathematicians in the world couldn't solve Erdős' planar unit distance problem for 80 years, and OpenAI's models knocked that out a couple months ago.

This stuff is moving fast, and if you aren't evaluating SoTA on at least a quarterly basis, you're going to have a bad time.

heresalexandria··on Please stop the AI confidence theater
The same attitude has been directed at points through history for people "who depend on the internet," "who depend on computers," and "who depend on machines."

I was told growing up "you won't always have a calculator in your pocket" and yet now my phone has an offline LLM on it.

heresalexandria··on Quake in 13 Kilobytes (2021)
This is really remarkable, and frankly more playable than a number of niche ports like this that I've tried - great work!
heresalexandria··on Please stop the AI confidence theater
My observation is that a lot of folks still discounting the capabilities or impact of AI either aren't working with frontier intelligence or aren't using it right.

While the coding horse has been beat within an inch of its life already, I'd recommend throwing Codex on 5.5 high thinking with Computer Use + auto approve at the next thing you're about to spend 5+ minutes on to start to get a feel for how well it handles a broad range of work across literally any surface you interact with today. Use voice mode & mobile app for remote control to seriously watch the friction break down.

Is it always perfect? Maybe not - but for a dramatically increasingly slate of tasks it's becoming a no brainer to offload the busywork and raise the bar on what a single person can do.

It's natural to have hype when you see where this already and where it's going.