3,022 karma · joined March 14, 2009
Previously, I co-founded SpeakerText (speakertext.com), a interactive/social layer on top of online video, and Humanoid (gethumanoid.com), a crowdsourced virtual labor service that uses machine learning to track worker reputation on Mechanical Turk and ensure quality.
In previous lives, I fought forest fires in Montana & Alaska for the US Forest Service, worked as a freelance news reporter for the NY Times, and drove a 911 ambulance in Harlem and the South Bronx.
https://open-vsx.org/extension/Transcendence/gist-discover https://gist.is/discover
We literally posted our Show HNs within minutes of each other: https://news.ycombinator.com/item?id=48768342
Great minds think alike.
If you run a smaller whisper-distil variant AND you optimize the decoder to run on Apple Neural Engine, you can get latency down to ~300ms without any backend infra.
The issue is that the smaller models tend to suck, which is why the fine-tuning is valuable.
My hypothesis is that you can distill a giant model like Gemini into a tiny distilled whisper model.
but it depends on the machina you are running, which is why local AI is a PITA.
Here’s the trick: use Gemini Pro deep research to create “Advanced Hacker’s Field Guide for X” where X is the problem that you are trying to solve. Ask for all the known issues, common bugs, unintuitive patterns, etc. Get very detailed if you want.
Then feed that to Claude / Codex / Cursor. Basically, create a cheat sheet for your AI agents.
This will unlock a whole new level of capability.
I’m @mattmireles on Twitter — feel free to DM me.
The actual plan was to distill Gemini 2.5 Pro into the best on-device voice dictation model.
Pretty sure it would have worked. Alas.
Also, I had a huge head start, as I spent a month or two working on this in September 2025, shelved it and dusted it back off this weekend.
We've had no WW3 (so far) and no one here needs to worry about being drafted into a war. Gatling might have thought his gun would reduce the number of war fatalities, but but Oppenheimer thought he would end the world. Both were wrong.
Alternative take: Inventors are bad at predicting the downstream societal effects of their inventions.
CoreML is the crappy middleware they made for 3rd party devs, but never got much love and never took off.
Re: ANE — as stated, ANE is crazy fast when it works. Yes, it’s also more power efficient, but the reason I think it’s actually worth building on is being able to make consumer products where the entire experience depends on speed.
I think you can agree that 5 milliseconds to generate a 512 px resolution image is absolutely insane speed.
More recently, I personally tried to convert Kokoro TTS to run on ANE. After performing surgery on the model to run on ANE using CoreML, I ended up with a recurring Xcode crash and reported the bug to Apple (as reported in the post and copied in part below).
What actually worked for me was using MLX-audio, which has been great as there is a whole enthusiastic developer community around the project, in a way that I haven't seen with CoreML. It also seems to be improving rapidly.
In contrast, I have talked to exactly 1 developer who have ever used CoreML since ChatGPT launched, and all that person did was complain about the experience and explain how it inspired them to abandon on-device AI for the cloud.
___ Crash report:
A Core ML model exported as an `mlprogram` with an LSTM layer consistently causes a hard crash (`EXC_BAD_ACCESS` code=2) inside the BNNS framework when `MLModel.prediction()` is called. The crash occurs on M2 Ultra hardware and appears to be a bug in the underlying BNNS kernel for the LSTM or a related operation, as all input tensors have been validated and match the model's expected shape contract. The crash happens regardless of whether the compute unit is set to CPU-only, GPU, or Neural Engine.
*Steps to Reproduce:* 1. Download the attached Core ML models (`kokoro_duration.mlpackage` and `kokoro_synthesizer_3s.mlpackage`) 2. Create a new macOS App project in Xcode. Add the two `.mlpackage` files to the project's "Copy Bundle Resources" build phase. 3. Replace the contents of `ContentView.swift` with the code from `repro.swift`. 4. Build and run the app on an Apple Silicon Mac (tested on M2 Ultra, macOS 15.6.1). 5. Click the "Run Prediction" button in the app.
*Expected Results:* The `MLModel.prediction()` call should complete successfully, returning an `MLFeatureProvider` containing the output waveform. No crash should occur.
*Actual Results:* The application crashes immediately upon calling `model.prediction(from: inputs, options: options)`. The crash is an `EXC_BAD_ACCESS` (code=2) that occurs deep within the Core ML and BNNS frameworks. The backtrace consistently points to `libBNNS.dylib`, indicating a failure in a low-level BNNS kernel during model execution. The crash log is below.
Unlike Apple's formal "developer evangelist" and several others I contacted, the guy actually took the time to talk to us, and I was/am grateful for that. He's a cog in a very large corporate machine. Apple is Apple. He's not the CEO. He was doing his job and did me a favor. I am grateful to him.
Obviously, there was more to the conversation that what I wrote, but these are the actual words that I remember being said.
For more context, the PM at Apple in question was a former colleague of my then girlfriend. I reached out to her to have a friendly catch up. It wasn't positioned as meeting officially with Apple. I was literally just going through LinkedIn trying to figure out who I knew at Apple. So I hit her up on LinkedIn and asked to catch up, then told her about the situation. And this is how she responded. Worth noting: English is not her native language.
https://github.com/ml-explore/mlx-swift
Maybe I am working on a different set of problems than you are. But why would you use CoreML if not to access ANE? There are so many other, better newer options like llama.cpp, MLX-Swift, etc.
What are you seeing here that I am missing? What kind of models do you work with?
I'm telling you what I was told. It's a true story. I was there. It happened to me.
Why would I make up a detail like that?
The only reason to use CoreML these days is to tap into the Neural Engine. When building for CoreML, if one layer of your model isn't compatible with the Neural Engine, it all falls back to the CPU. Ergot, CoreML is the only way to access the ANE, but it's a buggy all-or-nothing gambit.
Have you ever actually shipped a CoreML model or tried to use the ANE?
We were just calling the iPhone's built-in face tracking system via the Vision Framework to animate the avatars. That's the thing that was running on GPU.
Honestly, as a writer, I sometimes spend a bunch of time writing an essay that's the result of a conversation, and I realize –– hours later –– that the conversation actually does a better job of exposing the idea and explaining it than the essay I worked so hard to produce.
I agree that Claude convos can be slop, but don't bash it before you read it!
Why?
Look inside any M-series Mac today. Nearly half the silicon - the Neural Engine cores - sits idle, designed for the pre-transformer era of AI. Meanwhile, my ninety-five-year-old father struggles with a seventy-five-inch screen he can barely see, finding brief moments of magic only when speaking to Siri.
The problem isn't just technical - it's ideological. Apple's privacy stance prevents them from gathering the real-world data needed for truly adaptive interfaces. But there's a solution: AppleCare Platinum, a premium human support service that could evolve into a radical new kind of UI - by serving as an Enders Game for AI computer use training data.