The best workaround is a third party app that lets you run Parakeet V3 or Whisper on your phone. There are quite a few.
6,632 karma · joined July 18, 2012
Backend-focused full-stack engineer with 8+ years of experience, mostly focused on Go, Rust, and TypeScript, but open to other options.
Resume: https://drive.google.com/file/d/1VNC272B3n7ZEfppMHkm2wGgaINwYl4Av/view
The best workaround is a third party app that lets you run Parakeet V3 or Whisper on your phone. There are quite a few.
If you're going to use OpenRouter to test reasoning levels, always make sure you are locking to the official provider instead of third party providers.
https://facebook.github.io/zstd/index.html
Pretrained dictionaries have never been intended to help with book sized or bigger compression. zstd automatically learns the most efficient dictionary it can within a few kilobytes. Pretrained dictionaries are only useful when you're independently compressing very small records.
In the benchmark, have you considered instructing the models to build their own SPICE simulations to test their work? Simply asking them to write and run simulations could improve performance, even without telling them what to simulate.
I really look forward to a hypothetical LFM3-230M, because LFM2.5-230M is so close to being usable, while FunctionGemma is miles away from being usable.
But, yes, still tangential to TERMy.
The rule has always been intended to cover types that have another word in them but still choose to pointlessly repeat the package name.
`uuid.UUIDGenerator` is a hypothetical example of the anti-pattern that would instead be better named as `uuid.Generator`.
A single six year old RTX 3090 works great: https://www.reddit.com/r/LocalLLaMA/comments/1vkm42m/muse_gl...
I fully expect Meta will release other, smaller Muse models in the near future too.
The 5090 is also supposed to be a $2000 GPU, not a $5000 one. The entire market is utterly distorted right now, which will impact cloud inference more and more over time too. They are not immune to the absurdly high RAM prices, so their prices will have to go up over time too until the RAM supply chain goes back to normal.
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
I have a good amount of background knowledge on all of this, and I've invested hours into pulling this short series together, so this was not just some single shot prompt and hope for the best.
As someone who has used LLMs since ChatGPT first launched, and used them locally since the first open models were released around the Llama 2 era, I have a perspective to share that I think offers some value. LLMs aren't magic. People who have only ever used full reasoning models with agentic tool calling loops may not appreciate how they work under the hood, so this article series tries to give you the foundational concepts and terminology to do more research. Hopefully, it also gives a basic understanding.
Part of the difficulty of putting something like this together is that I don’t know what people don’t know, so there will invariably be gaps that make some of this hard to understand.
I am sure there are ways this series can be improved, so any constructive feedback is welcome.
According to one user preference leaderboard, MiniMax H3 is already ahead of Seedance 2.0 based on thousands of A/B votes: https://artificialanalysis.ai/video/leaderboard/text-to-vide...
I haven't seen any user preference comparisons between Seedance 2.5 and MiniMax H3. As an upper bound, H3 cannot be more than 6 months behind Seedance 2.5 since H3 is already ahead of where Seedance was 6 months ago.
GPT-5.6-Sol is pretty competitively priced, but not all American frontier models are, and even 10% to 30% is still significant for any commodity that's as fungible as frontier models often are.
> as you say that could go down to 20-30% cheaper
I never said anything about 20% to 30%. We don't know how much it actually costs to host this model yet, and that will determine the final price. It could be just a little less, or it could be a lot less.
> once you account for quantisation
There will be no need to account for quantization. Kimi models have been 4-bit only since at least K2.5. They don't release or serve models in higher precision than that. This isn't one of those situations where LLM inference providers are debating between serving 16-bit, 8-bit, or 4-bit, and I have never seen a publicly hosted, paid model that was hosted in less than 4-bit, even if hobbyists will use sub-4-bit quantizations sometimes locally.
The model weights are supposed to release tomorrow.
Over the next several weeks, I would expect competition among open weight providers to drive down the cost, as I've seen happen with other open weight model releases.
On the more cutting edge front, Granite Speech 4.1 has proven to be a reliable workhorse for me, but it is larger than Parakeet. Cohere Transcribe is interesting, but how strong it is seems to vary more from task to task.
Parakeet Unified 0.6B came out a few months ago, combining both online streaming and offline transcription into one model, and that is one that I need to test more, but it seems promising.
As others have mentioned macOS 27/iOS 27 is supposed to have a new model, particularly on devices with 12GB of RAM or more. I have not actually seen the option to enable that new model yet, though, despite being on the beta on a device that meets the requirements. Maybe a benchmark would reveal that it is already active?
I use both vLLM and llama-server. vLLM is very painful, even with the Spark community docker image. It is slow to start, it does not support 3-bit dynamic quants well, and it takes a lot of tweaking to get it to run well for each model I want to try out, which is made worse by the slow starts.
I’m glad you’ve had a better experience? I can only speak to the experiences that I have had repeatedly. For at least a month, people on the official Spark forum were claiming you just couldn’t run MiMo-V2.5 on a single Spark, because they refused to use anything other than vLLM, while I was doing it just fine on llama-server with 200k+ of context.
And llama-server is “worse” in what specific ways? I was specific with my comment. The usual complaint was the lack of MTP/Eagle3 support in llama-server, but that is solved now. Now the main difference is a minor hit to prompt processing speed, at most, if you’re using a single Spark.
Too many people on the Spark forum are closed minded to the idea that vLLM is not the solution to every problem.
llama-server also comes with a truly excellent built-in web chat interface these days, which includes the ability to connect to MCPs so the models can be used agentically through a conversational interface even from my phone. What does vLLM offer? Yeah… nothing. And options like Open WebUI seem really bloated.
For a cluster of multiple Sparks, the pain of vLLM is still worthwhile, as I already said before. Or if you’re running some kind of major production workload, I guess? Instead of a single user, few agent setup like most people.
I would recommend using llama-server if you're just on a single Spark. You get access to dynamic quants like that more easily, the performance is not that different from vLLM most of the time these days, and it is much faster and easier to switch between models.
As far as intelligence goes, Qwen3.6-27B is much smarter than the 35B-A3B model, but that's also not the sort of thing to argue with an AI model about in the first place. Just open a new chat and try again.
Gemma-4-31B is not as good at agentic use cases as Qwen3.6-27B, but it is a fairly balanced model overall, and worth trying out too. Its MTP can nearly triple the performance of the model, where the benefits of MTP or Eagle seem more limited for Qwen3.6-27B in my testing, maybe doubling the speed.
If that's accurate, then you must be doing something wrong/weird. On a single RTX 3090, I'm seeing substantially higher performance. Dual GPU won't necessarily give a ton of performance improvement, but it shouldn't hurt performance.
With llama-bench, I just measured Qwen3.6-27B at 41 tok/s and Qwen3.6-35B-A3B at 153 tok/s on one RTX 3090. (Those results are without MTP. With MTP, I'm seeing about 65 to 70 tok/s for Qwen3.7-27B.)
I'm using the unsloth UD-Q4_K_XL quant. If you're using bf16 for some reason, that could explain the low performance and inability to have enough context despite having 48GB of VRAM, I guess, but... don't do that.
Are you running with MTP enabled? I have seen some people on M5 hardware report 20+ t/s on Qwen3.6-27B using MTP... and I think that was a regular M5, not even M5 Pro.
I think you meant 5.5.
I agree it is probably the same size model. It's probably exactly built on top of 5.5, just with more training, or else they would have bumped the version number to 6.
The person I replied to was acting as if Taalas was ancient history. I was pointing out it has only been a few months.
It has only been four months since they unveiled their first prototype. I don't understand your confusion. Chip development does not happen overnight...?
Their initial blog post laid out a roadmap, so theoretically they should have another thing to demonstrate this summer.
You would think they would support their own GPT-OSS model, but, not really anymore. I wish they would release a GPT-OSS 2, but this doesn't fill me with confidence.
If someone wanted to fork Codex and make a community-maintained version that supports third party models, that would be great, because I liked Codex better than OpenCode for the most part.
Maybe you've found workarounds. Maybe you're using an old version of Codex. Maybe you have your own soft fork. I don't know. But I used to be able to use Codex with self-hosted models, and I gave up on that about a month ago as they kept breaking that.