Whistle: Speech to Text in 16.9 MB
cactuscompute.com
cactuscompute.com
I chuckled at this because my inner voice had an accent as I was reading your comment, due to your writing style.
With a restricted grammar, built in Windows voice recognition, all on device, has managed this exact use case quite well for over a decade. I used it to try and build a clone of the various paid apps that allow you to issue orders to Arma soldiers with voice commands
#JSGF V1.0;
grammar commands;
public <action> = open | close | delete ;
public <object> = file | window | application ;
public <command> = <action> [the] <object>;- my whistle setup (tested can run on echo show with <1s response): 206 correct out of 208
These were not tested on echo show yet - on my pc for now:
- Vosk small, phrase grammar: 184/208
- Speech-to-Phrase 1.4.3 (Kaldi): 118/208
- suggested sphinx: 63/208 and on some examples took 13s on i7 14700k pc
looks like my customized whistle works better for me than these alternatives. But more testing wouldn't hurt.
XML Web Services were awesome, but the SOAP and XML (where the hard parts that should have been given some batteries included defaults) I think was too much of a boat anchor to overcome.
And then the ruby/rails wave came and made 'rest' JSON the Silicon Valley hearthrob, and WSDLs were out like yesterdays trash.
Are they still used in businesses IRL? Yes! I worked for one that has probably (probably) moved off of them by now, but as recently as 2023 there were still some bank-facing power with SOAP RPC .asmx endpoint servers.
I follow hypermedia.systems and htmx/datastar/alpine.js because I'm still trying to get back that powerful _web_ tech advancement rather than squeezing everything into the javascript client side workaround. Typescript is great, a good poor man's F# and leagues ahead of ES3 (which I dabbled in / torture LLMs with so my old retrocomputers can still do web things), but its still bringing along so much baggage that could be lighter weight for older devices AND fits the tech utopianist promise of the original web (which is half of why I play with computer stuff).
By law, every single government form you're able to file is supposed to have an XML Schema available in a centralized registry. This regulation is widely ignored in practice, particularly by local / municipal governments.
This used to be more important in the past, as such a form could automatically / semi-automatically get an entry in the EPUAP form catalog and be made electronically fillable (EPUAP being the now-deprecated centralized government bureaucracy portal basically). As far as I understand, the way that worked was through XSLT and XML Forms. There was some weirdness about each document having a fillable / form view and a preview, I think XML Forms was somehow used to generate the XHTML form, while the XSLT sheet could only generate an XHTML preview of a complete, filled-in version. There were also some custom annotations for auto-fill and such, that was partly done by most documents relying on standardized schemas for entities like "person", "address" or "company".
Since we moved to E-Deliveries and lost a central place for these forms to live in, this is (AFAIK) a bit less important and less common, but internally, things are still XML. If you're filing something like an ID renewal application, even through a newer, more user-friendly frontend, it's still just XML underneath. If you sign something through podpis.gov.pl (the standardized e-signature solution for government paperwork), you can even download that underlying document, both signed and unsigned, and see what that XML is. The pre-signing document preview generated by that site still comes from the XSLT I believe. The signatures themselves are, unsurprisingly, also done via the XML signatures spec.
Incidentally, the European E-Delivery system itself also relies on XML, WSDL and Soap pretty heavily. For those unfamiliar and/or not in Europe, it's basically "email but for the government", with all the guarantees and legal obligations of physical mail, cryptographically-attested confirmation of receipt, proper identity verification and assurance, cross-provider address portability, deployed to a lesser or greater extend in many EU countries and set to replace physical mail.
Much larger though, I think I went full precision and its around 2GB.
> You can handle the truth. We're living a world that has walls and all those walls have to be guarded by men with cannons
This is disturbing. "can handle" instead of "can't handle" may be due to my pronunciation, but "cannons" instead of "guns" is pure hallucination, as well as "all" in "those walls". So maybe don't rely on this for a faithful transcript of what was said.
(For the record, here's the text reference: "You can't handle the truth. We live in a world that has walls, and those walls have to be guarded by men with guns.")
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...
Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)
Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.
The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.
I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.
Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.
"what are your observations on feeling as if sharing that output though?"
and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."
could you rephrase the question
for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
Yes, for doing extremely mild edits (like just removing uhh's etc.) this might be true
but I sometimes feel as if writing can allow me to shift paragraphs, so if I am writing para 1, para 2, I can shift back to para 1 and write another sentence in it and edit some parts of para 1 to include that point
but when I am doing STT, although I can move towards the other para, I find myself just speaking in a complete flow and just write in para 2 only.
Thus when I ask AI to write, I would prefer it to move the statements to appropriate paragraphs and in just general, create a more comprehensible viewpoint from all the STT text that I had written.
This does generate AI generated text which can be detected as such. Uploading it on blogs makes me feel as if people might read what they might consider "AI slop" and so the ethics part (as I myself don't wish to read AI slop)
the problem with AI written or edited texts is that I am unsure of how much effort the other person has put in (just a single prompt or a detailed thought was put in), and I feel as if, others feel the same way.
Should one try to show the rough draft as well to try to show that it was an effort which was human generated or that human effort was used, but that means having a proper disclosure that it was AI-generated/AI-assisted, which I feel as if offputs a lot people (including me) as because of the above logic, that there's still friction for the user within testing if real effort was put into place and I am unsure how effective sharing drafts of it could be.
I don't want my blogs to be tainted and treated as AI-slop because I care about them so I am unsure of what to do. I have multiple things that I have written which if I pass through AI can create some meaningful blog piece but as it stands, they are rough drafts and I find myself putting low efforts or being lazy in actually editing them myself as well (and potentially putting in multiple hours) when AI can be used to help create a more polished version just as well and get across my point.
You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.
You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.
I prompt it to:
"Attached (or underneath) is the transcript of a self recording i've done with tons of rambling and some incorrect words transcriptions, please do a pass clearing out and arranging any typos or possible misunderstandings. Keep original in parenthesis when not sure if it's a misunderstanding. Do not summarize or alter the nature of the content, simply tidy the transcript."
In case it's helpful to anyone else using it, at first it felt a bit slow to me, because there was a noticeable pause after I finished a message before it would quickly type it all out. I changed the input method from direct to clipboard and it's way faster now, almost instantaneous.
I am especially interested in this part. Could you please share the prompt you are using to instruct the LLM to clean up the dictation? Thanks in advance!
And even when it's getting it right, and if there's only one Walmart in Beaverton, it still to this day needs to ask "One option is Walmart on Expressway Road in Beaverton...." Maybe it's correct in its 0% confidence level there, since it's so bad, but... I don't get how you could design something that bad, even before LLMs existed. I feel like I could do better, even using their Speech-to-text engine, with the processing backend built of pure regexes and if/elses.
If you can do something with an extremely limited vocab, voice recognition was fine using off the shelf microchips in the 70s, where you wired in a microphone connection and had discrete pins for output actions.
LLMs are basically only useful for utterly free form transcription, but that doesn't actually help you turn that into tasks to perform and parameters for those tasks
The core "problem" in voice recognition is that freeform speech is an abysmal UX paradigm and provides zero discoverability, and LLMs IMO have not improved the situation of actually doing anything with the resulting text.
The other day I tried to prompt Gemini 3 times to tell me what the heck the business with a weird sign I saw was. The first prompt worked with a stale location context and therefore was way off, the second prompt had to reach out to google servers, and came back with recognizing the physical space I was discussing, but told me that I was talking about an event that takes place in the museum next door that I had told the model was next door to the business in question, the third try it still seemed to understand where I was referencing, but insisted I couldn't possibly be talking about anything there.
It took 1 second on google maps to find exactly what I was referring to, which was the business in Google's system located at the exact map location the model had found.
I'm sick and tired of people turning to LLM and "AI" tools to pretend they are better, when the problem is that these companies don't even use existing good solutions because they just don't care.
"Hey Siri, play [song]"
Leads to, take your pick:
- "You'll need to unlock your iPhone first."
- "I couldn't find [song] on Podcasts" (??????)
- "Playing [a totally different song]"
- "I couldn't find any music by [song, but it thinks it's a band]"
- "Playing music by [song, again it thinks it's a band]"
And don't forget whatever the current phrasing is for "I'm sorry, my shit's all fucked up" and "My network connectivity had a blip and I'm unwilling to retry" and "Even though I have on-device STT models, and now LLMs too, and an on-device database of your music, which is downloaded, I won't bother without the cloud.
but once you calm down and stop hyperventilating from my suggestion, you'll see that the only reasonable and pragmatic course of action is to move to android and linux.
It's also really terrible at recognizing names of my contacts, probably because those names are not represented in the training data.
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
The usecase for small models like this is making on-device STT/TTS more accessable. This is important if your usecase is sensitive to either privacy or latency, but this comes at the cost of quality.
My experience has been that these small TTS models are unexpectedly good if your audio is in distribution (western accents, higher quality audio, common vocabulary), but pretty quickly degrade as you move outside of that. They often dont support more complex features such as diarization, multilingual, or realtime streaming either.
In some cases, this may improve function for a few hours. Best regards =3
It's a very compelling aspect of the problem.
If you can get a model under certain size thresholds, that means you can eliminate latency domains. For example, if the model is able to fit entirely inside L3, the latency of servicing requests drops by an order of magnitude (or better) compared with a model that resides in L3+DRAM.
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
I don't really understand how streaming would work compared against my normal flows. When I dictate, I set a toggle and then do stream of thought as I poke around between windows. When I'm ready to 'flush', I navigate to some target and give it focus for the text to flow.
Does dictation software now keep sort of unfocused floater previews and come with to-clipboard shortcuts or similar?
Corrections based on larger context should also be part of the streaming output -- maybe include replacement text for previous chunk/s identified by chunk ID.
In my case I'm already using one core to run DSP for a beamforming mic array (which works amazingly well for noise cancellation!) so I don't have huge amounts of free processing though.
This definitely seems lighter and faster. How does accuracy compare?
I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.
I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).
If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.
If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with
I like Parakeet because it's good enough, relatively light, and fairly fast. I'm using this for things like meeting transcription and dictation. Since I'm sending most of the text to an LLM to clean up afterwards, it works well enough.
I'm hoping for something the size of Parakeet (or smaller) but better quality. It feels like with all of the advances in making smaller models better in the LLM space, someone should be able to come up with a lightweight and better quality speech to text model.
The real verdict will be after several days of use with dual language dictation, but so far it looks really good!
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.
[1]: https://fr.wikipedia.org/wiki/Langage_siffl%C3%A9_d%27Aas
Highly recommended. I see there is a "new" (2009, I getting old) edition & translation - https://www.amazon.com/First-Circle-Aleksandr-I-Solzhenitsyn...
But Chinese is in another league; the same spoken word may have five meanings, or ten meanings (open zhongwen.com and check out), and you have to build the complete sentence as you parse the sounds, asses its meaning (or several possible meanings, maybe in the context of a few previous sentences), and choose the written word that would match the meaning. You need to carry a lot larger "sense-making" model along with your phonetic, grammatical and syntactic models.
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
So limited selection of languages, not stellar performance?
Kokoro is good, but can't change the pitch easily.
Built my own in a day that blows them out of the water. You can quite literally pick any sufficiently good local model or API provider and combine it with Cerebras for cheap and very fast AI post processing and formatting.
opening this thread for questions/feedback if you have any
Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.
https://github.com/futo-org/android-keyboard/blob/master/LIC...
https://github.com/futo-org/voice-input/blob/master/LICENSE....
...Okay that was pretty good.