Numen: Voice Control for Handsfree Computing
numenvoice.com
numenvoice.com
Using your computer or programming with it works like a charm, with some interesting and impressive projects based on it coming out as well, like Cursorless[1].
There's a great strangeloop talk[2] demonstrating talon and the actual state of voice coding, which is how I discovered it (hint: it's much better than you'd expect, and straightforward to learn at that).
[1]: https://github.com/cursorless-dev/cursorless
[2]: https://youtu.be/YKuRkGkf5HU
Disclaimer: not affiliated, just a happy occasional user
[0]: https://twitter.com/lunixbochs/status/1574848899897884672
Due to whisper's weakly supervised training on a large amount of automatically scraped data and reliance on a bigger language model, it's far more likely whisper had seen some of the test data before.
It's also worth saying that if you only tried things out briefly, there are a handful of reasons recognition may have seemed worse. Talon uses a strict command system by default, because that improves precision and speed for trained users, but the tradeoff there is it's more confusing for people who haven't learned it yet.
For example, Talon isn't in "dictation mode" by default, so you need to switch to that if you're trying to write email-like text and don't want to prefix your phrases with a command like "say".
The timeout system may also be confusing at first. When you pause, Talon assumes you were done speaking and tries to run whatever you said. You can mitigate this by speaking faster or increasing the timeout.
The default commands (like the alphabet) may also just not be very good for some accents, and that will be the case for any speech engine - you will likely need to change some commands if they're hard to enunciate in your accent.
I recommend joining the slack [1] and asking there if you want more specific feedback. I definitely want to support many accents and even have some users testing Talon with other spoken languages.
Glad to hear Talon is still around! Their slack has grown and they really seem like they have a product now.
I'm curious if anyone has ever tried implementing the latter, or compared the two approaches. I'm sure there would be many obstacles I haven't considered.
Assuming you mean speaking in natural language, that's slower to say, and likely less precise and predictable if you want to be able to just say "anything" any have a result.
You need a command system either way. If you want to express some precise intention, you need to understand what the command system will do.
There is a combined "mixed mode" system I've been testing in the talon beta where you can use both phrases and commands without switching modes.
I wonder if we could replace mouse with eyetracking? I wouldn't expect it to be accurate enough though, give micro movements that eyes do.. and in general erratic movements.. but i'd love to be wrong.
I haven't used eye tracking but I'd imagine that commands could be given in the short time that an on-screen element is focused... and the rest of the time the cursor jumps erratically.
So the problem with eye tracking is what's called the "midas touch" problem. Everything you look at is potentially a target. If you were to simply connect your mouse pointer to your gaze, for example, any sort of hover effect on a web page would be activated simply by glancing at it. [1]
Additionally, our eyes are constantly making small movements call saccades [2]. If you track eye movement perfectly, the target will wobble all over the screen like mad. The ways to alleviate this is by expanding the target visually so that the small movements are contained within a "bubble" or by delaying the target slightly so the movements can be smoothed out which naturally causes inaccuracy and latency. [3] There are efforts to predict the eyes movements to give the user the impression of lower latency, but it's imperfect solution.
Another issue is gaze activation. Computers can't read our minds, so systems which require one to stare fixedly at an object in order to activate an interface are common. The problem with this is the both the delay and effort required. You can easily get a headache from the effort of trying to fixate your eyes on a target. Eye tracking in VR and AR have similar problems.
There are other forms of activation - if you open your iPhone's accessibility menu in the settings, you'll see a bunch of options including head nods, facial gestures, eye blinks and more. [4]
The future of eye tracking is definitely multimodal. A specific gaze target combined with a gesture or hotword is the way humans naturally interact with other humans (you look at a person, get confirmation through eye contact or a nod, and then speak or gesture.) What's amazing is the amount of redundant effort being made in this area. Some of this stuff has been known a decade or more. There are tons of both research papers and thousands of patents to explore which cover the topic in great detail. There is very little that hasn't already been solved.
1. https://uxdesign.cc/the-midas-touch-effect-the-most-unknown-...
2. https://en.m.wikipedia.org/wiki/Saccade
3. https://help.tobii.com/hc/en-us/articles/210245345-How-to-se...
If you encounter what feels like poor recognition in Talon, I recommend enabling Save Recordings and zipping+sharing some examples on the Slack and asking for advice.
The current command set is definitely harder to learn than a system designed for chat/email where "what you say is what you get", but it's much more powerful for tasks like programming once you learn it.
I'm dubious about what kind of general command accuracy Numen is able to get with the Vosk models, as Vosk to my understanding is more designed for natural language than commands.
Point is, I've been bugged with this problem.
" I need a dictation software to read me back what it understood and typed". ALL the software either assume you are looking at the screen and like the win7 (scratch that) I don't want that.
Let me say "I was walking and running besides the train." <pause> "I was walking and besides The train." Would be response so I would say "scratch that." And I would repeat it or ask for help and all.
Why isn't such a system there?
Think of it as a person doing the typing. You write a line, they read back what you said, okay, next. Otherwise fix that like this
All big techs use of voice has so far required Internet access and is creepy. Googles is apawling in that it changes so things that did work, stop working.
What voice needed was for humans to adjust a little to make the computer work easier. e.g. "Computer" "file save" is much more efficient all round than sending off audio to the bork for AI to try work out what it means.
Here's the original screencast on a peertube: https://diode.zone/w/7ZjccgJ5EJCsES3x3yrkpQ
There's also videos of my phone and pi: https://peertube.tv/w/uzMMQ5nbmsHMkDGGVcS1ZB https://peertube.tv/w/miurmjVygd6C71EfPk19QU
Reminded me a bit of those scenes on Blade Runner where Deckard is asking the computer to zoom in a certain area and enhance image :D
It does. Bloody awesome. I'm re-watching this video trying to understand some of the shorthand being used. There's "bang" for exclamation mark; "cap drum" (?) for `cd`. I can't figure out what words he uses to invoke `git clone` at 1:27 but it's incredibly futuristic. I wish my daily driver wasn't a Mac these days =(
Edit: the default 'phrases' are here: https://git.sr.ht/~geb/numen/tree/master/item/phrases
Plus, the offline part could make a good starting point for a DIY personal assistant.
That said, their "getting started" sounds...esoteric.
>There normally isn't any output but you should be able to type "hey" by saying "hoof eve yank" and transcribe a sentence after saying "scribe". You can terminate it by pressing Ctrl+c or saying "troll cap".