Speech Dictation Mode for Emacs
lepisma.xyz
lepisma.xyz
I use it transcribe audio then copy into an LLM to get notes on whatever it is. Helps me decide to watch or listen to something and saves a bunch of time.
Her tweet: https://x.com/JustineTunney/status/1825551821857010143
Instructions from Simon Willison: https://simonwillison.net/2024/Aug/19/whisperfile/
Command line options: https://github.com/Mozilla-Ocho/llamafile/issues/544#issueco...
I am also impressed by the advances in technology. 20 years ago, I had severe RSI problems and worked on "vx-mode", a package for interfacing XEmacs to Dragon NaturallySpeaking, the best speech-recognition solution available at the time. My goals were similar, although the result was nowhere near what the OP has done. Also, speech recognition tech was nowhere near what we have now: I still remember buying good microphones, worrying about microphone placement relative to mouth, endless training and re-training…
This kind of software can make a huge difference for many people.
Plus you can always just enter the command instead of using the key stroke for it. Again, the default UX for that is a bit weak, but with a few packages it becomes pretty strong.
Rejoice! The excellent which-key package that does this comes bundled with Emacs 30! (Emacs 30 will probably be released soon.)
> enter command… default UX is a bit weak
Agreed: the packages Helm, Ivy, and Vertico make this interface much nicer. I use Vertico [1] personally. Though, from Emacs 29, there are some really nice options you can set. I used the following in my Bedrock starter kit [2] to get nicer tab-completion: as soon as you hit TAB twice you'll get bumped into the Completion buffer to select something with your cursor.
Here's the relevant config:
(setopt completion-auto-help 'always) ; Open completion always; `lazy' another option
(setopt completions-max-height 20) ; This is arbitrary
(setopt completions-detailed t)
(setopt completions-format 'one-column)
(setopt completions-group t)
(setopt completion-auto-select 'second-tab) ; Much more eager
;(setopt completion-auto-select t) ; See `C-h v completion-auto-select' for more possible values
There's more configuration options, of course, but this is helpful:[1]: https://github.com/minad/vertico [2]: https://codeberg.org/ashton314/emacs-bedrock
They really are not.
(use-package visual-regexp
:defer t
:bind (("C-c r" . vr/replace)
("C-c q" . vr/query-replace)
("C-r" . vr/isearch-backward)
("C-s" . vr/isearch-forward)))
(use-package visual-regexp-steroids
:defer t)answer:
"I did it. Please note that you're using a Microsoft protocol. Microsoft has a long history of attacking the 4 core freedoms of the Free Software movement which are
The freedom to run the program as you wish, for any purpose (freedom 0). ..."
That said, I don't think this is the way the FSF evaluates software, or that they'd treat an open protocol like this. I could imagine a warning like this about integrating with a proprietary language server in particular, though— and I'd be grateful for it! A locally-run AI assistant that cared about things like that would be super cool.
Also it rewrote all of the legacy Emacs' Elisp into manageable Emacs Guile (with an uberfast JIT and/or libre Guile microcode from the FSF).
I wrote a small follow up trying to write and speak at the same time here https://lepisma.xyz/journal/2024/09/13/can-i-output-two-stre...
https://blog.nawaz.org/posts/2023/Dec/cleaning-up-speech-rec...
I now use Whisper with a much expanded prompt and have the flow integrated both in Emacs and my WM.
Prior HN discussion:
https://news.ycombinator.com/item?id=40174921
I've since done hours of transcription with it - often transcribing whole emails. The challenge is that my brain thinks very differently while talking compared to while typing. As a result, my output is very verbose, and is very different from what I would have typed. I haven't figured out how to speak as if I'm typing.
ELPA installed s/w suite: "I'm sorry Dave, I can't do that"