Mimic 3 by Mycroft
mycroft.ai
mycroft.ai
...well, that and a bunch of stuff with phonemes. But I'll do that part :)
For good text-to-speech you need 1 person speaking different phrases but very consistently. Here's an example dataset from Thorsten a German open voice enthusiast: https://openslr.org/95/
... Hi from Darwin :D
There are several ‘layers’ to a voice assistant;
Wake Word - that detects when you are speaking to the device. This is local to the device and we use PocketSphinx.
Speech to text - that detects what you say to determine Intents - we currently use a cloud service for this
Intent matching - this is done locally using our own open source software - Adapt and Padatious
Skills - Intents then match to Skills. Some Skills require internet connectivity.
Text to Speech - We use our own software called Mimic for this, it’s local to the device.
> Mimic 3: Mycroft’s newer, better, privacy-focused neural text-to-speech (TTS) engine. In human terms, that means it can run completely offline and sounds great. To top it all off, it’s open source.
If "skills" are what I think they are (something like external commands, for example "Play X on Spotify"), then my understanding would be that everything but those runs offline and local-only.
But if things like `speech to text` requires internet connection and sends the data to some cloud service, then the entire value proposition of this product falls apart.
I hope that's really not the case, as that would be outright lying and false advertisement.
It's very unclear and misleading to put "it can run completely offline" if you're not 100% sure it can actually run "completely offline", hardware be damned.
There are some options for that available if you're running Mycroft already: https://mycroft-ai.gitbook.io/docs/using-mycroft-ai/customiz...
But these are not what will be shipped on the Mark II...
"*When you use our Services including the Mycroft Voice Assistant, your voice and audio commands are transmitted to our Servers for processing.*"
So it appears that STT is still cloud-based, which is a pity. That's the only thing keeping me from ordering one today.
Note: I'm from Mycroft
I’ve seen you say that Mark II is going to be able to be run non-internet.
Will I be able to be run on my local network, even if I have to dedicate a machine to just speech to text (STT) ?
As will a new privacy policy that better reflects what we actually do.
Before the Mark II ships in September, local STT will be available by default!
> In order to provide an additional layer of privacy for our users, we proxy all STT requests through Mycroft's servers. This prevents Google's service from profiling Mycroft users or connecting voice recordings to their identities. Only the voice recording is sent to Google, no other identifying information is included in the request. Therefore Google's STT service does not know if an individual person is making thousands of requests, or if thousands of people are making a small number of requests each.
Well, unless Google does voiceprint analysis. But Google wouldn't do that, would they? /s
Beyond that, if I'm reading right, local STT will still require a separate STT server. It won't run on the Mark II itself, right?
Daily uses:
1) in the AM i set up all my needed reminders for 5 minutes before every meeting I have
2) it's connected to my hue bridge so I can turn off/on lights by asking while laying in bed, which is wonderful.
3) I play music all day.
4) It reminds me 10 minutes before every sunset to go outside for a walk.
My main software issue is currently how to replicate the music functionality. Playing music at the satellite that requested it, lowering the volume when it recognizes the wakeword. Preselection of "commands" for band and genre names should be easily scriptable afterwards.
In a quiet room, I have no issues with wakeword detection using a playstation eye camera (I wanted the seed USB microhphone array, but between discovering it and starting with buying hardware the supply chain bit once again)
In a simple fashion you can think of it as subtracting the audio being output from the audio coming in from the microphone.
I have not yet managed / worked enough on it (the lack of HW making everything theoretical, which kills my motivation). The way I understand it, is that there’ll either be a casting server on the satelite, or a pulse audio/pipewire server reachable via network. But I have next to no experience with consumer linux, so the configuration of those parts is… hard.
But there are many tutorials for playing multi-room audio (with icecast or something), I just assumed it would be easier without multi-room as I don’t need it, but it turns out it’s not ;)
I've got a home server and a seed array so would be ideal to split that mic (rasp) and processing
Therefore, I'll ask the question I always think of when I see smart speakers: what exactly is their use case? I've never used voice assistants. I've never had a PA. I have a variety of good, dumb speakers. If I am cooking, I have the radio on in the background and a smartphone in my pocket if I desperately wish to change something. I've always thought that the voice recognition was cool, but I've just never quite recognised a position where I would use it!
For the record, I live in a house with at least two raspberry pis on all the time (one as a DTV tuner) so I am far from a luddite in that regard. I just genuinely don't really know what use-case a smart speaker solves. Please enlighten me!
- Making animal / fart sounds for his grand-children.
- Timers for cooking ("Hey Google, set timer to 4 minutes).
I recently spent a month at my parents place. I really, really miss the timer thing.
Additionally, I can see voice assistants as a pretty good interface for Home Assistant. At least for some parts.
End of the day feels like it's the only thing it's good for and it often fails at even doing that.
And to play music. Just asking it to play a track and then asking it to play similar music (very hit and miss), for example, and then asking what's playing, all without having to reach for my phone, finding an app etc.
My experience was that I bought my first one mostly because I wanted something to play music on in the living room anyway and didn't really care about getting a full on stereo setup as I'm not very picky about the sound quality, but I was curious. I never use voice assistants on my phone. But I found myself using it more and more as I got used to being able to turn things on/off without reaching for anything or when my hands where otherwise full.
It's not something I'd have the slightest difficulty of living without, but it feels like it's decreasing friction for a lot of small things.
I now have four - one by my desk, one in the living room, one in my sons bedroom and one in mine.
It’s definitely something I could live without, but even Apple’s speaker is pretty cheap. Especially so if it’s your main speaker (if you care about audio quality it might be a problem, but I don’t so it’s not).
What's the BOM look like? I'd love to understand more about the design. The software's open source, right? After a brief skim I didn't see a repo link. Does anyone know where the source is? Do they use an AI accelerator DSP/TPU or just plain-old-software-on-a-CPU?
Regarding the BOM, I assume you mean the Mark II? That you can find here: https://github.com/MycroftAI/hardware-mycroft-mark-II/tree/m... We actually ended up designing our own RPi daughterboard called the SJ201. It's mostly an audio front end with an XMOS XVF-3510 and dual mics, but also includes a 23W amp, some LEDs for feedback, buttons, a hardware mic switch, GPIO breakout and power management (amongst other things).
I'm going to assume this isn't our developer Mike's alternate account congratulating himself lol
But all I see is documentation, discussion of what they used to build it, and.... where's the actual softare?!?
and the source code is available at: https://github.com/MycroftAI/mimic3
There is currently a Docker image, DEB package, and PyPI release.
I would however point out that the Mark II is completely hackable. So whilst it's not an SJ201 on its own, you can absolutely pull the whole thing apart, use it in other enclosures and even put it all back together again.
Another one of the reasons we made it is because having a single daughterboard greatly simplified production and made the Mark II more robust overall. There's no longer the possibility of a loose wire to the power supply or amp after it gets kicked around in the back of a delivery van. They're all in one and connected via the 40 pin GPIO header, but absolutely removable from the Mark II unit itself.
I've previously used rhasspy (by the author of this feature apparently, he got a job at mycroft and stopped development of rhasspy) with some custom scripts as glue to home assistant as my local-only smart home solution.
Can I do this with mycroft? Does it maybe come with a Home Assistant integration? My main worry is that mycroft seems to be doing a bit too much for my taste. Is there functionality overlap/conflict between HA and Mycroft?
docker run -it -p 59125:59125 -v "${HOME}/.local/share/mycroft/mimic3:/home/mimic3/.local/share/mycroft/mimic3" 'mycroftai/mimic3'
UI loads but when I click on the speak button I get this error: PermissionError: [Errno 13] Permission denied: '/home/mimic3/.local/share/mycroft/mimic3/voices'
fixed it for me
'chmod a+rwX'
The capital X is meant to only add the execute bit to directories.
Note: I am am the main developer of Home Intent.
edit: Ah I see they are in speakers.txt in the git repo for the voice.
The other option is to jump on https://mycroft.ai/mimic-3/ and use the "Hear my voices" tool. There are three drop-downs: - language - model - speaker
Full license text here: https://github.com/MycroftAI/mimic3-voices/blob/master/LICEN...
and an explainer on what the legal mumbo jumbo means here: https://creativecommons.org/licenses/by-sa/4.0/
On that topic, you might be interested in this tool I've worked on which works with an earlier TTS project the primary Mimic3 developer created:
https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to...
To be fair, they've communicated fairly well recently. There were periods when things went too quiet, but in the last year or so we've been getting monthly updates that have real information buried in the usual hype and boosterism.
I think the shipping schedule they're working towards now (September) is unlikely to actually hit the promised date: they've used up most if not all the slack already, but it seems like obstacles are falling regularly at least. The latest update (this week) mentioned that they need to redo their FCC certification tests, for instance, but they've found and fixed the problem already.
I'm somewhat confident of a late-2022/early-2023 delivery (vs the original campaign's promised Dec 2018).
https://github.com/MycroftAI/mimic3-voices/blob/master/sampl...
If you know of any good Portuguese voice datasets we'd love to train one. At some point we'll start getting new professionally recorded data for each language too.
But if it’s AGPL, I’ll move on. They’re too idealistic to deal with megacorp. They don’t sell through intermediaries, can’t/won’t be insured enough, etc.
Inevitably the small company sees $$$ and goes nuts trying to sell even when warned not to. Then the risk management and contract stuff kicks in, drags out, and kills it.
The small company blew a bunch of cash on nothing. Or they base sales projections on it. I’ve seen a couple go under.
Those are people’s livelihoods and I can’t do that to them.
That’s not limited to AGPL, but the license is a signal of that type of business. And they never do the smart thing with megacorp and sell through an established intermediary.
Takes a lot of good software off the table.
Most companies are fine with BSD/MIT software with absolutely no support. Your argument holds absolutely no water compared to what they actually want which is to take other people's work and close it up for profit.
If they want a license to close it up, they can pay for it.
Guess both sides have strong feeling why something is "bad".
I personally only contribute to agpl software, try to use agpl/sspl and similar licensed software.
Guess the " corporate mentality" and "Foss mentality" don't meet unless it's a corp built around Foss products and ideology.
Look at valve and their deck compared to Nintendo switch.
Darn
Use others to gain an advantage, then pull up the ladder after you.
I'm not surprised that large corporations want to have their cake, eat it and not pay for it but it's not a compelling argument for people producing opensource software