Here's another project that builds on top of VOSK to provide a tighter integration with Linux: https://github.com/ideasman42/nerd-dictation
But really, the situation is pretty good, with a lot of code and dataset available as opensource. Notably, if you're not constrained to smartphones and the like, you can run on your computer quite a number of modern models, see for instance https://github.com/coqui-ai/TTS/ (which itself contains many different models).
The work that needs to be done is """just""" to turn those models into something suitable for smartphones (which will most likely include re-training), and to plug them back into Android's TTS API.
* Larynx: https://github.com/rhasspy/larynx/
* OpenTTS: https://github.com/synesthesiam/opentts
* Likely Mimic3 in the near future: https://mycroft.ai/blog/mimic-3-preview/
Larynx in particular has a focus on "faster than real-time" while OpenTTS is an attempt to package & provide common REST API to all Free/Open Source Text To Speech systems so the FLOSS ecosystem can build on previous work supported by short-lived business interests, rather than start from scratch every time.
AIUI the developer of the first two projects now works for Mycroft AI & is involved in the development of Mimic3 which seems very promising given how much of an impact on quality his solo work has had in just the past couple of years or so.
https://www.eff.org/deeplinks/2005/09/authors-guild-sues-goo...
https://www.cnn.com/2011/09/13/tech/web/authors-guild-book-d...
https://popula.com/2022/01/22/what-kind-of-writer-accuses-li...
The difference between you reading and the software is one is 'mechanical', whether you want to constrain someone doing that in the rights you grant is debatable.
Copyright enables you (gives the creator the 'freedom' to choose) to make such choices, its what people choose to limit is the problem, as they tend to be very stingy.
The "outcome" (text is read aloud) is the same if you read it aloud to yourself. Really the difference is that when a publisher releases an audiobook they hire someone (sometimes multiple someones) often an actor or the original author to sit in a recording studio and recite. They pay for things like studio time, sound engineering, editing, the narrators time, sometimes music or foley, etc. It's very much a different product.
If you have a book and you recite it, or if you pay someone to come into your home and read it to you, or you get a bunch of software and have your computer read it for you, that's your right and at your own expense in terms of time, money, and effort. Some kind of text to speech software is expected on pretty much every device. Including such features in devices or using those features (especially the accessibility features) of your own devices isn't copyright infringement, shouldn't open you up to demands for payments from publishers, and is in no way comparable to a professionally produced audiobook. Maybe one day the tech will advance to where a program can gather the context needed to speak with and convey the correct emotion for each line and will be capable of delivering a solid performance, but right now we're lucky if more than 2/3 of the words are even pronounced correctly and the inflection isn't bizarre enough to make you question what was being said or distract you from the material.
As a follow up it would be cool, TTS did use different voices for Narrator and characters in a book ... if someone patents that your welcome!
Honestly, if AI ever gets good enough at crafting films from literary source material that the AI movies has any chance of competing with a hollywood production the entertainment industry is screwed. I'm positive that by then whatever crazy stuff that AI is putting out will be everywhere and playing with it would be way more fun than a movie theater ticket.
Different voices would be cool, but who is speaking which line can sometimes be ambiguous even for human readers. I'd be happy with just one voice that didn't sound like a robot or like a human voice spliced together from multiple sources.
Yeah ... so you might be getting into performance rights then ;P
I still expect it'll mean a lot of very amusing outbound voice mail messages.
We recently listened to How To Train Your Dragon instead of reading it to our kids ourselves specifically because it was narrated by David Tennant.
One of the vast litany of non-AI related Siri flaws that make you wonder how a $2t company can achieve such tragic levels of incompetence.
That said, having only skimmed the paper I didn’t notice a discussion of the compute requirements for usage (just training), but it did say it was a 28.7 million parameter model, so I recon this could be used in real-time on a phone.
[0] judging by the videos of Dr. Sbaitso on YouTube, it was only one step up from the intro to Impossible Mission on the Commodore 64.
Microsoft and particularly Apple used to make cutting edge (at the time) TTS available as a marketing gimmick, but then did not develop this further because TTS was only relevant for screen readers and people with impaired vision weren't really their primary customers. Then TTS made advances and companies started selling high quality voices at high price points. Now companies want to make as much money with it as they can. Improving "free" TTS for ordinary customers is not really a top priority of Microsoft and Apple. Moreover, since network speeds have increased, any improved end-consumer TTS will send the text to a server and the audio back, so that a company can collect and analyze all the texts and make money with this spy data. That's how Google's free TTS server works.
That's exactly what I meant when I wrote "...strive to keep the status quo of being way beyond the capabilities of the people".
Monopolizing the means of maximizing the profits to maximize the power - now that's some deep conspiracy territory, indeed.
Honestly, are those really different? How is it not in their interest for money?