Mozilla's Speech-to-Text Engine Is at Risk Following Layoffs
phoronix.com
phoronix.com
They list their 5 areas to focus on .. and Firefox isn't mentioned by name anywhere. For example:
"New focus on product. Mozilla must be a world-class, modern, multi-product internet organization. That means diverse, representative, focused on people outside of our walls, solving problems, building new products, engaging with users and doing the magic of mixing tech with our values" (emphasis mine)
> Mozilla exists so the internet can help the world collectively meet the range of challenges a moment like this presents. Firefox is a part of this.
They follow up on this by saying that they need to go BEYOND the browser but it doesn't mean that they want to drop it.
But well...you never know with that hollow techmarketing talk...
I wasn't trying to suggest there was no mention of Firefox in their roadmap blog - though it is telling that the Firefox mention you pointed to is in context of an argument for developing non-Firefox products. Nor was I arguing that they will drop Firefox. Even if they wanted to, they wouldn't be able to because Firefox is their sole revenue driver and the only thing that gives them a seat at the table for developing open-standards.
To me the telling part was when they listed their five areas of focus, and Firefox was not one of them.
"Not dead yet"
// Change these to proper private fields once Firefox supports themI'm a heavy Firefox user, and have given money to Mozilla in the past, but their leadership has dropped the ball and I'm sorry to say I'm reluctant to give them any more $.
My hope is that they focus efforts on Firefox as an open-web, security and privacy-promoting system (which it does a lot of already). I do think things like Rust have been a distraction, the world doesn't need yet another programming language (compared to Modern C++, Rust doesn't seem worth the effort?).
I disagree there. I think the world needed a systems-level language to replace C++ with safer constructs, more consistent syntax and no historical baggage. It's an open question whether Mozilla was the right sponsor for the development of that language, but nobody else was there (Sorry Dlang =) ).
Which may indicate that the world really doesn't need another programming language.
I'd dispute private fields being "essential". People have been writing a lot of OOP code in JavaScript without them for a long time now.
Just because a company’s mission is positive or aligns with ones values does not mean the company itself is good or is actually carrying that mission out.
I am also always surprised that Mozilla receives so much praise for its technology stack, when their core source code for Firefox and Thunderbird is such a buggy pain to work with.
I understand that they inherited a legacy codebase, but I'm not sure that playing catch-up with Google on (in my opinion) useless features like WebUSB was good technical management. In fact, I would argue that letting technical debt accumulate to the point where you need to rewrite large parts in a completely different language and deprecate technologies still in use (e.g. XUL) is more like an admission of failed technical leadership.
So I am honestly surprised that Mozilla is cutting projects and firing developers, yet nobody seems to take action to reduce the overhead of their apparently incapable management.
Nobody would switch to Firefox for WebUSB, which is why I chose it as the example. There's only crickets on Twitter: https://twitter.com/hashtag/webusb
Instead, people switch to Firefox for privacy, safety and reliability. Mozilla should prioritize those areas highly. But in the past, they haven't. Instead, they burned a lot of money on acquiring 3rd party services that many Firefox users see as nothing more than unnecessary bloat. Like $30 mio for Pocket.
That's why I point the blame at their management, not their code base or developers.
And just FYI, I'm pretty sure I could hire a world-class team and build an excellent Speech-to-Text Engine just from Mozilla's management budget for a single year. Because AFAIK, their (in my opinion incapable) management pocketed $80 mio in 2018.
https://assets.mozilla.net/annualreport/2018/mozilla-fdn-201...
Just because people call things that aren't a meritocracy a meritocracy does not mean that's what it means. It's like saying we don't want to say "democracy" anymore because China (PRC) calls itself a democracy and they actually aren't. Or if we turn the whole thing around, saying that "feminism" is offensive because TERFs use the term too.
Like, are there actually people who actually think like that? How do they communicate??
The meaning of words is constantly perverted
"We aspire to the idea that a human being can demonstrate competence and earn respect and leadership and authority based on what you do and how you do it."
Video was posted on Nov 8 2018. I think everyone can draw their own conclusions.
I remember just yesterday there was a great article on HN that discussed bullshit jobs and counted corporate lawyers towards the "goons". The rationale was that if nobody has them, nobody else needs them.
https://www.businessinsider.com.au/bullshit-jobs-australia-2...
"Hate the game, not the player", might be appropriate here, but even in corporate games players can simply be so ineffective that they should be replaced by someone who can actually play the game.
Every alternative I've been able to find has been either poor quality or too shady for me to ever actually want to use.
You can set it up yourself really easily (got it running locally on Mac in 5 mins), throw an auth'd gateway in front of it to keep it away from bad actors: https://github.com/mozilla/send
have you tried https://catbox.moe/ (or others https://archiveteam.org/index.php?title=Pomf.se/Clones)? Not sure what you mean by "shady", but most of them provide a direct download without any interstitials.
Now that you've made a start you'll probably find you start adding services to that box under the stairs. Why pay for Dropbox when you can use your very own box with far more storage for far less money? Run something like Airsonic and you can stream your own media. Instead of having Xi Jinping watching your goldfish through that IP camera you can put the thing in a dedicated network zone, run a VPN on the box and make the camera only accessible through that VPN. The possibilities are close to endless, the costs are low and coming down all the time, power consumption can be negligible if you choose the right hardware. If in the end you don't like running your own services you can always use something run by someone else so why not give it a try...
Managing backups and data integrity is no joke. Almost everyone does a shit job and don't know it until the data is gone.
All that just for few hundreds of euros; all data are private and nobody is trying to constantly upsell or trying to figure out how to monetize something taken for granted until now.
If you have family or friends you can make a deal with them: you get to hang a backup storage device off their network, they get to do the same on your network. Send encrypted backups and your data should be safe, and so should theirs. Problem solved, everyone happy and nobody gets to mine your data.
--
I just ran a speech-to-text converter on a very clear clip of former Doctor Who actor Tom Baker talking in an interview.
The DeepSpeech converter uses the very latest AI deep-learning advancements to 'listen' to the audio and output the spoken words as text.
After 3 long minutes of running it on a 30-second clip, it printed out its interpretation:
"hooloomooloo how booboorowie i have a honeymoon"
Did you maybe not convert your WAV to the correct sampling rate?
I found that the language model they supplied was trained data that did not contain the words I needed, and got significantly improved results when making my own language model using the kenlm[1] tools.
Besides the hypothesis that DS sucks, the software could also very well be just fine and you made methodological errors.
E.g. some very active projects are:
* Kaldi (https://github.com/kaldi-asr/kaldi/) obviously, probably the most famous one, and most mature one. For standard hybrid NN-HMM models and also all their more recent lattice-free MMI (LF-MMI) models / training procedure. This is also heavily used in industry (not just research).
* ESPnet (https://github.com/espnet/espnet), for all kind of end-to-end models, like CTC, attention-based encoder-decoder (including Transformer), and transducer models.
* Espresso (https://github.com/freewym/espresso).
* Google Lingvo (https://github.com/tensorflow/lingvo). This is the open source release of Googles internal ASR system, and used by Google in production (their internal version of it, which is not too much different).
* NVIDIA OpenSeq2Seq (https://github.com/NVIDIA/OpenSeq2Seq).
* Facebook Fairseq (https://github.com/pytorch/fairseq). Attention-based encoder-decoder models mostly.
* Facebook wav2letter (https://github.com/facebookresearch/wav2letter). ASG model/training.
* Vosk (https://github.com/alphacep/vosk-api). Offline lightweight speech recognition API with support for 10 languages.
And there are much more.
I switched away from Firefox to Brave yesterday. Became tired of how Firefox threads would consume 99% of my CPU endlessly until they're manually killed. This happened about four or five times a day, I'd only notice when performance of everything else tanks. I wish Mozilla would direct more resources towards Firefox than Lockwise and all the other half-baked junk they're spending developer time on.
Quoting their mission statement: "Our mission is to ensure the Internet is a global public resource, open and accessible to all. An Internet that truly puts people first, where individuals can shape their own experience and are empowered, safe and independent." https://www.mozilla.org/en-US/mission/
However Mozilla seems like it did too many wrong bets and didn't execute well (see also FirefoxOS)
That being said, I think the current open source options are actually okay and constantly getting better. They're certainly a lot better today than they were 3 years ago.
They were also doing this. And made an easy to use website, where you could contribute and people did. I cannot imagine it was soo expensive that they now have to throw it under the bus.
I very much hope they don't abandon it, the common voice project is arguably more important than their speech to text engine. There are competing open source speech to text engines. There is no other project like common voice that I am aware of.
It totally is and therefore also a very good crowdsourced project. And they set up already a nice website with badges and other things motivating people to contribute. Daily 15 minutes of a lot of people would mean lots of validated data. Because with user accounts you can also validate userquality etc.
What I cannot confirm is the relation of this to the speech-to-text engine. Reader mode most likely utilizes text to speech.
Now they can bring people back.
Or they can increase executive compensation.
How far in advance do you think these deals are decided before public announcement?