Codec2: A Whole Podcast on a Floppy Disk
auphonic.com
auphonic.com
But looking at this page Codec2 really holds its own when compared to AMBE and especially MELP, two of the most prominent ultra low-bandwidth speech codecs used today: https://www.rowetel.com/?p=5520
In digital amateur radio communication, currently the most widely-used codec is AMBE. But AMBE is a proprietary codec, covered by patents, unhackable - the counter-thesis of amateur radio. Codec2 was born to bring freedom to digital amateur radio communication, and technically even better than AMBE.
Question for IP experts: now that I have heard of the existence of WaveNet and a rough idea of how it works (training a neural network to decode low-bitrate speech data with as much fidelity as possible to the original), would I be prohibited from selling a similar product built with the same technique? How about if I had never heard of WaveNet and went about doing the same thing?
BUT: patents are far more specific than just "a neural network to decode low-bitrate speech data with as much fidelity as possible to the original)". Starting with that goal, you are unlikely to recreate WaveNet's specific structure that is patented.
In fact, WaveNet describes a more general method to efficiently work with sound signals, somewhat comparable to convolutions for images. It's also not impossible to work with sound using alternative MM structures that are not patented, and might actually perform better than WaveNet.
Welcome to software patents.
Reminds me of a story about a copying machine that had a image compression algorithm for scans which changed some numbers on the scanned page to make the compressed image smaller. (Can't remember where I read about that, must have been a couple years ago on HN)
And yes, I think this is a relevant comparison. As the entropy model becomes more sophisticated, errors are more likely to be plausible texts with different meaning, and less likely to be degraded in ways that human processing can intuitively detect and compensate for.
My understanding of this fault was that it was a bug in their implementation of JBIG2, not the actual compression? Linked article seems to support this.
[1]: https://www.xerox.com/assets/pdf/ScanningQAincludingAppendix...
I did check some of the sources, but was not able to find the one I remember which had statistics on it.
The xerox FAQ to it does lead me to consider that I might be confusing this with some other incident though, as they claim that Scanning is the only thing that is affected.
I'd believe him more then any other source.
But the nature of the algorithm means that you have this danger by default. So it's fair to put some blame there.
Assuming 150wpm and an average 2 bytes per word (with lossless compression), we get about 5bps, which makes 2400bps look much less impressive. Add some markup for prosody and it will still be much lower.
This codec also has the great advantage that you can turn off the speech synthesis and just read it, which is much more convenient than listening to a linear sound file.
If you have such a codec, it would be worth testing the word error rate on a long sample of audio. e.g. take a few hours of call centre recordings, pass them through each of {your codec, codec2}, and then have a human transcribe each of:
- the original recording
- the audio output from your proposed codec (which presumably does STT followed by TTS)
- the audio output from CODEC2 at 2048
Based on the current state of open source single-language STT models, I would imagine that CODEC2 would be much closer to the original. And if the input audio contains two or more languages, I cannot imagine the output of your codec will be useful at all.
A few of us have a contact on Sunday mornings here in Eastern Australia and it's amazing how the ear gets used to the sound and it quickly becomes quite listenable and easy to understand.
Are you using Codec2 over radio?
[1] - https://freedv.org/
So a codec that agressively throws away data but still gets good results must somehow enbody sophisticated facts about what human minds really care about. Hence "artifacts that are natural".
In the normal codec2 decoding it sounds like "seventy" but muffled and crunchy.
In the wavenet decoding, the voice sounds clearly higher quality and crisp, but the word sounds more like "suthenty". And not because the audio quality makes it ambiguous but it sounds like it's very deliberately pronouncing "suthenty".
It's as if in trying to enhance and crisp up the sound, it corrected in the wrong direction. It sounds like the compressed data that would otherwise code for a muffled and indistinct "seventy", was interpreted by wavenet but "misheard" in a sense. When wavenet reconstructs the speech, it confidently outputs a much clearer/crisper voice, except it locks onto the wrong speech sounds.
With the standard "muffled/crunchy" decoding, a listener can sort of "hear" this uncertainty. The speech sound is "clearly" indistinct, and we're prompted to do our own correction (in our heads), but also knowing it might be wrong. When the machine learning net does this correction for us, we don't get the additional information of how its guess is uncertain.
This is exactly the sort of artifact I'd expect with this kind of system. As soon as I heard the ridiculously good and crisp audio quality of the wavenet decoder, that fidelity just isn't included in the encoding bits, that's impossible. It's a great accomplishment and just impressive, but it has to "make up" some of those details in a sense very similar to image super resolution algorithms.
I'm just thinking we should perhaps be careful to not get into a situation like the children's "telephone" game, if for some reason the speech gets re/de/re/encoded more than once. Which is of course bad practice, but even if it happens by accident, the wavenet will decode into confident and crisp audio, so it may be hard to notice if you don't expect it.
If audio is encoded and decoded a few times, it's possible that the wavenet will in fact amplify misheard speech sounds into radically different speech sounds, syllables or even words, changing the meaning. Kind of like the "deep dreaming" networks. Sounds like a particularly bad idea for encoding audio books, because small flourishes in wording really can matter.
Edit: I just realised that repeated re/de/re-encoding can in fact happen quite easily if this codec is ever implemented and used in real world phone networks. Many networks use different codecs and re-encoding just has to be done if something is to pass through a particular network.
But the whole thing is ridiculously cool regardless :) And I wonder if they can improve on this problem.
Edit: nm should have read all of your comment before replying!
Can confirm. I spend a lot of time in fringe reception areas, but every now and then I get a good, strong signal and the HD Voice kicks in between my iPhone and my wife's and it sounds like she's standing right next to me. It really is something to experience, especially if the previous phone call was over regular tech.
Back when AT&T was running the "You get what you pay for" ads to combat SPRINT and MCI, it had a service you could sign up for that would give your landline phone calls amazing quality.
Sadly, a majority of people would rather pay less for crap than more for quality; even back then.
Oh well, we will probably never get that kind of quality, which is only possible with QoS on the whole path, if there is any congestion. That is the one thing something like rocket.chat and discord can't provide.
Edit: the way to do this is to force quality upon people, wherever you won't drive them away with the cost this incurs. That way people will associate your brand as a whole with the quality, i.e., in that case, people will associate AT&T with quality, not AT&T premium. Normal people do not even know what kind of plan they are on, except for about one hour before and after they sign the contract.
Except Germany. Which had probably the best telephone system in the world when they deployed ISDN nationwide.
DECT codec is useless if it's transmitted through analog lines, as it gets converted to standard 3.4k-Hz quality anyway. Except for in-house calls of course.
It only seems to work on mobile.
But there were some companies that advertised a lower frequency range as "HD".
1. https://en.wikipedia.org/wiki/Adaptive_Multi-Rate_Wideband
What "value" / "uses" does this bring us?
It cant be used in podcast because as shown it isn't very good with Music. And many podcast has Music in it.
While Codec 2 with WaveNet can have a 2-4x reduction in bitrate. I cant think of a application that benefits from this immediately.
The other thing I keep having in my mind is convolutional neural networks on Codec in general, Music, Movies, etc. What sort of benefits it bring us.
Maybe not too much for "us" with LTE and 128GB storage on our phones, but in cases of low bandwith (think digital police radio), or when you have low storage availability, that's really awesome.
https://virtuallyfun.com/wordpress/2017/04/20/getting-first-...
Of course sound quality is lacking, but it was really cool at the time, but the amount of of time and resources needed was insane for the time.
By pairing the audio with the text, you would almost certainly convince the listener that they can understand it.
Edit: typo
Then the player uses both as inputs to ai (some hand waving), which now has enough to put the pieces together and produce something intelligible again, in the speaker's voice.
They make voice-like synth sounds, different for each character, that are about the length of the text they're saying. It adds prosody and intonation to the text-based dialogue of the game.
Edit: oh sweet there's intonation too. Were these all made manually?
Sine-Wave Speech Demonstration https://youtu.be/EWzt1bI8AZ0?t=74
> Sine-wave speech is an intelligible synthetic acoustic signal composed of three or four time-varying sinusoids. Together, these few sinusoids replicate the estimated frequency and amplitude pattern of the resonance peaks of a natural utterance (Remez et al., 1981). The intelligibility of sine-wave speech, stripped of the acoustic constituents of natural speech, cannot depend on simple recognition of familiar momentary acoustic correlates of phonemes. In consequence, proof of the intelligibility of such signals refutes many descriptions of speech perception that feature canonical acoustic cues to phonemes. The perception of the linguistic properties of sine-wave speech is said to depend instead on sensitivity to acoustic modulation independent of the elements composing the signal and their specific auditory effects.
The post-show of this podcast talks about these and other issues in detail - Marco is on both sides of the issue as a podcast producer and podcast app developer: http://atp.fm/episodes/182
That moment is 15 years in with no signs of losing steam[1]. AAC effectively replaced MP3 for most online audio use cases, with podcasting as a notable exception[2]. And of course, AAC is the audio format for all basically all online video distribution.
[1] Apple kicked off the transition in 2003 with the introduction of AAC-based digital music sales.
[2] Because podcasting is a decentralized medium, and the vast majority of podcasters don't know much (if anything) about media encoding.
Considering also that YouTube uses WebM, which very explicitly is only Vorbis or Opus for audio, "basically all online video distribution" must exclude the web's most popular video distribution site...
That's because AAC is the only format you can count on to work on all devices, and to be hardware decoded on all devices where battery life matters.
Nope! Podcast episodes can be encoded using AAC (which is as ubiquitous as MP3) without issue.
That won't realistically possible with Opus until Opus hardware decoding has available in mobile devices for 5-10 years.
In conjunction with youtube-dl I could listen to pretty much anything I wanted, using almost no data.
These days I use it mostly for audiobooks, if storage is limited.
The N-Gage QD removed the MP3 decoder that was present in the original model. And you could install a software player, and it would struggle with bitrates above 128kbps :D
Modern phones can decode video in software (sucks for battery life, and framerate/resolution are more limited than with hardware, but it's possible). Audio is nothing for them.
I guess it's irrelevant you feel overwhelmed by how long your phone can go on a charge. Plus, low-power/low-CPU requirements are an order of magnitude more critical in devices like smartwatches.
What do you know, it's sort of like Pied Piper without the magical compression or cloud handwaving.
It is impressive how far one can compress speech.
Then I think of the possible negative applications.
a noation of 100m people, talking an hour per day on phone or other audio channel, could be stored on 100m * 365 * 1.5 MB of storage annually: 54 PB.
In raw storage, that's less than $2 million. Far below national actor budgets.
The authors are not from Cornell. I think the author made this mistake because the paper is posted on arXiv, and that’s what’s it says at the top of every page?
Makes it very hard to evaluate claims of codec quality, which seems like the primary purpose of the blog post. :(