Lyra audio codec enables high-quality voice calls at 3 kbps bitrate
cnx-software.com
cnx-software.com
The Lyra version is clearly much louder. This is a serious problem and it borders on being reasonable to call it "cheating".
It's well known in the audio biz that if you ask people to compare two experiences, and one of them is a bit louder than the other, people will say that the louder one was better, or came through more clearly, or whatever it is you're trying to market for. For the purpose of comparing artifacts in two samples, it's absolutely crucial that they be the same volume. You might as well compare two image compression codecs where one of them "enhances" the colors of the original image.
Note: I took the clips for this comparison from the "clean speech" examples at the original source on Googleblog, not the blogspam.
I guess that can be tweaked either way but they're going to tend towards that exactly because it sounds louder and thus clearer.
1. Lossy codecs will use a low-pass filter to get rid of hard to compress higher frequencies. This is often inaudible, but even when it is, it should lower the volume, unless you're applying some kind of compensation for it.
2. It's true that lossy codecs compress different frequencies differently, but that's not usually done in such a way that amounts to applying EQ to the frequencies.
3. Even if the relative balance of frequencies did shift as a result of applying lossy compression, this is still done in a way that the overall loudness of the audio does not change. In this case the Lyra output has changed significantly and in an easily audible way (about +6 dB). You could easily get the same effect in Opus just by amplifying (or applying compression to) the result, but Opus is doing things correctly.
> It's well known in the audio biz that if you ask people to compare two experiences, and one of them is a bit louder than the other, people will say that the louder one was better, or came through more clearly, or whatever it is you're trying to market for.
As a former mastering engineer, you're absolutely right that this is well understood in the audio industry. I used to present my clients with level-matched comparisons of source audio vs. processed so they would understand exactly what was being done, aesthetically.
File 'reference.flac': loudness = -25.7 LUFS, sample peak = -9.2 dBFS
File 'lyra.flac': loudness = -19.9 LUFS, sample peak = -4.1 dBFS
File 'opus.flac': loudness = -25.9 LUFS, sample peak = -9.7 dBFS
So that also matches pretty closely what my ears heard.It's also interesting that the Lyra clip is ever so slightly longer than the other two. The Opus clip has exactly the same number of samples as the reference. Maybe they didn't use a decoder for Lyra at all, just played the file on one system and recorded it using a line-in on another?
Again, assuming I understand correctly, that isn't a "trans coder" that is "here is a seed, generate what this seed creates." kind of thing.
Another way to look at it would be to think about text to speech. That takes words, as characters, and applies a model for how the speech is spoken, and generates audio. You could think of that as really low bit rate audio but the result doesn't "sound" like the person who wrote the text, it sounds like the model. If instead, you did speech to text and captured the person's timbre and allophones as a model, sent the model and then sent the text, you would get what they said, and it would sound like them.
It is a pretty neat trick if they are doing it that way since it seems reasonably obvious for speech that if you could do it this way then the combination of model deltas and phonemes would be a VERY dense encoding.
It would be interesting to see how well it does in other languages.
What I'm getting at is this, do they use the sample as a training data set with a streamlined model generation algorithm so that they can send new initial model parameters as a blob before the rest of the data arrives?
It has my head spinning but the possibilities seem pretty tantalizing here.
If they use any generative model in their codec, they had to train it first, offline, on some dataset. They can't possibly train it equally well on all languages, so we should be able to tell the difference in quality when comparing English to more exotic languages.
> If they use any generative model in their codec, they had to train it first, offline, on some dataset.
One thing I'm wondering if they have a model that can be "retrained" on the fly.
Let's assume for this discussion that you've got a model with 1024 weights in it. You train it on spoken text, all languages, just throw anything at it that is speech. That gets a you a generalized model that isn't specialized for any particular kind of speech and the results will be predictably mixed when you generate random speech from it. But if you take it, and ran a "mini" training system on just the sample of interest, so you have this general model, you digitize the speech, you run it through your trainer, now the generalized model is better at generating exactly this kind of speech agreed? So now you take the weights and generate a set of changes from the previous "generic" set, you bundle those changes in the header of the data you are sending and label them appropriately. Now you send only the data bits from the training set that were needed to activate those parts of the model that are updated. Your data product becomes (<model deltas>, <sound deltas>).
What I'm wondering is this, if every digitization is used to train the model, and you can send the model deltas in a way that the receiver can incorporate those changes in a predictable way to its local model. Can you then send just the essential features of the digitized sound and get it to re-generate by the model on the other end (which has incorporated the model deltas you sent).
Here is an analogy for how I'm thinking about this, and it can be completely wrong, just speculating. If you wanted to "transport" a human with the least number of bits you could simply take their DNA and their mental state and transmit THAT to a cloning facility. No need to digitize every scar, every bit of tissue, instead a model is used to regenerate the person and their 'state' is sent as state of mind.
That is clearly science fiction, but some of the GAN models I've played with have this "feel" where they will produce reliably consistent results from the same seed. Not exact results necessarily, but very consistent.
From that, and this article, I'm wondering if they figured out how to compute the 'seed' + 'initial conditions', given the model that will reproduce what was just digitized. If they have, then its a pretty amazing result.
If you can fit a large generative model (e.g. an rnn or a transformer) in the codec, you might be able to offer something like "prompt engineering" [1], where the weights of the model don't change, but the hidden state vectors are adjusted using the current input. So, using your analogy, weights would be DNA, and the hidden state vectors would be the "mental state". By talking to this person you adjust their mental state to hopefully steer the conversation in the right direction.
Title RMS Peak Diff
clean_p257_011_lyra -20.07 -1.13 18.93
clean_p257_011_opus -26.07 -6.65 19.41
clean_p257_011_refer -25.77 -6.15 19.63
PSD (Welch's method, window=213)It's also more generally useful because I can host other files too, not just images.
[1] https://www.cloudflare.com/distributed-web-gateway/
[2] https://ipfs.io/
(I can imagine a server package that can modify index.html sub-resource URLs depending on current server load, preferring private, locally hosted sub-resources but willing to use 3rd party solutions like Cloudflare, too, if required by a black swan event.)
> the ideal situation in terms of simple, decentralized file hosting solution
Not sure what you mean by "decentralized" if you are in fact hosting it yourself.
> What are the downsides of this approach?
Well, for the casual person it has the obvious downside that you have to have your own VPS. Most people don't have those. Even if you do, IPFS has a couple of advantages: you can host images anonymously, and anyone anywhere in the world can "pin" the image to make sure it stays live. If you're using a server and you forget to pay DO your $5 one month, all your images go poof into the ether.
Imgur these days is slow and riddled by ads. A page show will sometimes load many times the image size in Javascript, stylesheets and images. It also doesn't allow the user to just view the raw image, going as far as redirecting requests to the raw image to a web page if you directly access the URL.
The only downside I see is that the URL is less user friendly without the IPFS toolset installed. Sounds like a pretty good idea to me.
Cloudflare as a gateway is distasteful and this won't last long, but for now at least when you click an ipfs image over cloudflare you get an image and not javascript code.
But its also the only way normal users can see the content.
The Google blog post links to the Lyra paper[1], and Section 5.2 of the paper says:
> To evaluate the absolute quality of the different systems on different SNRs a Mean Opinion Score (MOS) listening test was performed. Except for data collection, we followed the ITU-T P.800 (ACR) recommendation.
You can download those ITU test procedures[2], and skimming through that, it does mention making "the necessary gain adjustments, so as to bring each group of sentences to the standardized active speech level" and a 1000 Hz calibration test tone related to that. (See sections B.1.7 and B.1.8.)
So, if I skimmed correctly, and if the ITU's method of distilling speech loudness into a single number is an effective way to match the volume levels[3], then it seems like they did what they could to avoid cheating at the listening tests.
It is still interesting that Lyra makes things louder, though.
---
[1] https://arxiv.org/pdf/2102.09660.pdf
[2] https://www.itu.int/rec/T-REC-P.800-199608-I
[3] and even for speech that passes through different codecs before its loudness is determined
The part about matching volume levels in the ITU recommendation seems to be talking about making sure the source recordings were balanced. All their clips might well have been exactly at the ITU recommended level of -26 dB, but if Lyra introduced a level mismatch this would have to have been corrected at a later stage, and it's at least possible that it might not have been. The Lyra paper does explicitly say that they didn't follow the ITU rec for "data collection".
Interestingly, the Opus and Reference sources are almost exactly -26 dB relative to full scale (according to several measurements of loudness), but the Lyra clip is about 6 dB hotter. So the source (the reference clip) exactly follows the ITU rec. Did they remember to fix the levels on the Lyra clips? I hope so!
Because as far as I can tell, nobody cares in the slightest about latency.
Phone calls are getting to be like writing postcards to each other. Speak in a whole paragraph. Wait several seconds for the latency to clear. Then the other party responds with a whole paragraph, waits several seconds for the latency to clear...
Improvements to fidelity are nice-to-have, but I would like some real-time in my real-time communications, please.
It was fantastic.
You don't appreciate how much latency is destroying our ability to communicate verbally until you go back to the old way.
One example is arguing. It's no wonder people used to be able to argue with one another on a telephone. You could raise your voice and still hear the other side and adjust your speech in real time. Today it's just one party shouting over the other to drown the opponent out.
I'm rooting for something to replace phone communications. Any chance that Matrix can do better on any of those fronts? Especially on fidelity and latency since they're germane to the high-level subject of this discussion.
Sounds like a run-the-mill everyday business decision.
For many people, end-to-end audio latency in a 1:1 conversation becomes noticeable/annoying at 200ms. And in a multi-participant conversation, talking over each other becomes noticeably more common even at 100ms compared to 50ms.
One other big factor is consistency: if the variability is due to compressor overhead which is constant, the effect will be noticeable but less distracting than if it's varying due to something like wireless conditions.
Musicians have that internal metronome to compare things with.
I would've thought it would be closer to the 200ms range, but I don't have any data to support that.
I know americans said 'I still don't need(it)' but I still do miss it :-)
Funny thing was I had better(and cheaper!) calls to the US using calling cards, dialing into Frankfurt, and from there to the US than using the native offer of my telco.
https://i.imgur.com/vR7NSpG.png
Or are you talking about something else?
A more annoying "feature" of many VoIP systems is that they mute the other person while you're talking. You literally can't interrupt people because they won't hear you.
I presume this is done in order to reduce feedback, but it still sucks.
Most of the problems with voice over mobile networks is caused by frame drops and wacky inter-arrival times. A wired IP network just doesn't have those problems.
Going back to the article/paper, I'd love to hear more about how Lyra interacts with Duo's machine-learned voice interpolation that fills in for dropped frames. Do they complement each other, or interfere?
From the CO to a tandem through the core network your call was digitally switched to the endpoint where it was converted back to analog for the callee's phone.
> This trick enables Lyra to not only run on cloud servers, but also on-device on mid-range phones in real time (with a processing latency of 90ms, which is in line with other traditional speech codecs).
Does that not cover it?
- stream your voice through an encoder: X ms
- send encoded packages over the network: 20-100ms latency (fiber vs mobile phone)
- potential decoding + encoding (if receiver does not support the senders codec, e.g., a landline phone using old codec)
- stream packages through a decoder: Y ms
If you are aiming for 60ms audio latency, which is what I would consider "good", then in the best scenario (20ms network latency; both using same codec) the latency of the encoder+decoder has to be max 40ms (e.g. 20ms for encoder, and 20ms for decoder).
It should be obvious that a decoder that does not meet the 20 ms budget, but takes 90 ms instead which is > 3x the budget, can produce better audio (ideally 3-4x better).
Latency wise, everything below 60 ms is really good, 60ms is good, and the 60-200ms range goes from good to unusable. That is, 200ms, which is what this new codec would hit under ideal conditions, has a latency that humans consider "unusable" because it is too high to be able to have a fluent conversation.
For me, personally, if latency is higher than 120ms, I really don't care about how good a codec "sounds". I use a phone to talk to people, and if we start speaking over each other, cutting each other, etc. because latency is too high, then the function of the phone is gone.
Its like having a super nice car that cannot drive. Sure its nice, but when I want to use a car, I actually want to drive with it somewhere. If it cannot drive anywhere, then it is not very useful to me.
> The overall algorithmic delay is 90 ms
I think the answer for Lyra is that latency is a concern, but maybe at this stage not as much of a concern as it could be. I'm only guessing, though based on this [0]:
> The basic architecture of the Lyra codec is quite simple. Features, or distinctive speech attributes, are extracted from speech every 40ms and are then compressed for transmission.
That sounds like the minimum frame size for Lyra is 40ms. For Opus (the audio codec used for most WebRTC applications), the default frame size is 20ms [1], and most implementations support frame sizes of 10ms [2].
Of course, your favorite web browser might not default to 20ms frames for Opus. And by "most implementations" I meant Google Chrome. :-)
[0] https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
[1] https://tools.ietf.org/html/rfc7587#section-6.1
[2] https://chromium.googlesource.com/external/webrtc/+/HEAD/mod...
That is, with no networking, and no processing, it takes 20ms for any information to from microphone back out to speakers.
https://docs.microsoft.com/en-us/windows-hardware/drivers/au...
The huge gaps between people speaking, and the complete change in conversation flow since you have to speak in huge continuous paragraphs. The constant "go ahead" , "no you go" etc... Ugh.
But I agree with you... nobody seems to care except me.
Well, there's part of your problem right there. No need to even mention Bluetooth or device-related delays.
South Australia to North Britain is the better part of literally half way across the globe. It's either 64000 km at lightspeed (via satellite) or roughly 35000 km via optic fibre cable at 60% lightspeed (e.g. equivalent to ~60000km at lightspeed).
That's 200 ms one way latency just from the distance alone (best case scenario, no less), so 400 ms of latency just from distance alone. Even with something like Starlink we'd still be talking about at least 100 ms latency.
The whole latency from wireless protocols and codecs are just the cherry on top.
What kind of lag, though? Input lag has actually gone up in the past 15 years (e.g. due to displays and USB device polling). These things add up quickly and it doesn't even have to be just the network that introduces lag.
Even 3G AMR, I think that was pre 2000 Speech Codec started at 5Kbps. With a latency of only ~20ms. If I am reading correctly the encode due to ML nature would take at least 40ms and up to 90ms for Lyra.
I am sure there are some specific usage that would be a great use case. But for most consumer consumption I cant think of one on top of my head. One should also be aware of the current roadmap in 5G and the on going work in 6G. We still have a long way to go in maximising bandwidth / transfer per capita. i.e More bandwidth for everybody.
It seems to be the case with ML they want to take these speech codec to new low bitrate. While it is fun doing it as a research, I much rather they push the envelop of at least 6 / 8Kbps if not even higher closer to perfection with even lower latency ( 10ms if not lower ).
On that sample, I felt [2] that the Lyra version exaggerates the pronunciation of the phrase 'with chocolate' in a way that meaningfully differs from the speaker's original. It weakens the voiced 'th' to nothingness, and overshoots both the lead consonant and first vowel of 'choc', and then proceeds to wash the entire rest of the sentence with a peculiar brightened voice that's high, lacks consonant definition, and is close to ringing.
I'm guessing it's actually style transfer, because though the result sounds not much like the speaker's original, the result is reminiscent of the speech pattern and accent that people with East Asian and Southeast Asian ancestry adopt when speaking American English. It was surprising, given that the speaker doesn't sound like that in the original. I wonder if others hear this too.
While Lyra sounds richer and wider-band than Opus or Speex at these bitrates, the degradations and artifacts of those codecs are universally recognized (through years of familiarity with telephones) as compression artifacts and not innate features of the speaker themselves. Therefore listeners can be expected to be sympathetic to the quality issues and not attribute the whole of the sound on the speaker's person.
If AI-trained voice synthesizer codecs become the norm, and it performs well on most speakers, that expectation will go away, and the resulting audio will be attributed wholly to the speaker. That increases the impact of mistakes and misrepresentations introduced by the codec, unbeknowst to the speaker and listener.
[1] https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
So while Spanish, French, or German might get there eventually, don't even try Polish, Czech, or Farsi (Persian) dialects.
I honestly don't hear a 'th' in the original.
> It was surprising, given that the speaker doesn't sound like that in the original.
I disagree. Note that the speaker says "these bread". The three possibilities for those two words—"these bread", "thiiiis bread", and "these breads" with a dropped "s"—would all be weird things for a native english speaker to say for different reasons relating to either wrong pronunciation of "this" or "breads" or the fact that bread is its own collective noun and therefore we typically require separate qualifiers like "these buns" or "these loaves" when separating multiple individual "pieces" (another) into a non-collective. We ask for "some bread" or "a piece of bread", but we don't say "a bread" or "some breads" unless we are discussing categorical types of bread ("ciabatta and rye are breads") rather than instances of such, and only one type of bread is represented in the video.
The Lyra reproduction has a band-pass filtered quality to it, but I find it still remarkably representative of the reference.
The word "looks" sounds completely wrong for me with Lyra. To the point of completely not understanding what this word is supposed to be (first example with your [1] link).
I actually listened to the Lyra version first, and thought the speaker said "when a lad looks for something beyond his reach"
10+ years ago I worked in a small voip shop, where we had very high quality (low jitter), but low bandwidth connection. I researched many codecs of the time (2010-ish).
We liked speex, because it can be used "without strings attached". Also, I can choose the quality depending on the bandwidth. Although for low bandwidth g729 was better. Which we couldn't use because of royalties (but allowed myself to test it).
We chose alaw/ulaw when bandwidth was not a concern, and speex when it was.
Since it does not mention usability outside of google, I also find this comparison unfair or incomplete: if you are comparing a proprietary codec, compare it to g729. If you are comparing a codec to speex, it should be open/free.
Edit: grammar
It is sad that we have to think about licensing and patents of technologies instead of only how good or advanced they are.
If there's a free software implementation, and a company offering the service based in the EU (or shop around and find any other jurisdiction where software patents don't matter), it's often YOLO - but call that "legal arbitrage" if you want to sound fancy :)
These days it's more or less the standard for realtime voice. WebRTC uses it, most of the popular realtime voice applications use it as well, as does Signal.
https://web.archive.org/web/20170202062530/http://www.sipro.... https://en.wikipedia.org/wiki/G.729
[1] - https://www.mumble.info/
[1]: https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
Am I the only one? It is a little odd to me to see the praise here and on the previous discussion.
To be fair, I am convinced I have APD (and, to be fair again, I have never got it checked out).
E: Just realized there is a third example. Perhaps it is not as strong a statement due to Opus' doubled bitrate, but it is still far scratchier. Yet, it is more decipherable than the Lyra codec to me.
I'm almost certain that Lyra has increased the volume on the first sample too. It's quite audible, although I haven't confirmed this with Audacity.
Through good quality headphones, I actually find the Lyra artifacts rather piercing and think I'd pretty quickly get fatigued through having it in my ears over a long conversation. Maybe they would handle this better with a bit of a lowpass filter added.
It's more apparent in the video example from the Google post, I won't spoil it but there's a word that starts with B that sounds very funny in the Lyra version (listen to that one first): https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
Take a look at this too. Also runnable on low power devices. And there was some work of using AI to enhance the codec2 encoded bits too.
http://www.rowetel.com/downloads/codec2/hts2a_1300.wav
Codec 2 does a better job of isolating the parts of sound which are most necessary to intelligible speech, without necessarily caring too much about preserving the original qualities of the speaker's voice or environment.
Fun fact: Codec 2 can be used to transmit voice over IRC:
The screen readers run at 2x normal speed (at least) and to my untrained ears the robotic noises just sound like a garbled mess instead of intelligible speech. The blind person using the phone, however, seems to have no problem understanding it. Fascinates me every time.
I'm not sure exactly why - someone mentioned that it might be because deep submarines have very little bandwidth and need to use them, but I don't have a reference.
Nevertheless, wouldn't be surprised if google didn't want the hassle.
Maybe N years from now, your {Skype,FaceTime,Zoom,Jitsi} call starts by transmitting a pre-trained auto-encoder that can reproduce your speech and visual appears with a "good enough" margin of error from a few kbps worth of data.
Not for audio yet, I think?
https://techcommunity.microsoft.com/t5/microsoft-teams-blog/...
Over time, as these "synthetic accents" gain wide spread adoption, would the main stream pronunciation also adopt such an accent? It's certainly interesting to think about.
It sounds like he said "Someuve" instead of the "Some" clearly audible in the other versions.
They remind of the visual glitches of AI generated images, and I really hope Lyra does not become a common codec!
Having the Lyra sample louder than the others is cheating and gives a false impression.
I was thinking this could be most relevant for something like digital wireless transmissions.
RTP, UDP, IP, and Ethernet overhead are what - 60-ish bytes?
- Not everyone has such good connectivity.
- So you can handle many streams at once, like in a big meeting, without a server mixing them.
- So you can have decent quality video on the same limited connection.
- So you can archive large amounts of speech.
- To advance the state of the art.
- etc.
Why doesn't this article mention LPCNet or Codec2?
This is an effect even with regular telephony these days. A smartphone, carefully held a little way from your ear because you don't trust it to mask touch events when used as a phone, and using 16kbps audio, is not as understandable as an old fashioned hardline phone. Ironically, higher fidelity audio via app (e.g. Whatsapp phone calls) scores better, despite the occasional glitches.
https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
It is unfortunate that for now this appears to be proprietary, closed source and being treated as a google competitive advantage over others, unlike opus which is fully open.
While it is theoretically possible to process audio with something compiled with webassembly, data can't be marshalled into/out of a webassembly worker without the main browser threads help, and that tends to be too janky to use for realtime audio on most platforms.
That pretty much forces googles hand - if they want to use it in their web-based products, it must be opensource and available to all competitors.
They might say this feature only works with their mobile apps though.
"Restricted to a narrow frequency range of 300–3,300 Hz, called the voiceband, which is much less than the human hearing range of 20–20,000 Hz"
Note that this isn't cheating in any way, the source is the source, so it's just a quirk from their conversion process. Probably the tooling around Lyra is pretty rudimentary and the decoder could only output a 32 bit file.