YouTube audio quality – How good does it get? (2022)
audiomisc.co.uk
audiomisc.co.uk
According to their documentation, this is "Upper bound of 256kbps AAC & OPUS": https://support.google.com/youtubemusic/answer/9076559?hl=en...
EDIT: FWIW, I'm a Premium member. I'm not sure if this is a standard feature.
Even if receiving the same sources, the encoding process is their own and aims to have zero clipping.
Due to what I assume are music licensing oddities, one song in my Spotify now has an entirely different singer. It was taken down for a while, then later reuploaded as a new recording with new vocals. I’ve also seen copies of my songs changing to remix/cover versions, presumably due to metadata adjustments by the artists. Although usually I can re-search and find the original, adding it back to my library.
A few songs just have new outros/intros. One song in my library now has an additional several seconds of silence on the end. One has a new longer intro that I think is really from the music video version of the song.
As convenient as these music services are, I hate that my library changes beneath me without my control. There are dozens of songs just straight up missing from my Spotify library now. And a small handful that have changed audio. These are almost always in indie songs, and Spotify just hides deleted songs by default in the UI so most users don’t notice.
-F (uppercase!) lists all available audio and video formats to download.
"I decided that analysis should focus on the higher, more conventional rates – 48k and 44k1" - opus is always 48khz, so that doesn't mean much.
1. Poor sensitivity in bass and treble. See:
https://en.wikipedia.org/wiki/Equal-loudness_contour
2. Limited ability to hear multiple sounds simultaneously, or almost simultaneously. See:
https://en.wikipedia.org/wiki/Auditory_masking
Bernhard Seeber has some videos on Youtube with demonstrations of auditory masking:
https://www.youtube.com/watch?v=R9UZnMsm9o8
https://www.youtube.com/watch?v=bU0_Kaj7cPk
The only fair way to evaluate lossy codecs is with double blind listening tests.
There is such a broad spectrum in thr word mp3 that you need to be more specific. I can absolutely pick the difference between certian high bitrate mp3 encodings and wav files, however a different mp3 with exactly the same bitrate, effectively indistinguishable.
Bad mp3 encoding is a problem, not one I have experienced recently though. I think the bigger issue is people will rip a music video from youtube, then instead of extracting the existing audio stream into it's own container will reencode it. Mp3->mp3 encoding will be lossy just like any other encoder.
And who are therefore forced to hear terrible audio because the compression method only considers the majority.
Description:
High-Quality Audio
Available until February 22
With high-quality audio, you can listen to music on YouTube in the best audio quality.
How it works: Watch an eligible music video on YouTube and enjoy the benefits of higher-quality audio.
Only available on iOS and Android.
It's wild how many details sound people have to keep track of. I know when I upload to Youtube things get smoothed noticeably compared to say Soundcloud. Probably because I've mastered over their -14 LUFS requirement.
I wonder how much of the modern 'everyone needs to use closed captioning watching TV now' comes from the streaming services Codecs and other decisions and not just the A/V sound peoples' decisions. Do movies sound people now need to listen through on something like Streamliner above for every decision?
That said, technical aspects matter, recording practices, mastering for loudness with little dynamic range (so that compressed voice blends with foley and the rest), and indeed encoding can definitely affect the ease of understanding human speech. Speaking of encoding… I couldn’t find it on the site, is Streamliner a one-time purchase?
[0] Well, not every film has to (Upstream Color comes to mind, as an example of a film with relatively little plot-advancing dialogue), but it seems that the majority of productions rely on dialogue and it isn’t going away.
https://www.slashfilm.com/673162/heres-why-movie-dialogue-ha...
Upload 320kbit encoded MP3? Sounds great.
Upload a high-khz WAV? It gets butchered, the top-end turns to glittery noise.
Maybe others have different experiences, but honestly it felt like I'd been duped when paying for the subscription but still got trash quality audio, only to have to pay more.
I would take a guess that a higher bitrate = longer loading times, and viewers care far more about an extra few second of buffering than they care about audio quality, especially when they don't have the original to compare to.
To support seeking you could encode a low bitrate stream, and a high quality stream, and then a number of ramps between these. So when you seek you start with the low bitrate stream and then after a few time units go on the ramp to the high quality stream.
… while yeah… a lower bitrate upfront might lower the required bandwidth and thus, latency, to get enough of a buffer to start playback … all the bloat on the page would be a better first port of call.
How much is loaded for a navigation from one video to the next?
With newer codecs, doing 4k with 2Mb/s isn't unheard of.
For audio, on the other hand, 32 or 64kbps per channel isn't unheard of.
It does not. That's a recommended bitrate for a live streamer (Streamer -> youtube).
Coming out of youtube, the numbers are quiet different.
For example, a random 4k video I just pulled had a video bitrate of 4.5Mbps.
A 720p video I pulled of a talking head had a 369kbps video stream.
That is to say, a podcast style video is likely to have a nearly 50:50 split on audio/video.
YouTube uses variable bitrate for audio, which can vary dramatically in size. Your example of podcasts or "talking heads" is actually perfect. Most encoders are extremely efficient at compressing voices, as they will only have to encode 30-300Hz, and voices have less data variation than images.
Image encoding is just very complex. It'll get better and better, but audio encoders of the same generation will also improve.
Why doesn't this huge AV platform use a better audio time stretch algorithm?
Does anyone know a youtube-frontend that lets me change the playback speed in smaller steps using keyboard?
Also, your use of need is odd as well, and seem to have convinced yourself that the world is wrong and only you’re right. If it needed, the creators would have made it that way
It's called having personal experiences and an opinion. You should try them sometime. It's almost as if appropriate tempo was in the "eye" of the "viewer", and so I was very clearly not suggesting my needs and essentials are universal objective truths in the first place. Getting extremely tired of having to insert "I think", "I believe", "in my opinion" to signal subjectivity in what - I think - are ostensibly subjective contexts, just so that I can avoid subsequent bikeshedding like this.
> You’re now applying to something else, so that’s a bit of goal post moving.
This will be crazy I know, but instead of this delightfully malice-assuming explanation, I simply missed the words where they said "music practice". Didn't help that "timestretch for music practice" is not a feature of YouTube, only timestretch is (as part of the playback rate adjustment feature), so when you were (according to my personal impression of the wording of your previous comment) generally addressing the feature, I replied in kind.
If I was being extra prickly, I'd accuse you of intentionally writing in a way so that you could accuse me of strawmanning you later (goalpost moving is a very loose fit here) for an easy dunk, but of course as someone who reaches for fallacies immediately, you wouldn't do that, right?
This is what distrust sown between people, as well as just plain not being able to know your discussion partner looks like. It's been increasingly frustrating me, and it looks like it's having an effect on you too.
That’s the most bizarre comment I think I’ve ever seen. You didn’t have to reply to my comment. You’re now saying that I assume that the reading comprehension of everyone is so bad that I make comments specifically as gotchas. WTF is that logic? People that post comments in threads without taking the whole thread into consideration are like people that butt their way into a conversation based on the last sentence heard. It never goes well. It’s called social etiquette.
Just admit you didn’t read the full thread and that based on now understanding the full context of the conversation that your comment is out of place and have a nice day
> the reading comprehension of everyone is so bad
No, I am not saying this. This is just your headcanon, it is not even a remotely necessary presumption to have.
> You didn’t have to reply to my comment.
I felt compelled to after being told that I'm "moving goalposts". Obviously. Again, a subjectively perceived need.
> based on now understanding the full context of the conversation that your comment is out of place and have a nice day
Even with the additional context your original comment still rings unreasonable and self-absorbed. It is true however that I do not need you to explain why anymore. Have a nice day indeed.
Besides, their Android and iOS apps do slow music as bad if not worse than on web.
[1] https://bungee.parabolaresearch.com/compare-audio-stretch-te...
If there's reason to believe this is a useful way to handle time stretching, then there's reason to believe the same browser could do it natively just fine.
Oversampling can be a useful internal detail for ADCs and DACs. For example, if the digital audio stream is mathematically converted to 96 kHz with a high-quality FIR low-pass filter and then fed to a DAC, then the analog low-pass filter can have a much shallower roll-off and be more easily designed. Same goes for ADCs, where the analog filter can be simple and gentle, then digitized at 96 kHz, then downsampled digital to 48 kHz with high-quality but more computationally intensive filters. ( https://en.wikipedia.org/wiki/Oversampling , https://en.wikipedia.org/wiki/Delta-sigma_modulation , etc.)
But yes, listening to or distributing audio at anything over 48 kHz is a complete waste of resources. Monty@Xiph.Org explained very well in: https://people.xiph.org/~xiphmont/demo/neil-young.html
Now I just check on YouTube again and they are now back to 128 / 130 Kbps for AAC-LC.
While this does retain the majority of useful information, it explains why the youtube version of your song feels just a little more 'lifeless' than the high quality version you have elsewhere.
The original recording contains high frequency detail that got lost. Your human body uses that high frequency detail to orient itself in space with respect to sound sources (like reverb, reflections, or ambient sounds).
It is interesting from a data storage point of view because this could result in massive savings. Consider audio is recorded at 44.1khz or 48kHz but is actually stored at 32kHz. They have effectively saved 25% in audio file storage at marginal customer experience.
Having hearing sensitivity over 16 kHz is unusual. If you're under 15 years old and kept your ears pristine by not listening to loud noises, you might be able to hear it. Older people are out of luck.
Moreover, even if you can hear above 16 kHz in loud pure tones, there is so little content in real audio/music above 16 kHz that it makes no practical difference.
> massive savings ... effectively saved 25%
Not really. Going from a 48 kHz sampling rate to 32 kHz is indeed 2/3× the size for uncompressed PCM audio. But for lossily compressed audio? Not remotely the same. Even in old codecs like MP3, high frequency bands have heavy quantizers applied and use far fewer bits per hertz than low frequency bands. Analogously, look at how JPEG and MPEG have huge quantizers for high spatial frequencies (i.e. small details) and small quantizers for broad, large visual features.
Good point about the savings. I was using uncompressed format as the reference, but it is indeed unlikely that YouTube serves out lossless audio.
I also should have used the word "delivery" instead of data storage. Those are two separate problems: where the original asset is stored (and how, if they don't store raw originals), and also how the asset is delivered over the web.
If you put something above 16 kHz at full scale and/or if you play it extremely loud then maybe. With typical music content at typical volumes, I doubt it.
Maybe with a browser that doesn't support Opus and gets AAC instead (Safari?). With Firefox or Chromium on Linux I get up to 20 kHz, which by design is the upper limit in Opus codec.
Take unprocessed audio and process it. Then take both the processed and unprocessed audio and add them to your audio software (e.g. Audacity). Now flip the polarity of one of the audio signals.
This allows you to listen to the difference between the two signals and if there is nothing there, guess what, they are the same or the differences are so small that theg are inaudible.
This is a great way to anger people with expensive hifi gold cables, because what is true in the digital also works in the analog.
You only need to make sure both ajdio signals are at the same level (by minimizing the level of the difference).
Server CPUs can encode audio at hundreds of times faster than realtime so there’s no need for hardware acceleration. Back in the iPod era DSPs were used to decode MP3/AAC but now only the most CPU or battery constrained devices like AirPods need hardware acceleration.
Modern CPUs seem to be able encode opus at around 266x speed. In other words they can encode 266 seconds of audio in 1 second. The tests results also don't scale with core count, so the program itself is probably single threaded and therefore you could encode even faster if you have concurrent streams. It's highly unlikely that even with VCUs, a server can encode video streams at several thousand times faster than playback speed.
It’s a noticeable problem in audio production if e.g. a filtered kick drum goes out of phase and sucks amplitude when mixed with the original.
EDIT: And I think with the two-pass approach you need to calculate the filter such that you get the desired effect after two applications instead of one.
In all seriousness, every aspect of this comparison is somewhere between deeply flawed and invalid. No point dwelling on just one part.
[reads]
...Jesus H. F. Christ....
Every generation thinks they discover sex and audio analysis for the first time.
[And don't call me Shirley]
I got better audio quality ripping songs from limewire or Napster in the 2000s.
Why do we settle with this substandard quality? Oh wait, YT barely has any competition and subsidized by G. No need for competition. Just shove ads down users throats and sell of their usage data.
Nah, you didn't, at least not reliably. Half of that were recodes and upcodes that used Blade, FhG or Xing.
Only with torrent technology and community-driven trackers we got reliable distribution that surpassed the official non-physical channels.
Good. YouTube audio quality is crap. Plain and simple. 320 bps MP3 sounds better than anything Youtube offers. And 320 bps MP3 is not even "quality".