Streaming services' obnoxiously loud ads become illegal on July 1 in California
arstechnica.com
arstechnica.com
Well, stop "trying" and fix it already. It's your own damn system.
loudness doesn't have a standard. different streamers want different loudness settings. ad platforms would have to have multiple audio streams limited to the different loudness settings. it's not a hard problem in the least, but one of adding to the complexity of content management.
or the streaming platforms could take over ad delivery and take it on as part of their internal content management.
Then the streaming services are failing to punish their ad providers, so they still have culpability.
Bonus: the clients can sample the advert audio and send back the computed ReplayGain, compared to the claimed ReplayGain. Add a penalty clause in the contract for every single advert delivered louder than claimed.
It literally does. Audio players were implementing ReplayGain support 25 years ago.
> ad platforms would have to have multiple audio streams limited to the different loudness settings
This is trivial to solve. You have a standardized volume and a scale factor for each service. Or you use ReplayGain or similar for the same.
> it's not a hard problem in the least
No, it’s really not.
The problem with video content is there can be extended very quiet parts. If you try to bring the “average” volume of an ad to the same “average” volume as the last 30 minutes of a drama it could end up being insanely low. The average level of the drama being low due to long quiet parts.
Not always the case but this problem isn’t as simple as it might seem.
My main hope is that this doesn’t kick off a “loudness war” for tv/movie content on streaming in an effort to get its average volume up as high as the ads.
The streaming service decides where the ad breaks are. For video on demand, which is most streaming, they can precompute the ReplayGain for the 30s of video leading into the ad break.
For live video, the client can compute it.
Either way, the client can then adjust the volume of the ad to account for any difference in perceptual volume.
Conceptually I don't see how hard it is for the stream provider to implement this. Whether they want to implement is another story.
If you can't create, copy: https://en.wikipedia.org/wiki/EBU_R_128
Ad is coming from provider. Volume is not.
This is just as much the fault of streaming platform exclusive content having no time to breathe. Nobody sane really wants to sit through an episode of anything that lasts longer than 20 to 30 minutes. The binge watching era died when the ads became forced upon all tiers, but there is a compromise that is proven to work. Just go back to the old episodic format of broadcast/cable already.
I think it's quite obvious that you can have a system to represent 17 stops of dynamic range without anchoring brightness to an absolute scale, and that you can have a system to represent 6 stops of dynamic range that is anchored to an absolute brightness scale. The two properties are separable.
The only reason that you can get away with saying "half the point of HDR is that the filmmaker can say I want this bit to be this bright" is because of exactly the conflation I'm criticizing: the "HDR" label as used for marketing was redefined to mean high dynamic range plus a standard representation that encodes absolute brightness.
I couldn't find a Youtube source for this but it's mentioned extensively online: https://audio.rswaver.com/blog/youtube-loudness-standards
It’s right and proper YouTube do not do this to people’s work.
Though as those are rare as hen’s teeth, perhaps you might as well.
If it’s a video of someone speaking in a quiet voice it shouldn’t have to have an average volume level of a new Metallica video. Forcing people to use heavy dynamic range compression to get it there would not be good.
What I find most bothersome is the timing. On linear TV the ad breaks are planned to fit with the show. On YouTube they can happen at pretty much any time and often step on a dramatic moment or compelling scene and totally break the mood.
With their ability to automatically make transcripts of video, and their AI models, surely they could make something that could look at the transcript ahead of time and figure out places where ads could go that would avoid this problem, couldn't they?
[1] For several months I've started the day with ad blocking off on YouTube. If they annoy me too much it goes on for the rest of the day. I follow these rules. (1) Ads that are relevant to me do not change my annoyance level, or maybe even lower it. (2) If the ad that interrupts what I'm watching is skippable in 5 second or it is non-skippable but not over 6 seconds and is not followed by another ad it does not change my annoyance level. (3) If there is a second ad and it is skippable in 5 seconds or non-skippable but not over 6 seconds and not followed by a third, it will raise my annoyance level, but they can get away with this a small number of times. (4) A 15 second non-skippable ad will raise my annoyance level enough that as soon as I get back to what I was watching I note the time, turn on the ad blocked, hit refresh, and seek back to where I left off if the refresh loses my place. (5) Too many ad breaks will also raise my annoyance level enough to turn the blocker on.
For the first few months this worked great. It was is if their algorithm had figured out what I was doing and adapted. I'd always get 5 second skippable ads, and they would be spread out far enough apart that most days I wouldn't turn ad blocking on. But lately, over the last few weeks, they are doing a lot more non-skippable 6 second ads following by skippable ads or a second 6 second non-skippable ad, and they are more likely to insert way more ad breaks than they used to. They now almost always are in the ad blocker by the middle of the day.
Protect your attention (and frustration levels) by running adblock full time. If you want to influence their behavior so everyone stops running annoying ads. You have to give the ad sellers a reason to re-evaluate what ads they're willing to show. The only way you can do that is increasing the number of users who don't view ads. Starting with in off, makes you a user who views ads, not the kind making a point.
It's unlikely that you're meaningfully able to influence the system. You're much more likely to respond and adapt to the system, so it's more likely that youtube has trained you to watch ads.
You don't need to pay YouTube protection money. Get a different browser.
Boo-freaking-hoo. Cry me a river, poor streaming services without the technical know-how to calculate an ad’s volume. We can’t expect them to know how audio works!
> Additionally, as the opposing groups previously pointed out, streaming services must contend with a broad range of output devices, including TVs, tablets, and phones.
See, that’s just flat-out lying. What’s this mythical circumstance where playing audio A at the same volume as audio B on one device will magically make A louder than Bon another? Especially when dealing with server-side ad insertion, as the article discusses, where the service has full control of the input files and the output stream? This reads like a restaurant trade group claiming that it’s impossible to know how much salt they put in the gravy.
TL;DR distract the enemy with resource sinks
I guess the solution is to switch to a proper ad insertion company that normalizes to -24 like you’re supposed to, but that’s not cake either. Especially if contracts are signed.
What people are now talking about here is to normalise the “average volume” (by using compression to increase it, or just lowering the volume to decrease).
The problem is the “average volume” of 30 minutes of video, depending on the content, could be quite low.
So matching the ad to that level might not work. The nightmare response we might get is the average volume of TV/movies will be pushed to the limit in future, removing any dynamic range or impact. To comply with the rules for ads. Sigh.
Oh, the humanity...
Regarding your second point: as any audio engineer or electronic musician knows, the same exact audio absolutely will sound very different on different speakers, depending on how well they replicate various sounds, what level of gain is being applied, and the volume (which is different from gain, although people confuse the two).
That's even before you get into the fact that many modern devices, like smartphones, will apply their own compression or sound processing before playing the sound, sometimes to compensate for those deficiencies and make them less noticeable, and sometimes to "enhance" the sound.
Loudness/volume (technically different things but let's conflate them here) are also unintuitive because human ears don't have a flat frequency response curve, and some things will be perceived as louder despite being the same volume, or vice versa.
Advertisers actually can (and do) take advantage of this, by using sound engineering to make things feel louder while staying within the desired volume, by targeting the way humans perceive the sound.
This isn't a defense of the advertising/streaming companies here, because it is a solvable problem. But it is true that this is a problem that they need to solve.
Start at 1/4 the volume they use now.
After all, they don't need to approach compliance tuning and debugging from the loud side. They can start at a whisper and work up.
(I hope they get fined into bankruptcy, if they try to claim they're "working on it", but do so from the loud side.)
Regarding the perceptual volume differences: while true, that’s also a solvable problem. Output volumes can be calculated using standard curves. In any case, TV broadcasters have had to figure all this out years ago.
Sorry, but all of that is obtuse. The fact that some digital audio can be perceived as much louder than others –– yet it's all limited to the same digital range –– proves they aren't similar at all.
There is no such thing as a standard curve for compression. Source levels vary almost infinitely. Accurately separating and reducing sound after the fact, without turning the whole thing to mud, is considered to be an impossible technical challenge.
Next, TV broadcasters worked on a predetermined schedule with predetermined advertising. This gave them time to inspect and approve ads in advance.
Streaming ads are generally served just in time from third-party services to the streaming host. FFMPEG gets the output from the stream host, but the host has to combine content together from multiple sources (entertainment + multiple ad servers) into that single stream. Currently, sound-level is completely at the whim of each ad server, as well as each ad producer. Meanwhile, the final output is at the whim of the streaming host: 24-hour-news streaming sites probably have different audio standards than Apple TV+.
Ultimately, AI could potentially be used to solve it, since it can generate / make-up new sounds as part of reverse-compression. But it would still have to be done in advance by the third-party ad servers.
I read the article. It specifically talks about server-side ad embedding, i.e. where the service is inserting ad content into the streams, and therefore, by definition, has access to the ad content. They can do the calculations on their end during the embedding process and normalize volumes there before transmitting the result. To make things even easier, they don’t have to calculate the ad volume each time one’s streamed, just once per ad they’re going to serve.
And finally, all of this is a solved problem for TV broadcasters. They face the same problems: advertisers send them content to air, then the broadcasters are legally required to normalize the ad vs content volume, and they do. If this is an insurmountable problem that the streaming services face, they can drive over to their nearest TV station and ask them how they manage to pull off this technological feat.
Conflating DCT-based compression of audio data (like MP3) to dynamic-range compression of an audio signal (as done on an audio compressor during production) shows how little your grasping the problem.
It doesn’t work very well for TV in my opinion. Ads still sound louder as they are mastered differently.
I’ve no skin in this game, and no desire to ever see or hear ads. I just hope we don’t kick off a “loudness war” for TV/Movie audio by mandating the average volume across entire programs has to hit some high level like modern music or ads have.
Given how audio codec compression, eg MP3 files, work it’s easy to efficiently calculate the perceived human ear loudness of a sound. The hard work’s already been done. While of course there are intricacies and edge cases, it’s not impossibly hard to match the volume level of an ad to the volume level of the content immediately preceding it. TV stations already do this, by law.
There’s also a weird undercurrent in this thread conflating audio dynamic compression with loudness. Level compression does not imply loudness. It implies a constant volume, at whatever level the engineer picks. Compress the heck out of ads if you want, then match their starting volume to the preceding volume of the streamed content, and you’re golden.
The streamers should be responsible for the signal. If the device front end has crazy frequency response or the backend does weird DSP tricks, that’s on the device manufacturers.
Consider a streaming movie with surround sound with an inserted ad that is in stereo. I'm playing it on a 5.1 home theater system and you are playing it on a stereo phone. Your system is mixing the surround sound down to stereo.
When your device does that it applies attenuation to the program so that if several channels in the 5.1 stream have something loud all at the same time it won't be too loud in the down mix for stereo and clip. When the commercial cuts in your device recognizes it is ordinary stereo and it doesn't need to down mix. It goes straight through without the attenuation that down mixer applies to the program.
Whatever level the commercial is really at relative to the program, it is going to sound loader than that on your system because of that attenuation difference.
On my device it is not attenuating the 5.1 program since it has all the necessary channels. However, if the commercial is at the same level as the program it will actually sound louder on mine. That's because the same total level of sound split among 5 speakers perceptually seems less loud than the same total level coming from stereo speakers.
The streamer can do loudness normalization between the program and the commercial. It can calculate what the perceived human loudness will be at any time in the adjust the levels so that on my device the perceived level of the 5.1 program when it gets to the commercial will match the perceived level of the commercial on stereo.
But for devices that are down mixing to stereo there is still going to be the attenuation the down mixers uses, and that differs from device to device. That limits what can be done server side to get the program and commercial to match.
Some multichannel formats do include metadata for the device telling it how much to attenuate when down mixing to stereo. If all the device supported that it should be possible for the server to fully take care of loudness matching. Otherwise you probably need device side normalization.
Another approach would be to up mix the stereo commercial server side to whatever surround sound format the program is using. Then they could do server side loudness normalization between the program and the commercial without it being messed up by the difference in how stereo devices down mix.
I'm not sure why that is generally not done. LLMs are suggesting several reasons but I have no idea if they are reasonable. I'll leave exploring that to someone else.
Passing legislation for this kind of thing is almost impossible. This is the so-called "rule by man".
It is quite an interesting thing to see a country across the ocean solving this kind of problem through law.
Leaving room for nuance reduces the seeming capriciousness seen in the enforcement of some laws that look heavy-handed when applied strictly, while said underspecification can allow for abuse instead.
As long as people are individuals with their own volition this tension will exist.
Though, I don't even know if Apple TV has an ad-supported plan. This is mainly wishful thinking here :)
Apple TV one of the few streamers with decent quality audio.
We don’t need a “loudness war” for TV and movie content.
Two reasons:
Highest quality available for every media. Bluray remuxes are a game changer, when available.
Every media in one app.
I also self-host Navidrome in my homeserver.
They liked the "background noise". They'd read with it on, have conversations shouting over it, and so on. Baffled me. I often wondered, why not just plop down a food blender and leave it on?
Why pay for cable?!
Mindless scrolling is the modern version of this, but it's worse because there isn't even a shared experience that might spur a conversation.
People like you and me are quite the opposite: we hate external structure and long to be left to our own thoughts and devices. It's not too dissimilar to micromanagement in that respect. What's the point of having a brain (and the rest of the body, that matter) if you can't use it?