Video encoding requires using your eyes
redvice.org
redvice.org
The non-zoomed image looks fine to me, and I (to some extent) know what I’m looking for. Some private torrent trackers that pride themselves on having transparent encodes will look for this kind of stuff; you have to do multiple test encodes tweaking various parameters to ffmpeg, agonizing over A/B screencaps, only to inevitably be told you either missed some minuscule detail in a single scene, or that your encode is bloated.
There used to be a legendary blog called "Diaries of an x264 developer" by Fiona Glaser [0] where she'd go on long rants about various ways to cheat in encoder comparisons [1], much like this.
[0] https://web.archive.org/web/2012/http://x264dev.multimedia.c...
[1] https://web.archive.org/web/20130215095527/http://x264dev.mu...
It does remind me of how stereo & speaker manufacturers sometimes boost treble a little bit (rather than being perfectly "transparent" to the original signal) because it gives the impression of clarity. But ideally each step in the processing chain "colors" the signal as little as possible, because those little differences can add up.
Of course, you won't get a sound as if you're in the same room (without a very fancy setup), so you'll generally want some sort of transformation to get an acceptable output. And artists often want to aim for a certain effect on top of that. But with how things currently are, many of the decisions going into the final sound are very opaque.
It should be valid because it's "neutral". IIRC it's basically a conversion to simulate how a neutrally tuned speaker would sound if you were in the same room.
There are many reasons objective headphone measurements aren't actually objective for you though. The biggest one is that they're taken in a silent room, so a single CPU fan or anything near you makes it invalid. Noise cancelling can mean a lot in practice.
The other reasons are that different people have different ear shapes, some people wear glasses so the headphones can't get a seal, your amp isn't electrically compatible with the headphone, your music is badly mastered so you prefer a headphone badly tuned the opposite way, etc.
Is it, though? Blogspam posts about it waffle over the exact definition, but Olive's original post [0] gives the methodology, "A panel of 10 trained listeners rated each headphone based on overall preferred sound quality, perceived spectral balance, and comfort," and a later Harman post seems to cite the original methodology without comment [1].
Unless the subjective part was just to select between different headphones that had been calibrated to simulate neutral speakers? The posts don't make it entirely clear where the curves originally came from.
[0] https://seanolive.blogspot.com/2013/04/the-relationship-betw...
[1] https://pro.harman.com/insights/akg/defining-the-standard-th...
> People who work deeply with codecs are usually hypersensitive to these sorts of issues that mere mortals like us need to try to see.
I think that kind of shows that the author is unfairly critical.
They're saying "this should not have shipped", when it seems just fine to us "mere mortals".
Yes, video encoding requires using your eyes. But it also seems like it should use normal eyes, not hypersensitive eyes...?
A consumer's perceived preference can be modulated by any number of issues with which the consumer is not properly informed, such as preferring one sweetener over the other, even if there are valid health concerns, thanks to generational brainwashing and an intentional lack of consumer training. What we have today is the consumer market equivalent to the unsophisticated investor market. People getting economically exploited and feeling like it was their own idea.
In the case of video encoders... Actual domain experts should not be forced to cede to the whim of an untrained eye simply for the purpose of capitalists extracting more money from them.
It's one thing for Netflix to experiment with the latest advances in machine learning, as a lot of us do. But when the opening sentence is to their blog is "When you are binge-watching the latest season of Stranger Things or Ozark, we strive to deliver the best possible video quality to your eyes", it's hard not to find issue with A) the commodification of entertainment and patronizing consumer speak and B) a misleading preposition about commitment to quality which the cited response article makes clear is untrue in the case of this product.
Deferring to domain experts in this case is better for every consumer, as they will not be duped into exchanging an increasingly degraded experience in exchange for increasing monthly rates. People feel very strongly about film quality, and for good reason. We're talking about the preservation of art and culture, but Netflix doesn't see it this way, as to them, it's all a commodity, and they manufacture consent for commodification. If the average user can't see this and push back against the degradation of quality that comes with commodification of art, I have a hard time deferring to them over an actual expert.
My suspicion is that none of this mattered though, because the evaluation was probably "perceptual equivalence" vs bitrate. I can easily believe it might be a marginal win over traditional algorithms from that perspective.
That was in the 00s. It is not the encoder's job to remove or filter out all the details. Background or not. There are some caveat to that but that was comparatively speaking at the time, say RMVB from Real Media or WMV.
Worth remembering it wasn't really the internet era back then. People encode so they could fit more things into CD and DVDs. At least that was how it started.
Somewhere along the line Internet, or mobile internet aka iPhone happened. Now everyone watches on a small screen. With all details washed out, people just want to consume. None of the details mattered. What we would only used to do in AVISyth Filter are now done automatically with Encoder. The 10 min Youtube video doesn't care about any of that. And then the 3 min, now the attention span is basically 30s TikTok or Instagram Reels. Worst of all a lot of these attention to details are also gone when doing Netflix or other long from of movie streaming. VMAF 90 is good enough, lets try to minimise the bit rate as much as possible.
We need higher / best quality at minimal bitrate, instead of having bare minimum / good enough quality at lowest possible bitrate. The two are very different.
While the march of internet / tech giant on video codec means someday we may lose out. Somewhat fortunately we still have a small group of old people in the movie production, broadcast industry, and private torrents release group still cares about these.
Hopefully, someday, especially the west, could move back to celebrate greatness rather than mediocre.
That's all to say - I also could not find all the things they are talking about, probably a combination of not being trained, not working a lot with video codecs, and not having the best monitors - but I get the authors frustration, and I'm glad there's people who care about these things! But yeah, I hunted for that color shift and just not seeing it...
For the people who are sensitive to a lot of these, it is more of a curse than gift. Some cant taste the difference between Corn Fed ( or Finished ) and Grass Fed Beef. The colour shift in this article, or how the latest TV perform between OLED, QD-OLED, Four Layer OLED, Mini-LED with different brand.
It turns out being able to "compare" is a skill set in itself. I would assume comparing is also a function that requires more brain power / energy, and most people's natural state would be to conserve that energy.
I have been thinking about this for quite some time. Most people dont know how to compare, or what to compare it to. And precisely because most people dont know how to compare or how not to compare, we need marketing. And I think most successful founder are very good at comparing things. Steve Jobs would be a prime example.
Though, even if I could, this is a new way to preprocess an image before feeding it into an encoder, and the examples have both been fed through the new downsizer, then the standard encoder, presumably at the standard Netflix bitrates and then (I think) upscaled back to the original size.
So if this didn't look a little compressed then that would be a methodological mistake, as you don't use downscaled encoding unless you've already decided that a full size encode has too much quality for your task.
And Netflix generally has incentives set up to reduce quality until their customers notice. That's why they quote stats based on that.
In web dev terms it's like reducing the size of your product images until it hurts sales. You're almost guaranteed to have artifacts visible to image compression experts before you hit the point that it affects your bottom line. If you are targeting customers on slow internet (and again if you are downscaling then you basically are) your sales are likely to initially rise as you get usable pictures to people faster.
I worked for a cable channel (TechTV) in the late 90s early 2ks (until Comcast bought it, laid everyone off and turned it into G4) this was the early days of cable VOD. At that time you had to pay a service by the minute to watch your video before they’d distribute it. That was the QA forced on you by the cable companies.
The fun note is that those services charged double for “adult” content.
What’s the best (computationally not that more expensive than Lanczos) option?
Edit: also some CV researchers write like that (the Netflix writing) — bicubic is like a flag in opencv that they just use. Probably those researchers were much more preoccupied with the researchy problem than actual wide deployment, which is what many researchers do
I personally like the Spline family, and I default to Spline36 for both upscaling and downscaling in ffmpeg. Most people can't tell the difference between Spline36 and Lanczos3. If you want more sharpness, go for Spline64, for less sharpness, try Spline16.
Edit: As far as I'm aware though OpenCV doesn't have Spline as an option for resizing.
If there is user feedback about the quality, then by all means listen to users and at least have an "advanced settings" menu in the app to let users toggle between encoders if they really care.
And do not rely on user feedback. Again. People don't know whats wrong, they don't even know what's right. They just feel that something is off. And only a very small percentage of users would write an email and even fewer would get through the automated AI bullshit or the snotty person who would dare to hand you a HOW TO USE NETFLIX pdf when you want to report something.
If you can't get any user feedback on your products, that is its own problem.
To my eyes it's more pleasant than Lanczos, which has too much ringing.
The Netflix post is sort of bizarre. They claim to be optimizing for minimum mean-squared error (MSE) given a conventional (bicubic) upscaling process [0], but... that should have an analytic closed-form solution, as this post states? You definitely do not need multiple layers of neural networks to achieve it. Then they present VMAF results, but VMAF is very much not equivalent to MSE, so you have no idea if they even improved the metric they optimized for. Subjective results are similarly unpersuasive: it isn't clear if "the deep downscaler was preferred by 77% of test subjects" means they thought it was closer to the original or simply "better" than Lanczos[1]. Netflix may not care: longer watch times are longer watch times. But as an engineer you might want to know if that is because you actually achieved the thing you were optimizing for, or it is due to an artifact of the process that might go away the next time you change something to actually improve what you were optimizing for. You can famously make people prefer one audio track over another by making it louder, and video has similar things around sharpness and contrast (and now, thanks to ML, hallucinated detail).
I agree that you can do better than Lanczos for large downscale factors [2]: you need to do something area-based, like you suggest (I have not looked at OpenCV's implementation, but it could be fine). The biggest thing to get wrong is handling gamma incorrectly, but the right thing to do depends on whether you intend to display the result at the downscaled size or upscaled back to the original size as seems to be assumed here (and whether or not your upscaler is gamma-aware, which it probably isn't).
As an aside to those struggling to see the visual differences, make sure you are looking at the image in its original (1874x1596) resolution: https://redvice.org/assets/images/netflix-downscaler-compare... (or right-click, View Image on the original page). Otherwise you also have your browser's resampling algorithm in the mix. To my eyes on my display there is also a pretty big color shift in the featureless pink background of the painting on the right wall, but when I look at the actual pixel values, that appears to be an optical illusion. Subjective comparison is hard!
[0] Unlike the post, I think this is reasonable. In the past, we did experiments that showed that optimizing for bilinear when nearest-neighbor is used (for chroma) is worse than optimizing for nearest-neighbor when bilinear is used. I suspect something similar will be true for bicubic and bilinear, but these days it may be safe to assume that you will get at least bicubic upscaling (for luma), because bilinear luma looks really bad. I haven't done a recent survey of actual playack devices, though.
[1] It's also not how you report subjective test results: what was the statistical significance? There are standard protocols for these kinds of tests and it would have been helpful to cite which one was used.
[2] Nobody ever says what downscaling factors are being tested here. The example graph shows 1080p to 342p, or ~3.16x, but Netflix goes as low as 144p (from, e.g, 2160p), so they can get pretty large (15x) in practice. A 6-tap filter is not going to cut it.
If you're looking for examples of ringing and hallucinated details, they're really obvious in the framed picture on the right on respectively the character's shirt and the frame.
Also, side-by-side comparisons are hard, its best to flip back and forth between the two images, like opening them in an image viewer and pressing arrow keys. Or cross eyes like with magic eye stereograms so you see them "layered".
But yes, more fundamentally, I think you're right that this is not really image content dependent, it doesn't need any image prior, if all you want is to minimize means squared error after upscaling with a fixed bicubic interpolation.