Netflix: 30% bitrate saving using Film Grain Synthesis
waveletbeam.com
waveletbeam.com
I've always found it funny that the "HBO" branded intro to all their shows is produced of pure visual static. ([1] for anyone unfamiliar.) They literally couldn't have picked something worse to encode for streaming.
I'm also amused when they drop confetti and the screen turns into blocky slush.
Another surprisingly hard to encode scene is water. Wave motion and shimmering is almost impossible to handle efficiently.
Which is not quite as absurd as it sounds. Have a look at the reflections in https://www.youtube.com/watch?v=udPY5rQVoW0
Capturing depth in a satisfying way in a photograph is a skill that can be learned, but often it's impossible. (The skill is really about learning to see the small set of perspectives that do translate meaningfully to a 2D projection.)
My photography teacher told us "If you can't see the same thing (that you see with your eyes) via the viewfinder, don't take that photo". It took me a decade to completely understand what she meant.
One day it clicked at a very mundane moment, but it was pure enlightenment.
(Bonus: the parallax effect can be exploited to convey depth also on a print, actually. A comparatively long exposure time will, depending on light levels, emphasise or de-emphasise motion closer to the camera in motion blur effects, that vary in size and intensity depending on the two things above.)
I'm often watching videos at a resolution larger than my screen can display just to get the extra bitrate.
Every time a movie came up 10 minutes short of the hour or half hour, they padded it with tiny documentaries and previews for other movies.
Noise detection could be useful, if you can get noise that's close enough, it'd be a more than adequate substitute.
It'd also be useful for rain, which video codecs also struggle with.
Another thing is, a lot of modern video was generated, and composited - why not have that reflected in the codec?
You could directly encode the elements as understanded by the process used to generate them, rather than attempting to reconstitute things from a stream of bitmaps.
You might like Gan Theft Auto: https://www.youtube.com/watch?v=udPY5rQVoW0
On the other hand, deep-learning based compression with autoencoders or something similar is pretty promising, since it can learn the constituent elements in a more general way, independent of what program was used to create it.
First, set up the scenario. You're going to binge-watch an HBO Max series that is 10 x 1 hour episodes. Each episode has the HBO static card (7s), a 1m45s intro sequence, and a 2m30s trailing credits sequence where the first 30s is episode-dependent but the last 2m is constant across the series.
This is a best-reasonable-case scenario, since we know that people do sometimes binge-watch a series like this. We can construct unreasonable scenarios which would do better (10 minute loop of an aquarium played for 24 hours) but they would be unreasonable. More often, I think, people tend to watch one movie or one episode of a show and then switch to something else or leave.
The receiving device needs to have 3m51s of storage available for reuse across episodes. That's not unreasonable. At a 4K streaming rate of 10GB per hour (Youtube 4K uses more than this, Netflix uses less) that would be 260MB or so.
That number does sound unreasonable, though. A typical streaming device has between 1 and 4GB of RAM available -- an Amazon FireStick 4K has 2GB, various modern Roku devices have 1 to 2 GB -- and eating a quarter of that speculatively does not make sense to me.
Maybe it's just an illusion. But I swear I can see the seams where the widescreen pixels start
(Modern video formats are not deterministic in that the same source always encodes to the same output for the same format. They are more like programming languages that they provide as set of tools you can use to compress various different kinds of redundancy out, but it's the job of the encoder to find that redundancy. For example, one of the tools typically available is to provide a source image to an off-screen buffer, and then when rendering the screen, occasionally provide offset/length pairs into that buffer instead of pixels. This could be used for efficiently encoding the duplicated noise.)
For some reason ffmpeg's H.265 implementation doesn't have a film profile, I wonder why?
It would be nice to have a fantastic implementation of an h.265 encoder like x264, but alas.
Still, it was really interesting that
ffmpeg -i bluray:[filename] -c:v libx264 -crf 22 -tune film -preset slow -nr 500
produced a smaller file with higher visual quality than ffmpeg -i bluray:[filename] -c:v libx265 -crf 25 -preset slow
for highly grainy inputsI remember noticing as a kid that some channels on cable TV were noticeably worse quality (far more blocky iirc, looked very MPEG artifacty) than the analogue broadcast equivalent and that others looked very different between even cable and satellite tv (iirc More4 was one of the more obvious ones, looked OK on Sky but awful on NTL).
The HBO static intro screen was worst case for compression but gave a pretty good indication of the quality (of the bitrate/compression) of the rest of the show; which worked for broadcast, recordings or downloads. Probably not deliberate, but actually kinda useful.
Update it every decade or so to reflect artefacts of whatever in-house codec they use at the time. The ramp-up would be a creative exercise for the techies working in the studio.
This reminds me of Steve Yedlin's, cinematographer on everything from Brick to Star Wars: The Last Jedi to Knives Out and its 2022 sequel, Glass Onion: A Knives Out Mystery, who showed through research that film can be completely emulated using digital techniques and mixed digital and film shots during Star Wars.
https://www.audiocheck.net/audiotests_dithering.php
https://www.polygon.com/2020/2/6/21125680/film-vs-digital-de...
https://www.yedlin.net/NerdyFilmTechStuff/MatchLensBlur.html
I’m not sure how apt this analogy is. Dither is only useful for mitigating artifacts of quantization noise; it does not generally increase perceived detail of any arbitrary signal. More specifically, it is introduced to reduce the harmonic content of quantization noise, at the expense of a higher overall noise floor. Suppose you have a 1kHz sine wave; quantizing it will introduce harmonics that peak at, say, an average of -70dB, with an absolute noise floor of -120dB. Adding dither will raise the noise floor to -90dB, but reduce the harmonic peaks to -100dB. So while it actually decreases the true dynamic range, it increases perceived signal quality by removing the harmonic content.
These spurious harmonics occur because quantizing a signal introduces periodic artifacts. For example, suppose our analog sine wave can continuously vary between 0-7, and we quantize it to 3 bits (discrete values 0,1,2,…,7). Any analog value 4.7 will always be rounded up to 5; in a sine wave, the value 4.7 will occur periodically, thus resulting in a periodic rounding artifact, leading to harmonic distortion.
In order to prevent these periodic rounding errors, dither needs to be added pre-quantization, so that 4.7 can sometimes randomly become 4.4 and get rounded down to 4 during quantization. Adding “dither” to an already quantized signal (e.g. digital video) would just make the apparent picture noisier.
As a test, try quantizing a full-color image to 16 colors, and then adding back some random noise. It won’t look any better. You need to strategically dither the 16 colors with knowledge of the original full-color image.
https://www.hdnumerique.com/dossiers/962-test-4k-ultra-hd-bl...
When those smooth-then-pixelated-then-smooth lines transition to even-smoother, like 4K / HDR consistently hits, grain looks like the muddying noise that it really is.
Lots of stuff looks wrong with higher resolution and framerate, it’s not the thing (like grain) is inherently wrong, it’s just being done poorly.
Foley sounds and lighting are the things that annoy me the most about modern cinema. Those two characters looking into the sunset in the background… where exactly is the cool white light illuminating their faces coming from?
The same place music in the film is coming from.
Presumably, "Rian Johnson's cinematographer"?
Added film grain to digital to me is very dumb. It is hard for me to think that it is a strong authorial intent and not just a legacy idea that is basically assumed by default because that is how it used to be. Same thing bothers me in games. I can’t even begin to understand why someone wants film grain slapped over a game. Another reason why I think of this as an assumed convention instead of a meaningfully intended or utilized convention in modern media.
As for the other reason, it's sort of like how we perceive people as older due to the clothing they wear and their hair style, our cultural reference for how old a piece of media is can be tricked by the mimicking the imperfections of the filming techniques of the time. Take a photo of old brick building with a Polaroid camera and ask folks when that was shot, I would bet that most folks would say 80s or 90s more than the 2020s.
This Polygon article touches on some of the process and considerations: https://www.polygon.com/2020/2/6/21125680/film-vs-digital-de...
Anything that went through a VFX pipeline was denoised (so that trackers, painters & rotoscoper can work more efficiently.) once the effects applied, then film grain was put back in. (even if it was to be lasered back out onto film.)
Film stock improved significantly from 2000 to maybe 2008. grain getting smaller. optical resolution getting better. (well maybe digital intermediary got better.)
Film grain doesn't mean anything, except what the public have been conditioned for it to mean back when it was a natural chemical process.
(You of course can just dither it first and save the result in 8 bit in encoded result; but that would dramatically increase bitrate).
I pay for Netflix Premium but currently don't have any 4k screens, but I imagine the extra bitrate would improve the quality at no cost (to me) when watching at e.g a 2560×1440 monitor.
Unless you're using Microsoft Edge, you can only get Netflix in 720p on Windows: https://help.netflix.com/en/node/23742
Conan The Barbarian DVD is the noisiest film I can think of off-the-top. Strong de-noising makes it look far, far better. But just about every movie out there benefits from modest de-noising.
Before I got my Shure SE215-es I ended up doing a lot of research to figure out which headphones were actually going to give me a reasonable balance between treble and bass.
The Aonic 3 is my daily, although using the cable that came with the 425 (EAC64BKS - more pliable than the transparent one)
Sure, but I don't see where I said the opposite! I didn't say the SE215 was the op of the line, just that I did a lot of research first :)
Part of that research included finding out that the SE215 was only 80 quid, but the SE425 was 300+ quid...
The meaning of art is up to the viewer, not the artist. There's nothing wrong with having a preference, like using things like temporal denoising or increasing/decreasing saturation. Even an IMAX movie theater isn't enough because that's not what the film was graded with. Pretending that art shouldn't be modified because it moves away from the artist's intent is a bad faith argument. There were certain House of Dragon color grading decision that made me feel very disappointed, and I was watching in Dolby Vision on a nice high dynamic range TV.
Artists' intent is useless if 99.999% of the population can't enjoy the actual art.
https://www.vulture.com/article/why-house-of-the-dragon-epis...
To address your second- I think that’s an example of the “slippery slope” fallacy. I’m stating only that film noise and grain is a big part of the atmosphere of many movies and shouldn’t be removed; I’m not suggesting that a viewing experience needs to be done in a world-class theater under perfect conditions to capture artistic intent. Anyone with a basic HD television can see analog noise and grain.
There is an endless list of poor Bluray and UHD Bluray transfers where the the film grain makes it unwatchable. The unfortunate thing is that the nicer the display is, the sharper the image and the more distracting the grain. In some cases, I would trust the AV1 endcoder decoder to do a better job at adding a more reasonable amount of grain than the 4k transfer did. It's not a slippery slope, bad transfers and poorly color graded films/shows already exist. Mistakes exist in both directions, the UHD remaster release of Terminator 2 is a exceptionally horrible example of too much DNR. In my opinion T2 is bad enough to bring real actors into the uncanny valley.
The developers aren't doing anything different from what films do when going through an analog to digital transfer.
Jesus. Every engineer who's worked on color implementation just choked on their drink.
You can literally scroll down a little further and find blog posts on how iMessage green vs blue saturation creates tiny changes in readability. The human experience matters a lot more than what engineers (who can't tell apart two shades of women's fingernail polish) decides matters.
Readability is dependent on screen brightness and difference in contrast between background and foreground colors. These decisions aren't based on what engineers assume matters, it's dependent mostly on A/B testing and perception psychology. It's the same kind of research that went into making the OPUS audio codec and deciding which parts people will notice missing and which parts don't really make a difference. Tiny to moderate changes in visibility/readability make a much smaller difference than you imagine.
https://www.colorblindnesstest.org/farnsworth-munsell-100-hu...
The hobbit doesn't feel like a film in 60fps. it feels like a play, or something from TV.
Now. to your point, that motion smoothing stops this, its not strictly true. There are still artefacts that happen because its film. this then becomes the marker for "film".
> The meaning of art is up to the viewer, not the artist.
yes, but the delivery is key. your point about imax doesn't really add up. Unless you are printing out IMAX prints (which most people don't) they you'll at least do a technical grade for the IMAX to make it look like the version that the DoP/director agreed.
The House of Dragons chose to make it look like that. Mainly because I suspect the director was trying to chase the fashion for removing colour from everything. Grading, like most things is prone to fashions, and pushing against "what came before" is par for the course. Some people do it well, others are shite.
I’d say it is more like asking musicians to stop adding feedback — another phenomenon that began as an error but became an artistic tool — to their songs. Or requesting that sound engineers stop the engaging in the loudness wars.
Same if someone wanted to make a silent movie or ask me to join a dating site where all the profile photos are daguerreotypes.
what's the benefit of doing this? you waste your time and you lose color option, I see no advantages over just adjusting saturation
If used as element to paint the scene noise is fine and can work well, but the trend to just slap the grain on anything "coz that's how movies looked before" is just silly.
(kidding)
Total tangent here, but in my opinion, most HPPD are optical/perceptual defects that everyone experiences all the time, but normally our minds block, filter, or compensate for. Taking psychedelics makes people aware of these defects, and once you notice them, you can't "unsee" them.
My personal experience with this is that very mild astigmatism causes diffracted halos around light sources. I never noticed these until taking psychedelics. Now I always notice them, but they're certainly caused by defects in my eyes' lenses and not by some permanent direct effect of psychedelics. Similar with floaters, visual snow, POV effects and so on.
Many of these effects/defects become exaggerated whilst tripping and hence get noticed, then when people sober up and still notice them, they think it's something new, rather than realizing it's something that's been going on unnoticed the whole time.
That is an interesting way to frame it. I see it more as an expert use of 35mm film and lighting, and the grain is just present because it is inherent to the medium.
[1] https://www.hdnumerique.com/dossiers/962-test-4k-ultra-hd-bl...
Alternatively you can paint/airbrush every set piece to add that used/worn look, but I don’t think it’s economically feasible.
He told us how shabby the set looked like in reality. But none of that was visible on PAL tv.
I guess that's the opposite of the effect you were talking about.
It leaves gaps in the image for you brain to “fill in the gap” and add perceived detail.
Interleaved scanlines are a great trick to increase the temporal resolution, but don't do much for (or against) spatial resolution.
Or what do you mean by scanline effect?
E.g. the non-digital aesthetics of Tarintino films contribute a lot to their style in my opinion.
Plastering fake film grain ontop of something is not really my ideal, but going full film is also probably not something people can really afford in terms of time and budget...unless you're someone like Tarintino who's already at the top
Adding extra fake noise to the movie solves a few problems such as ensuring a consistent look from scene to scene, as well as avoiding video artifacts like banding.
It’s also an artistic choice just like the choice to use telephoto lenses that blur backgrounds into fuzzy balls of light.
2004 Collateral and 2009 Public Enemies were early all digital movies full of digital noise, it looks BAD.
https://tvtropes.org/pmwiki/pmwiki.php/Main/RealityIsUnreali...
Digital production methods are great at producing data in very predictable, reliable ways. But our brains our great at spotting patterns, and they can tell really quickly if something is unchanging and can be ignored. So while the challenge in the analogue world was creating order (a faithfully reproduced signal) from chaos, in the digital world it's flipped. We need to find ways of adding some chaos to make the resulting signal interesting to our brains.
Netflix (and most of the other streaming services) are dying because they no longer focus on customer experience, but on management goals. The recent move to declare password-sharing illegal is evidence - the nail in the coffin of the recording industry was when they started suing their own customers too.
Netflix is only optimizing the pipeline here, trying not to mess with the artistic intent.
The level of grain you get from digital cinema photography is mostly by artistic choice and added through the camera or in post production. Sometimes the movie is shot digitally and transferred to film, then scanned. This was done for Dune (2021), for instance.
Also, ARRI (who make the most renowned cameras for cinematic use) now specifically let you choose a grain texture on the Alexa 35 that is imprinted in the digital material. [1]
[1]: https://www.arri.com/en/learn-help/learn-help-camera-system/...
That is so sad. Such a waste if it can just be done during decoding.
One could call it something like Dolby Vision Film Enhancements, that would just mandate decoder support. Those who don't support it get a fallback, like the usual.
Grain is probably something that will disappear in a few years as the old film-oriented directors die off. Along with 24FPS.
24FPS is gonna stick around for a while. I've not seen any high framerate movies that didn't make the movie set look like a...movie set. This is great for sports and documentaries where a better visualization of the subject is always better, but in movies it just makes everything look like it was manufactured as quickly and as cheaply as humanly possible--because it was. I feel like high framerate, even more than 4k, lets you "see" the set and the makeup and the props as a hastily crafted illusion. Everything looks fake. Even the body language looks like acting in high framerate.
Meanwhile, in sports and documentaries everything looks more real, because it is real; there's no set, no props, no act--well less acting. More acting in soccer.
I'm sure at some point in the future set designers and the props departments and all that will figure out how to make 48FPS and above look good, but we don't live in that future yet.
That happened when HDTV came in. Sets that looked OK at 525 lines and NTSC color looked terrible at 1080 lines and hard-edged color.
I doubt it. Oil paints are still here, even though digital painting is easier and quicker.
> Along with 24FPS.
24fps is film. 24fps will not die in our lifetime. In the same way that vinyl records are still here and thriving, despite being a shite medium.
24fps I suspect will be displaced by some sort of VR as the predominate high-end story medium.
I seriously doubt it. For one, VR headsets are broadly incompatible with the theater business. You can make it work in smalls scales or in amusement parks (where the park admission and food sales are where the money is made) but broadly it doesn't work because the headsets are expensive and you need to disinfect them to prevent lice. It would be like bowling alley shoes except you wear them on your face... very gross.
At home, wearing a VR headset to watch a movie works for a solo experience, but ruins the fun of watching movies with others. VR is bachelor tech, doomed to be niche.
The beauty of film grain is back in the digital world. The grain will be reduced before encoding and transmitted via a grain table in the AV1 bitstream. After decoding the grain is added back and now it is possible to enjoy film look even at ultra low bitrates.
This is a very subjective thing to do! I don't like it.
So, for the majority of those that like film grain it feels deceptive, and for those that don't it's an added annoyance. It's only appealing to what I surmise is a vanishingly small subset of Netflix's customers that just want the visual artifact of grain and don't care where it came from or if it is original to the analog transfer.
Also with film evolution, the film grain has changed.
However I can't find any reference to it online. There are loads of "degrain" workflows for things like nuke and AE, but thats not really the kind of evidence I would accept if I were you!
Would be interesting, and to me a bit disappointing, if so.
So the golden standard (althought I am somewhat biased) for film scanning is this: https://www.ebay.co.uk/itm/282841082109 (https://www.filmlight.ltd.uk/products/northlight/overview_nl...) The Northlight film scanner.
The reason why its good is because its not a realtime scanner (like a spirit) so it has the time to do "perfect" registration. This means that you should be able to re-scan the same thing twice, and any VFX/added in stuff will still line up.
It has the light source well away from the film, so it doesn't heat it up. It also does an Infra red pass to allow for dust busting. (that is the removal of blemishes from the film)
This scanner is the one you'd use for restorations/rescans and "digital remastering".
With that in mind, its capable of scanning at 8k (which is then down res'd to 4k from memory).
If you are doing a straight film-to-digital with no remastering (which are pretty rare) then thats what you'd do.
Now, to answer your question!
For restoration, there might be a bit of degraining and re-painting & reconstruction, followed by a regrain. This is because film grain is really hard to rotoscope/paint over. Because the edges all move, and its a general pain. So for some limited restorations there will be some "fake" grain. But only as much as needed to make it look like the rest of the film.
Think of it as like a conserved oil painting, a light cleaning, some post processing to bring out the colours/contrast, bam off to digital.
for "re-mastering" there might be some upscaling. Now, its my understanding that upscalers remove grain, because otherwise you'd just get large amounts of noise on the screen, rather than detail. But there are a large range of ways to upscale. The gold standard is basically re-painting everything: https://www.youtube.com/watch?v=IrabKK9Bhds or with a 4k remaster of blade runner: https://youtu.be/tN0MtsIz6H0?t=15 where they have spent a lot of time removing film noise.
in those, the film grain is removed almost completely.
the BFI from what I have seen only do restoration and dust removal.
I really dislike chroma subsampling and Bayer-pattern sensors and all the other little steps in the video pipeline that sacrifice color density. I'm stuck in a world where color is always blurry but apparently, everyone else seems fine with it.
Thanks for trying Netflix, but I think I hate this procedural-photography future we are slipping into.
But for some time now that magic was employed with a perceptual goal to make the raw data look "correct".
I'm not sure if I would object to fake depth-of-field or fake film-grain if it were done absolutely convincingly or if I just dislike the whole idea of trying to add "artistry".
I think this article is just a plug for the company website and should be removed from HN.
Of course, film grain is a common effect added during the editing process.
>Easiest way to tell is by looking for film grain
>Industry finds way to synthesize this so money can be saved on bandwidth
The question is just how many bits per second you need.
Strictly speaking ("film noise"), I suppose you're right. But we are all aware that there is such a thing as CCD noise, right? That's not being synthesized.
I wonder if one is preferable to the other: that is, is CCD noise worse than emulsion noise? My sense is that the CCD, with the Bayer filter in place, gives you wild chroma noise while film gives more of a tonal noise.
I was making a more stupid point: these days image you are seeing on screen is always synthesized from 0s and 1s. No matter how that stream of data was original produced (ie by scanning actual film stock).
modern slow 35mm film (now there are loads of types and speeds) has an optical resolution of something like 5-12megapixels. a full frame CCD has easily got an optical resolution of 50mp.