Guide to Adopting AV1 Encoding
bitmovin.com
bitmovin.com
I'm probably not up to date with newer/hardware encoders and please let me know if my view is outdated.
In addition, there are fast threaded AV1 encoders you may be overlooking, like SVT-AV1. For non-realtime, my favorite is av1an, which also yields better quality than is possible from aomenc and works with pretty much any encoder/codec: https://github.com/master-of-zen/Av1an
x265 should be what is compared against AV1 when discussing quality and encode speeds.
We use x264 for compatibility purposes, if your device is intended to play video, it will decode x264. x265 decoders are in a lot of devices at this point, and AV1 is just now starting to see representation.
x264 is like .jpeg and will probably never die.
It will probably eventually be like mp3, where users are still reflexively encoding to it without a good reason.
Netflix also uses H.265/AV1. Amazon, HBO, Disney, etc also use H.265, but I believe only on higher resolutions.
On Apple world even phone pics uses HEVC (HEIC).
I think it's a total mix honestly (Because tiktok/X/Meta uses h264 as far I know).
More technically correctly, more than half the currently in use devices in the world do not support more than H.264
Like all my Smart TVs, my kitchen iMac, my parents’ phones etc. only my phone and my laptop support HEVC.
Netflix uses AV1 when on mobile data, even if the HW doesn't support as well (for TVs, it need HW decoding).
The only way to not use VP9 on Youtube is if you don't have the codec installed, which is super rare because the codec is open.
https://giannirosato.com/blog/post/nvenc-v-qsv/
x264 used the medium 10 bit preset, which is a bit of an oddball because 10 bit AVC is "unofficial," and many hardware decoders don't support it.
Specifically, it has a "VMAF target" mode that encodes a few frames from each scene as samples, measures their VMAF, and then boosts or reduces the encoding preset for that individual scene based on the result. It also has a "black boost" feature to allocate more bitrate to dark scenes, which encoders and various metrics tend to misrepresent.
Also, the possibilities with the VapourSynth support are infinite. Some simple examples include denoising a noisy video on the GPU for better quality than the encoders' CPU denoising, or deblocking with custom parameters, or upscaling with some model in PyTorch. And that's just the start:
A faster encode might allow using a slower speed/preset setting, which would increase quality per bitrate a little bit. But I don't really consider av1an a tool that increases encode quality, to be honest.
Do you happen to know if that's still the case?
(I guess for use-cases such as live streaming it doesn't matter that much, but for video that ends up in some archive, it's probably less acceptable)
Later on in the video, there are some graphs comparing Intel's AV1 encoder to SVT-AV1 at different speed presets. Even one of the faster presets (9) will comfortably stay above AV1 quality according to VMAF, and if you don't need real-time speeds you can lower the preset to get further ahead of the hardware encoder. (BTW: That video is >1 year old now, and SVT-AV1 had some significant updates in the meantime too. So the software side is probably looking better now.)
Don't underestimate dav1d. It's a highly optimized software AV1 decoder:
https://code.videolan.org/videolan/dav1d
On my nine year old system, 1080p60 AV1 video was unwatchable with early releases of dav1d due to too many dropped frames.
Eventually dav1d got enough AVX optimizations to play the same video on the same hardware with zero dropped frames.
It was an impressive demonstration of what can be achieved when software makes the most of the available hardware.
I encode x265 on a $2.5/month VPS w/ 4gb ram during the off hours. AV1 is almost an order of magnitude slower.
All the major codecs have flags for you to balance speed against quality and compression, and it's up to you to pick the right tradeoffs for your use case.
And for most purposes, you want to use software encoders because they're much more flexible in terms of flags/options than hardware encoders. (Hardware encoders are usually optimized for speed rather than quality -- they're for live capture more then for video conversion.)
-vf scale=1280:720 -c:v libsvtav1 -crf 30 -preset 7 -c:a libopus -b:a 96k -ac 22. ffmpeg -i infile.mp4 -map 0:v:0 -pix_fmt yuv420p10le -f yuv4mpegpipe -strict -1 - | SvtAv1EncApp -i stdin --preset 6 --keyint 240 --input-depth 10 --crf 30 --rc 0 --passes 1 --film-grain 0 -b infile.ivf
3. ffmpeg -i infile.ivf -i infile.mp4 -map 0:v -map 1:a:0 -c copy outfile.mp4
On the Jetson Nanos I was lucky to get maybe 1fps in ffmpeg using VP9. Multiply that by six boards and that's about 6fps in total; ffmpeg running x264 in software mode was getting around 11fps on a single board, not even counting using the onboard encoder chip, meaning that I was getting better performance from one board using x264 than all six using VP9.
Now obviously this is a single anecdote on specific hardware, so I'm not saying that this applies to every single case, but it's a big reason why I personally have not used VP9 for anything substantial yet.
That said, AV1 is very obviously the future, and I'm perfectly happy with it taking over the market from h264, and I think that due to the bandwidth savings it's only a matter of time before all the major video services make it the default, especially as the speed of encoders increases to a useable level, which I'm sure it will soon enough.
[1] I know the most recent Raspberry Pi doesn't have a decoder chip for h264, but I think it's fast enough to do it in software.
They've recently contributed non-trivial patches to Firefox to use the embedded Linux API for video hardware acceleration (V4L2, vs. VAAPI on desktop that we also support), and are shipping the h264ify extension with their Firefox build to get that codec often for their users so that the experience is good on older devices.
Maybe the 5 is that much faster than it's not needed as much, but h264 represent so much content that it feels a bit surprising anyways.
But I'm just a software person, hardware is complicated differently.
If only! Then the patents would have expired. But H.264 is newer than MPEG-4 Part 2.
But you're right: H.264 has had the advantage of time, to gain fast hardware support.
I am pretty sure you are thinking of H.263 if it was from the 90s. H.264 barely started in the 00s.
In adaptive bitrate world, you split a video up into fragments, say 2-10 seconds large, and encode each segment in multiple bitrates, so that every say 5 seconds the video player can make a decision to download a different quality for the next 5 seconds.
Ok, but why not split the file up for standard encoding? Well, you can't just concatenate two .mp4 together without re-encoding and have it make sense to most media players (as far as I am aware), and moreover, it's inefficient from a RAM perspective when doing that. 1 second of RAW uncompressed 4k (24 fps) video is about 600MB. Source content for a single episode/movie at Netflix (I don't work there, just something I read once) can reach into the terabytes easily.
You can't just literally `cat foo-[123].mp4 > foo.mp4` with old-school non-fragmented .mp4 files, but you just have to shuffle the container stuff around a bit. You don't need to re-encode.
One downside is if you decide ahead of time that you're going to divide the video into fixed 5-second fragments/segments/chunks to encode independently, you're going to end up with that-length closed GOPs that don't match scene transitions or the like. IDR frame every 5 seconds. So no B/P frames that reference stuff 10 seconds ago, no periodic incremental refresh, nothing fancy.
>Ok, but why not split the file up for standard encoding?<snip>
at this point, you would be better served by just writing an elementary stream rather than a muxed mp4 file since it's just a segment anyways so why waste the cycles on muxing? you then absolutely 100% can concat those streams (even if you did mux them into a container). if you think you can't, you clearly have not tried very hard.
>I don't work there, just something I read once
I don't work there either, but do have 30+ years of experience with this subject. Sadly, you're not as well informed as you might think. People don't tend to encode to AV1 from RAW. They instead are dealing with a deliverable file most typically a ProRes in today's world after the post process has been completed. No where near terabytes for a feature. More closely to a couple hundred gigabytes for UHD HDR content. You seem to be unnecessarily exaggerating.
Edit: it's a 10x increase in encode speed, not time. that would be opposite effect.
I've never tried merging streams across computers so was naively just thinking that your output from each computer would be an MP4 but that makes sense.
I pulled that info. from a Netflix talk, perhaps video cameras back from when that talk occured didn't compress the video for you? Besides, isn't IMAX all intra-encoded? It was my understanding that IMAX films are actually just a series of J2K images, so I would imagine that the video cameras used there would also be intra-encoded.
i was thinking increased the encode speed 9x, but typed increased encode time. i also swapped the number of segments by segment duration. 9 segments of 10mins = 9x increase in performance.
Sounds like you are confusing Netflix' recommended formats for acquisition vs delivery. Cameras capture RAW formats (rarely is it uncompressed though), and the post houses use that as sources. The post house/color correction will the create the delivery formats which is typically ProRes. RAW is not a friendly format for distribution in the slightest. The full workflow from camera original to what ends up being streamed to the end viewer changes formats multiple times through the process.
In my opinion it's still in the early-adopter phase though, and it's perfectly valid to use tried-and-true codecs for user-interactive rendering and encoding cases, or where the existing codec meets your requirements for the compute vs disk/bandwidth trade-off.
We could have EVC which provide same H.264 ( or x264 ) encoding time but 30%+ Bit-Rate reduction. But market seems to want maximum BitRate reduction regardless of computational complexity.
However, I ran into issues with decoder speed too; whatever codec Kodi was using 18 months ago struggled to decode high-bitrate 1080p AV1 video last time I tried. Maybe this weekend I'll try again on the current version of Kodi.
If you're on Windows I recommend using StaxRip for encoding.
AV1 does not outperform H265 at high bitrates (and in certain cases, medium bitrates). What is considered "high bitrate" is dependent on source content, but a good rule of thumb is 40 MBit (think BluRay quality) or more for 4k content almost always goes to H265.
For AV1, depending on what you are encoding, and how much extra time you want to dedicate to experimentation, take a look at grain synthesis (you will want to test decode capabilities on that one).
This will only get cheaper over time as hardware accelerators get implemented and are improved YoY.
Yes the 1080p@30fps stream takes 2mbit/s in the worst case, no I don't give a damn about making it smaller. That is literally pointless. Yes in theory I could add eight times the CPU power to make it stream 4k@60fps and still end up with a lower bit rate but I don't care, because the CPU is the bottleneck and that many cores is incredibly expensive.
Then, as soon as I accept the cookie consent form, the page immediately starts to reload. I cancel the reload/navigation, find the close button for the chat widget, and finally start to scroll and read the content after 20 seconds of fidgeting with the page.
Immediately as I start scrolling, I realize the app bar header banner thing at the top of the page is actually being rendered with position: sticky; or some equivalent, and to top things off they’ve actually done me the favor of including some marketing banner ad about their “video developer report survey” being available, and they went ahead and attached that on top of their app bar, and also made it sticky. I was delighted to see that there was in fact no way to close that ad, as well.
Seriously— why do we make sites like this? The top 1/3rd of my screen is a useless banner and app bar that I have no intent to use whatsoever, the bottom 1/5th of my screen is that stupid chat widget, all I have is… let me do the math here… give or take half of my phone screen to actually read the content of the site.
/rant
That is by nature, I put emphasise on quality, because you could arguably disable all the AV1 features and lower the complexity to VP9 or h.264 level. But then you dont get the bitrate reduction benefits.
Whether that higher computing requirement is justify will be up to debate. And depending on your usage people will have different trade offs.
Why don't I see AV1 in many Youtube videos though? Checking with yt-dlp. It looks like they were planning to use it, but didn't really roll it out.
The client readiness is there. Already at 75% (relative to H264 100%, HEVC 15%).
It is thus clear which of the new codecs won client-side; it will only grow from here.
Definitely worth it.
Also interesting: https://news.ycombinator.com/item?id=23747923
https://streaminglearningcenter.com/codecs/codec-royalties-o...
It's just not worth it. Royalty-free formats (like AV1) are the way to go on the web.
- FullHD: http://compression.ru/video/codec_comparison/2022/main_repor...
- FullHD 10-bit: http://compression.ru/video/codec_comparison/2022/10_bit_rep...
- 4K: http://compression.ru/video/codec_comparison/2022/4k_report....
SVT-AV1 has seen a number of speed ups in recent releases:
This means AV1 won, client-side.
Now it is about time for the likes of netflix, youtube, twitch to deploy AV1 server-side.
And we can say goodbye to MPEG and patent encumbered codecs.
Good riddance.
I assume you mean to be included in things where things is either software or hardware? For software you get that for free from pretty much every single shipping PC, Smartphone or Tablet. So you really shouldn't implement it yourself. For hardware you still need to paid a tiny license fees ( which is why RPi 5 doesn't include it )
If you ignore extensions like SVC (Scalable Video Coding) and MVC (Multiview Video Coding) where most people dont use it. The normal H.264, especially baseline and Main Profile will have 99% of its patents expired within next 3-4 years.
With mp4 and TS it's almost trivial to get subsecond latency on a websocket, I hope AV1 has something like it.
Shame since Opus is smarter with bitrate allocation with surround sound.