AV1 Video Codec
aomedia.org
aomedia.org
https://aomedia.googlesource.com/aom/+/refs/tags/v1.0.0
https://github.com/AOMediaCodec/av1-spec/releases/tag/v1.0.0
Though obviously other encoders (and decoders) are available for the AV1 format. I think for individual encode enthusiasts the old libaom codebase is still the goto, but if you are Facebook or Netflix you probably want to be looking at SVT-AV1.
ffmpeg -i input.mp4 -c:a copy -c:v libsvtav1 -crf 17 svtav1_test.mp4
Hopefully using the SVT-AV1 encoder helps speed up your encoding.
[1] https://github.com/AOMediaCodec/SVT-AV1/blob/v1.0.0/Docs/Ffm...
Also, having more cores helps a lot with SVT-AV1.
For last, say, 25 years I've seen many parties trying to hamper that, even in non-commercial setting (e.g. the original mp3 format encoder). If they don't, good!
These closed formats are typically patent protected, not kept secret through proprietary encoding/decoding software. As you probably know: (1) Those patents will eventually expire, (2) The price of a patent is immediate and complete disclosure of the technical method underpinning the claims.
Look at LED lighting, it's more energy efficient, so now we deploy more of it, ending up using exactly as much energy as before (but now everything is well lit).
The same with video codecs, as they get better, we generate more video.
Though with LED lighting, it really depends. I still am using far less power for all of my lighting vs 1 incandescent light bulb. I am thinking specifically of the room/desk lamp of my childhood. Though I haven't counted the blinkenlights on my router, switch, etc. Maybe all of those added up will tip the balance.
No way those total over a watt.
H.265 isn't interested in being royalty-free. But, luckily, AV1 is very interested in being royalty-free. So this time around we don't need to repeat the mistakes and compromises of the past.
No offense, but HEVC vs AV1 is not in any way consequential to this trivial scenario. Those photos and videos will be decodable and re-encodable for you _forever_. Don't think you need to store all that stuff in some half-baked encoding just because it's theoretically free-er.
It will be possible forever, but not easy. JPEG is built into the standard library of every scripting language, and viewers come with every OS. I'm just lazy, that's all. I've tried "better" formats, but my life is easier when I stick with JPEG, MP4 (even though it's non-free), MP3, FLAC, etc.
Just yesterday I came across an old backup I made of some CDs 20 years ago... in Monkey's Audio. The files were 2% smaller, but they also reached out from the past and annoyed the hell out of me. haha
ffmpeg can be compiled with support for your hardware encoder.
Linux distributions and Firefox do provide H.264 support via openh264; it is sponsored by Cisco, who are already in the flat license territory, so all they have to do is track binary downloads (that's why they are separate download, or repository respectively). It also plugs seamlessly into linux multimedia framework (Gstreamer), so it is on the same level of integration, as Media Foundation codecs in Windows. There is no H.265 equivalent, still.
The hardware based encoding is not equivalent to software one; it is optimized for dumping compliant streams with low latency, it does not concern itself with effective use of the bits available. I.e. exactly what software based, offline encoders are good at.
[1] https://apps.microsoft.com/store/detail/hevc-video-extension... [2] https://apps.microsoft.com/store/detail/hevc-video-extension...
This isn't my experience with NVENC for h.264. Sure, software encoding can do somewhat (but not a lot) better (quality per bits), but at a fraction of the performance and using much more energy. For most cases were you're personally encoding video it's likely better to use a hardware encoder.
For H.264, hardware can brute-force every possible Intra encoded macroblock and make rate-distortion-optimised QP/prediction decisions in real-time or better. This gets much harder for H.265, where brute-force is fairly impractical and the clever algorithm wins.
A well-designed hardware codec can also run inter-frame block matching searches vastly more efficiently than software can.
I agree that I would prefer to use a patent-unencumbered codec, but I'm not willing to use AV1 until it has common hardware decoding support. My battery life and power consumption is much more important to me than a theoretical, unlikely, future inability to decode it.
Other techniques involve using copyright (the bitstream decode tables could be considered copyrightable for example).
Future techniques might even involve trying to use trademark law (for example, the first frame of any encoded video is the 'MPAA' logo in big bold letters, but without decoding that frame you can't decode the other frames). Trademarks don't expire.
At most that means you'd need to use old versions of software. An old publicly available file can't violate a patent issued 5 years later.
For the rest, Sega v. Accolade? And just putting a trademark inside the video wouldn't mean a decoder is using the trademark...
Now extrapolate to today: today it sounds ridiculous to use XviD and even H.264. For ease of use I’d use x265-10 bit, for future proofing I would need to read up on av1. Think what it will look like in 10 years (2032) as you will have those files in 10 years for sure.
For long term archive work the above or AV1 (if you have infinite time / energy budget) are probably better, depending on settings.
The reason for this is the DCT transform results in 16-bit numbers for every pixel. The reason that doesn't make things worse is (partly) because it only sends the non zero values.
There's too many more reasons why, but that's the quick summary.
I haven't seen this argued as 16-bit DCT, but in color space conversion. The gist is that all 8-bit RGB values cannot be represented properly in 8-bit YUV420, so you're supposed to use 10-bit to get "proper" YUV values. But if you start with an 8-bit encode you've already thrown away the extra precision, so why waste the (considerable) extra compute on 10-bit just to make sure you don't truncate the already-truncated YUV?
I have a project in progress to measure all of the variations, but from quick testing with CRF encoding the same value results in much longer compute AND a larger file in 10-bit versus 8-bit. The larger file has a slightly higher VMAF score, as would be expected from spending more bits. The work is in finding a set of encoding parameters to measure the quality difference at the same output size, and to measure the relative improvement across CRF vs size vs bit depth.
The process is: Raw input pixels (8 or 10 bit) minus predicted pixels (8 or 10 bit) -> residual pixels (8 or 10 bit + 1 sign bit).
You take these residual pixels and pass them through a 2D DCT, then scale and quantise them. At the end of this, the quantised DCT residual values are signed 16-bit numbers - you don't get to choose the bit-depth here; it's part of the standard (section 8.6). For every 16x16 pixel input, you get a 16x16 array of signed 16-bit numbers.
The last step is to pass all non-zero quantised DCT residual values through an entropy coder (usually an arithmetic coder), then you get the final bitstream.
The key point is that it didn't matter if the original raw pixel input was 8-bit or 10-bit; the quantised DCT residual values became 16 bits before being compressed and transmitted. This is also true for 12-bit raw pixel inputs.
This seems impossible; for 8-bit inputs, you've doubled the size of the data (slightly less than double for 10-bits), so you must be making things worse! The key is that after scaling and quantisation, most of those 16-bit words are zero. Those that are non-zero are statistically closer to zero so that the entropy encoder won't have to spend a lot of bits signalling them.
The last part comes when you reverse this process. The mathematical losses from scaling and quantising 10-bit inputs into the transmitted 16-bit values are less than the losses for 8-bit inputs. When you run the inverse quant, scale and iDCT, you end up with values that are closer to the original residual values at 10-bit than you do at 8-bit.
https://meta.wikimedia.org/wiki/Have_the_patents_for_MPEG-4_...
https://openbenchmarking.org/test/pts/dav1d#results
Edit: It’s false! This is just a decoder, not an encoder! I just learned it.
Here is the link for the performance benchmarks for SVT-AV1, the most popular encoder as I just learned:
Back in those days, x264 felt to me like a software written by aliens from the future.
I don't know of any benchmarks, but rav1e does have --tune psychovisual, and there are issues raised against it, so it seems they take it seriously.
> Back in those days, x264 felt to me like a software written by aliens from the future.
Indeed it may be :) H.264 might not be the latest and best video coding standard anymore, but in my opinion x264 is, and will always be by far the best encoder ever written for a video format.
Which brings me to...
> Because back in the days, both Videolan projects x264 and x265 (especially x264) had much better psychovisual quality than commercial encoders.
Unfortunately x265 isn't a Videolan project (its developed by MulticoreWare Inc.) and it's a very mediocre encoder which doesn't hold a candle to x264 and IMO kind of a shame considering its legacy.
Also after 2018 it's practically became maintainenance-only and was surpassed by proprietary encoders in MSU encoder tests in the following years. That's a big loss considering x264 was still seeing significant efficiency and performance improvements as late as 2013 (when the H.264 format was 10 years old), so when compared to x264 I assume a good 4-5 years of potential improvements have been left at the table for x265.
I think the big issue is x264 was very obviously a labor of love from some very talented developers. I just haven't seen that sort of love dumped into other encoders.
Newer codecs have been relying on the format to provide more obvious tools for compression (and mostly giving benefits for HD+ resolutions).
Up until the end of development, x264 was hyper focused on getting the best possible subjective quality with the smallest possible bitrate. To date, the x264 CRF metrics are (IMO) unparalleled in consistency. With other codecs a similar CRF mode is simply, well, shit. I can't just set stuff to "CRF 20" and expect the output to hit roughly the same level of quality. VP9, in particular, is terrible with this. In VP9 CRF is more closely related to the bitrate than the actual quality of the scenes being encoded.
To be clear, even with these critiques you SHOULD choose x265, vp9, or AV1 over x264 for your encoding choices. They have better specs that allow for better compression. However, they are also leaving a lot on the table for what they COULD do.
I current do VP9 + vmaf on each scene to set a CRF value (using my own thing similar to AV1AN). That gives good consistent results at minimal bitrates. It's just a little terrible (IMO) that I have to do so much work that the encoder should theoretically be able to do better.
[1] https://web.archive.org/web/20100105000031/http://x264dev.mu...
What are up-to-date AV1 encoders still leaving on the table as far as optimization is concerned?
You are not mistaken. The difference is in how reliable the control is regardless of input video.
For libvpx, the CRF control is garbage. A CRF of 30 will be good for some scenes and horrible for scenes that are too dark or have too much motion. It means if you want to just use libvpx (or ffmpeg), you are often setting that CRF way lower than you need to so scenes where it fails don't end up looking like smooth color blobs. It's bad enough that they introduced a "minimum bitrate" flag.
x264 is not that experience. The amount of adjustment you have to do for CRF for a given input are extremely minor, I found between 20 and 24 to be more than acceptable. For vpx, you need to come up with a value anywhere from 10 to 50 depending on the source.
I get that a lot of this is subjective experience, but it's what I've experienced doing a bunch of dvd rips.
> What are up-to-date AV1 encoders still leaving on the table as far as optimization is concerned?
The biggest seems to be good quality controls that have been tuned by someone with a good subjective eye for that sort of thing. Beyond that, IDK, the bitstreams allow for a LOT more transformations than H.264 allowed for, yet the codecs don't seem to have the same level of complexity. For example, x264 came up with a bunch of motion vector search patterns over it's evolution. You don't see those sorts of developments with the other encoders.
Heck, you even saw that sort of care for quality output in the fact that x264 has tuning guides for (at the time) common objective measures of quality, SSIM and PSNR. (which returned worst quality than the x264 subjective quality metrics.
IDK, this may also be that I don't have as much time to geek out over video codecs :).
For other encoders that do not do parallel encoding that well there are things like av1an.
Rough example: `ffmpeg -i video.mkv -c:v libaom-av1 -cpu-used 5 av1_test.mkv`
I used ffmpeg "-cpu-used 8" for AV1 (higher than this and I get "Error setting option cpu-used to value X"). I removed the -cpu-used command for H.264 as the encoder defaults to autodetecting the number of threads to max out CPU usage. My CPU is an 8-core 16-thread AMD Ryzen 7 PRO 5750G. My source material was the first 20 seconds of the H.264 blu-ray rip of Titanic. top(1) shows the AV1 encoder uses around 500% CPU (meaning only 5 of the 16 hardware threads are utilized) while the H.264 encoder uses around 1600% (exactly 16 of 16 threads utilized). So there is a lot of potential parallelization optimization that could be exploited. But even assuming perfect scaling from 5 to 16 threads, the AV1 encoder would get to 1.1x of realtime playback speed. It would still be about 5 times slower than the H.264 encoder.
https://github.com/AOMediaCodec/community/wiki#how-to-make-e... explains the different settings that most change the speed of encoding.
I don't think AV1 will ever get to the encoding speed of a codec like h264, because in a very general sense; simpler math is easier to do. AV1 can encode the same information into fewer bits, but that efficiency has a computational cost.
It really depends on what you're doing. If you're trying to livestream it could be a big issue. If you're going to encode a video once and then store it for years, it's probably not a big deal. Totally up to you. For me AV1 has become a pretty good option.
For livestreaming, SVT-AV1 1.0 is now usable on good CPUs (eg. AMD 5800x+ desktop CPUs, Intel 12600+ desktop CPUs) at higher CPU presets (8+, depending on your CPU and what you're streaming), just currently no one allows for AV1 ingest for livestreams.
For example, a 2Mbps video stream will usually have a 4Mb video cache that can be used to preload data. Having this video cache means that a 2Mbps video stream can burst up to 6Mbps for one second while still maintaining a 2Mbps cap.
I retried with "-row-mt 1 -tiles 2x2" (keeping "-cpu-used 8" so the benchmark is comparable to my previous test): the encoding speed is 0.443x of playback speed; top(1) shows about 8 of my 16 cpu threads are utilized.
Without "-row-mt 1 -tiles 2x2" the encoding speed was 0.329x. So these options only increases speed by 34%. This doesn't match the increased in cpu utilization of +60% (5 to 8 threads). Contention on shared data structures? Looks like it's better to just spawn multiple ffmpeg instances working on different source files instead of leveraging the encoder's multi-threading. That way I could get close to 1x of playback speed.
I have 3500 hours of video content. At 1x I need 5 months to reencode all. Heavy. But doable I guess.
If we're going to do AV1 vs. H.264, and complain that AV1 is much slower, we might as well compare XviD and H.264 and complain about the latter.
AV1 is one generation newer. It was made by merging the projects for VP10, Thor, and Daala.
AV1 is computationally better than HEVC/VP9 and more comparable to H.264. Visually it’s better than either.
Comparation between codecs and encoder presets. At the same speed SVT-AV1 gives better quality vs h264, h265. M10-M8 look like nice spot
Newer codecs can generally take better advantage of that parallelism and SVT in particular has this as a core design element.
Still cool, but might not fully apply depending on your use case.
The trick is if your encoder can adapt the quantizer based off a metric to compare the encode quality you can save a ton more space without visual degradation.
AV1 is ~30% more efficient than h.265. And h.265 is ~40% more efficient than h.264. The lower the resolution of the video, the less savings you will get.
I personally will wait until more devices support hardware AV1 decoding. I believe only Intel 12th gen, and Nvidia 3000 support AV1 currently. I don't plan on purchasing a new laptop or smartphone for 3-4 years, at that point I may start encoding in AV1.
The take-away is that nVidia only has hardware support for AV1 on the RTX 3000 Ampere series cards.
Intel supported AV-1 decoding in its Xe-LP GPUs in 2020.
Intel also reached v1.0.0 of its open-source codec 2 weeks ago, here: https://github.com/AOMediaCodec/SVT-AV1
I am not aware of any other CPU or GPU that has an AV-1 codec built in.
JPEG XL looks pretty good, too, but current browser support is a flag in Chrome. I wonder if they can squeeze better quality out of low filesizes (0.3~0.6 bits per pixel) to compete better with AVIF.
The introduction appears to only be a legal disclaimer?
H264 and h265 are patent encumbered, so depending what you are doing you need to pay royalties to use them. You are free to use AV1 with paying any royalties. Also, AV1 can deliver better quality video at the same bitrate compared to h264/h265 (at the expense of encoding time). This makes it useful for things like netflix where they can put lots of effort into encoding something once and then reap the bandwidth reduction reward for each person who watches it.
Same as AV1 [1].
> use AV1 with paying any royalties
Same as H.264/5 for almost all normal usage.
https://www.streamingmediaglobal.com/Articles/ReadArticle.as...
It doesn't really matter though, because AOmedia's patent license here:
https://aomedia.googlesource.com/aom/+/refs/heads/main/PATEN...
has a defensive clause:
> Defensive Termination. If any Licensee, its Affiliates, or its agents initiates patent litigation or files, maintains, or voluntarily participates in a lawsuit against another entity or any person asserting that any Implementation infringes Necessary Claims, any patent licenses granted under this License directly to the Licensee are immediately terminated as of the date of the initiation of action
So if you sue anybody for patents you essentially lose access to all of the relevant patents by any of those companies:
https://aomedia.org/membership/members/
So please stop spreading baseless FUD.
However, make a competitor of YouTube, and I wish you good luck not paying those fees.
With regards to "from people who had nothing to do with development", there are many research centers, not patent trolls, who actually published publicly their researches. Maybe the people who did AV1 didn't read the literature, but I'll presume they did.
They don't indemnify you from patent claims.
They merely offer you the ability to use their patents to help defend yourself. But against the likes of NTT, Orange, Phillips etc patents aren't the issue, it's running out of money to litigate the issue.
I think it's worth tempering the language here.
These aren't patent trolls but companies who have been involved in codec design for decades and will be involved in ones in the future e.g. NTT, Dolby, Toshiba.
> So please stop spreading baseless FUD.
No one is spreading FUD. You have outlined a worthless termination clause given that it only applies to those who are interested in implementing the codec. That offers zero protection against third party claims like an indemnification would.
And that's assuming validity in the first place.
Normally I would agree with you, but not in this case. Software patents need to die, and anyone who tries to extract rent through them *is* a parasite. Especially if they employ mafia-like tactics like those patent pools. Oh, quite a nice codec you have there; it'd be a shame if anything happened to it.
Nowadays you can't even fart without infringing someone else's patent. So, sorry, but I'm not going to apologize for calling them parasites. This is one of very few strong opinions I have, and is a hill I'll gladly die on.
In terms of compression efficiency it's on par with H.265, but it's free to use.
My usual playback targets are almost exclusively a web browser or Chromecast, so storing media in formats that would allow for more direct play opportunities would be fantastic.
Almost every source you can get your hands on will be lossy compressed - unless you generate your own - but even then, most consumer devices don't even allow recording into lossless formats.
Uncompressed video is mind bogglingly huge.
According to its license which is very short and can be read here:
https://code.videolan.org/videolan/dav1d/-/blob/master/doc/P...
https://www.sisvel.com/blog/audio-video-coding-decoding/sisv...
It became a real legal issue before the first consumer AV1 hardware was ever launched.
and it's backed by these companies: https://aomedia.org/membership/members/
Yes, in the sense that some AV1 codec implementations will be faster than some HEVC/H.265 implementations.