Rough example: `ffmpeg -i video.mkv -c:v libaom-av1 -cpu-used 5 av1_test.mkv`
I used ffmpeg "-cpu-used 8" for AV1 (higher than this and I get "Error setting option cpu-used to value X"). I removed the -cpu-used command for H.264 as the encoder defaults to autodetecting the number of threads to max out CPU usage. My CPU is an 8-core 16-thread AMD Ryzen 7 PRO 5750G. My source material was the first 20 seconds of the H.264 blu-ray rip of Titanic. top(1) shows the AV1 encoder uses around 500% CPU (meaning only 5 of the 16 hardware threads are utilized) while the H.264 encoder uses around 1600% (exactly 16 of 16 threads utilized). So there is a lot of potential parallelization optimization that could be exploited. But even assuming perfect scaling from 5 to 16 threads, the AV1 encoder would get to 1.1x of realtime playback speed. It would still be about 5 times slower than the H.264 encoder.
https://github.com/AOMediaCodec/community/wiki#how-to-make-e... explains the different settings that most change the speed of encoding.
I don't think AV1 will ever get to the encoding speed of a codec like h264, because in a very general sense; simpler math is easier to do. AV1 can encode the same information into fewer bits, but that efficiency has a computational cost.
It really depends on what you're doing. If you're trying to livestream it could be a big issue. If you're going to encode a video once and then store it for years, it's probably not a big deal. Totally up to you. For me AV1 has become a pretty good option.
For livestreaming, SVT-AV1 1.0 is now usable on good CPUs (eg. AMD 5800x+ desktop CPUs, Intel 12600+ desktop CPUs) at higher CPU presets (8+, depending on your CPU and what you're streaming), just currently no one allows for AV1 ingest for livestreams.
For example, a 2Mbps video stream will usually have a 4Mb video cache that can be used to preload data. Having this video cache means that a 2Mbps video stream can burst up to 6Mbps for one second while still maintaining a 2Mbps cap.
I retried with "-row-mt 1 -tiles 2x2" (keeping "-cpu-used 8" so the benchmark is comparable to my previous test): the encoding speed is 0.443x of playback speed; top(1) shows about 8 of my 16 cpu threads are utilized.
Without "-row-mt 1 -tiles 2x2" the encoding speed was 0.329x. So these options only increases speed by 34%. This doesn't match the increased in cpu utilization of +60% (5 to 8 threads). Contention on shared data structures? Looks like it's better to just spawn multiple ffmpeg instances working on different source files instead of leveraging the encoder's multi-threading. That way I could get close to 1x of playback speed.
I have 3500 hours of video content. At 1x I need 5 months to reencode all. Heavy. But doable I guess.
If we're going to do AV1 vs. H.264, and complain that AV1 is much slower, we might as well compare XviD and H.264 and complain about the latter.
AV1 is one generation newer. It was made by merging the projects for VP10, Thor, and Daala.
AV1 is computationally better than HEVC/VP9 and more comparable to H.264. Visually it’s better than either.
The take-away is that nVidia only has hardware support for AV1 on the RTX 3000 Ampere series cards.
Intel supported AV-1 decoding in its Xe-LP GPUs in 2020.
Intel also reached v1.0.0 of its open-source codec 2 weeks ago, here: https://github.com/AOMediaCodec/SVT-AV1
I am not aware of any other CPU or GPU that has an AV-1 codec built in.
Comparation between codecs and encoder presets. At the same speed SVT-AV1 gives better quality vs h264, h265. M10-M8 look like nice spot
Newer codecs can generally take better advantage of that parallelism and SVT in particular has this as a core design element.
Still cool, but might not fully apply depending on your use case.
The trick is if your encoder can adapt the quantizer based off a metric to compare the encode quality you can save a ton more space without visual degradation.
AV1 is ~30% more efficient than h.265. And h.265 is ~40% more efficient than h.264. The lower the resolution of the video, the less savings you will get.
I personally will wait until more devices support hardware AV1 decoding. I believe only Intel 12th gen, and Nvidia 3000 support AV1 currently. I don't plan on purchasing a new laptop or smartphone for 3-4 years, at that point I may start encoding in AV1.
https://openbenchmarking.org/test/pts/dav1d#results
Edit: It’s false! This is just a decoder, not an encoder! I just learned it.
Here is the link for the performance benchmarks for SVT-AV1, the most popular encoder as I just learned:
Back in those days, x264 felt to me like a software written by aliens from the future.
I don't know of any benchmarks, but rav1e does have --tune psychovisual, and there are issues raised against it, so it seems they take it seriously.
> Back in those days, x264 felt to me like a software written by aliens from the future.
Indeed it may be :) H.264 might not be the latest and best video coding standard anymore, but in my opinion x264 is, and will always be by far the best encoder ever written for a video format.
Which brings me to...
> Because back in the days, both Videolan projects x264 and x265 (especially x264) had much better psychovisual quality than commercial encoders.
Unfortunately x265 isn't a Videolan project (its developed by MulticoreWare Inc.) and it's a very mediocre encoder which doesn't hold a candle to x264 and IMO kind of a shame considering its legacy.
Also after 2018 it's practically became maintainenance-only and was surpassed by proprietary encoders in MSU encoder tests in the following years. That's a big loss considering x264 was still seeing significant efficiency and performance improvements as late as 2013 (when the H.264 format was 10 years old), so when compared to x264 I assume a good 4-5 years of potential improvements have been left at the table for x265.
I think the big issue is x264 was very obviously a labor of love from some very talented developers. I just haven't seen that sort of love dumped into other encoders.
Newer codecs have been relying on the format to provide more obvious tools for compression (and mostly giving benefits for HD+ resolutions).
Up until the end of development, x264 was hyper focused on getting the best possible subjective quality with the smallest possible bitrate. To date, the x264 CRF metrics are (IMO) unparalleled in consistency. With other codecs a similar CRF mode is simply, well, shit. I can't just set stuff to "CRF 20" and expect the output to hit roughly the same level of quality. VP9, in particular, is terrible with this. In VP9 CRF is more closely related to the bitrate than the actual quality of the scenes being encoded.
To be clear, even with these critiques you SHOULD choose x265, vp9, or AV1 over x264 for your encoding choices. They have better specs that allow for better compression. However, they are also leaving a lot on the table for what they COULD do.
I current do VP9 + vmaf on each scene to set a CRF value (using my own thing similar to AV1AN). That gives good consistent results at minimal bitrates. It's just a little terrible (IMO) that I have to do so much work that the encoder should theoretically be able to do better.
[1] https://web.archive.org/web/20100105000031/http://x264dev.mu...
What are up-to-date AV1 encoders still leaving on the table as far as optimization is concerned?
You are not mistaken. The difference is in how reliable the control is regardless of input video.
For libvpx, the CRF control is garbage. A CRF of 30 will be good for some scenes and horrible for scenes that are too dark or have too much motion. It means if you want to just use libvpx (or ffmpeg), you are often setting that CRF way lower than you need to so scenes where it fails don't end up looking like smooth color blobs. It's bad enough that they introduced a "minimum bitrate" flag.
x264 is not that experience. The amount of adjustment you have to do for CRF for a given input are extremely minor, I found between 20 and 24 to be more than acceptable. For vpx, you need to come up with a value anywhere from 10 to 50 depending on the source.
I get that a lot of this is subjective experience, but it's what I've experienced doing a bunch of dvd rips.
> What are up-to-date AV1 encoders still leaving on the table as far as optimization is concerned?
The biggest seems to be good quality controls that have been tuned by someone with a good subjective eye for that sort of thing. Beyond that, IDK, the bitstreams allow for a LOT more transformations than H.264 allowed for, yet the codecs don't seem to have the same level of complexity. For example, x264 came up with a bunch of motion vector search patterns over it's evolution. You don't see those sorts of developments with the other encoders.
Heck, you even saw that sort of care for quality output in the fact that x264 has tuning guides for (at the time) common objective measures of quality, SSIM and PSNR. (which returned worst quality than the x264 subjective quality metrics.
IDK, this may also be that I don't have as much time to geek out over video codecs :).
For other encoders that do not do parallel encoding that well there are things like av1an.