Dav1d, fast AV1 decoder, version 1.0
code.videolan.org
code.videolan.org
What you should really look at on dav1d is the fact that the code has now around 186 kLoC of hand-written assembly in .S and .asm files…
I think this is quite a feat (this is more asm than the whole FFmpeg) and this is very rare those days to write so much asm.
This reads to me as it being a PITA to maintain. Cross platform code is usually a pain, cross platform with assembly optimizations is more of a pain. Optimization nearly always makes things harder to maintain, and this sounds like it was optimized to hell and back...
Not more than any other language: when it's well done, it's manageable. When it's spaghetti code, it's not manageable.
People don't really get how complex video codecs are.
That's for x86_64(SSE3, AVX2), ARMv7 and ARMv8 right? Is the 186kLoC equally distributed between those arch or is there one arch that has most of the optimization focus?
Is it all hand-written, or do you have some kind of script/macros to create a good part of this assembly (like OpenSSL does IIRC)?
Hand written with some macros, notably for the x86 mess (windows calling convention, ssse being 32 and 64b).
One of the bugs found by OSS-Fuzz: https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=11464
There's cpu = 'i686' in the config file which would in fact seem to imply none of the modern archs mentioned in a different subthread here ("x86_64(SSE3, AVX2), ARMv7 and ARMv8") are fuzzed.
but having a look at clang or GCC, I can confirm. still baby steps, years ahead for proper optims.
An AV1 example at 1080p50: https://www.youtube.com/watch?v=YhXtJWi2PjI
You can verify if YouTube videos are playing back in AV1 by right clicking on the video and choosing Stats for Nerds. The codec is av01 for AV1 video.
I know so little about this space so not sure how helpful the comparison is, but it is like magic what we are able to do with math.
I'm using Chrome on macOS, which seemingly supports this as well: Codecs av01.0.08M.08 (398) / opus (251)
What decoder is Chrome using?
I didn't see benchmarks mentioned in TFA, please let me know if I've missed or overlooked something.
How are you testing? Is that dav1d 1.0? 8bit or 10bit content?
If you tested before 1.0 and 10bit content, you should run it again, since we pushed new optimizations for this.
I should have included in the original comment this machines needs 60% to 70% CPU to play VP9 on similar bitrate / bit / frames. i.e Decoding on VP9 and AV1 are on a similar level of CPU usage. Considering the complexity of AV1, I thought this was an incredible achievement. ( It actually blows my mind this can be done )
I tested it on with 8 and 10 bit content the result aren't that much different. ( This is just eyeballing CPU usage in Activity Monitor ) The 3.3 Ghz was turbo, so it really should be 2.9Ghz Dual Core Broadwell CPU.
It's good and getting gooder.
Just give it a try using https://github.com/Alkl58/NotEnoughAV1Encodes or whatever, svt-6 is fairly quick.
I just transcoded some movies with vmaf of 98% from 4k hdr blurays this last week. Takes about two days of mostly single core work with cpu-used=4 on a 5950x but they get average bitrates of around 4400 kbps. Which is like, really, really good for hitting that quality target.
The trick is that since its mostly single core I just do 16 movies at a time, or 24 on my server.
Turns out it's $NFLXs OSS automated encoding quality assessment tool. Pretty cool!
https://github.com/Netflix/vmaf
p.s. zanny, nice casual stealth drop you did there ;)
If Netflix wants to hire me, dm me, lol.
We should be seeing better AVX-512 support with CPUs in the coming years though.
It’s physically disabled in new CPUs
https://www.tomshardware.com/news/intel-nukes-alder-lake-avx...
(disclaimer: I used to work there)
Even the newest Shield Pro is based on a 2015 SoC (Tegra-X1) with an integrated 2014 Nvidia GPU (Maxwell). Not surprising it doesnt support av1.
The danger is this block only implementing the decoding of a non-royalty free/patent encumbered format à la mpeg, presuming the "intellectual property" was not globally fixed and still toxic like nowdays.
SMPTE suggest that 4k for viewing is only worthwhile for people with 20/20 vision once the viewing angle increases beyond 30 degrees.
I'm currently looking at a 40cm wide screen, I'm 80cm away from it, which is about 30 degrees.
For my phone to fill that amount of space it would have to be about 10cm in front of my eyes due to binocular vision (if I just use one eye it's a bit further
When I watch something on my phone it's typically 50cm away, about 15-20 degrees viewing. 720p is more than enough at that size, 1080p if you have particularly good eyesight.
I do not know what phone you have but an iPhone 12 ( as an example ) is under 1300 pixels on the short dimension. So you are not getting much more than 1080 pixels no matter what. It seems any experiential difference would have more to do with compression quality than with resolution.
Speaking for myself, I do not think I could tell the difference between 4K and 1080p on a phone ( on a decent AV1 clip ).
But it does mean that there are reasons to want 4K playback on a phone.
The far more interesting question is if you accept you have Xmbit to play with, what is better on a given platform (screen size, resolution, viewing situation, how well compression works, how much battery is used in decoding, etc)
To a certain extent it's just a place holder so folks can start writing software against it so that when the CPU vendors ship everyone can hit the ground running.
This is the description the dav1d page gives:
dav1d is the fastest AV1 decoder on all platforms :)
Targeted to be small, portable and very fast.
Since they use the word fast twice I was wondering how fast.
Which is cool, but if you have hardware decode then you may not care, though software decode gets used in all sorts of weird places that you might not expect.
I dont think it was one single slide I was thinking of but you can for example find comparisons of vp8, vp9, h264 by the same videolan/ffmpeg devs that demonstrate this general principle, the decode fps is surprisingly flat between software decodes of different codec families and generations.
The earlier generations often hit bottlenecks though, I think ffmpeg VP9 decode (built by the same people as dav1d) was the leader for a while but hits a wall at 4 cores due to format limitations. HEVC and especially dav1d have some gains when you can throw a modern amount of cores at it, and even phones often have 8.
I like open and free standards, and I dislike royalties, but the political aspect is best foregone when one of the alternatives really is objectively so much better for everyone.
https://youtube-eng.googleblog.com/2015/04/vp9-faster-better...
https://engineering.fb.com/2018/04/10/video-engineering/av1-...
https://netflixtechblog.com/netflix-now-streaming-av1-on-and...
https://medium.com/netflix-techblog/bringing-av1-streaming-t...
https://bitmovin.com/bitmovin-improves-av1-video-encoding/
https://blog.webex.com/video-conferencing/cisco-leap-frogs-h...
https://blog.webex.com/engineering/the-av1-video-codec-comes...
Your problem is fundamentally an emotional one. Now that you suspect Device X will perform poorly at Task Y you feel buyer's remorse.
Try not to worry about it so much. Computer hardware will continue to be made obsolete by more demanding software for quite some time yet.
Of course there are. This is typical.
> I consider the MPEG family as objectively better on the technical merit of GPU/VPU hardware decoding being available everywhere
Then there will never be any codec development. It's a silly position to take.
Hardware decoding is available. And dav1d is a very fast software decoder. I don't have AV1 hardware and yet I play back AV1 on YouTube just fine via dav1d.
> I'm merely underlining that the MPEG family is the state of the art.
It isn't.
AV1 is at least as widely supported as VVC.
Nothing supports hardware decoding if mpeg EVC, and hardware decoding if the original mpeg 4 (xvid generation) was never a big deal.
Core 2 Duo are processors from 16 years ago...
Mobile phones are more powerful than those machines...
H265 or VP9 in software are the same… Here, we managed to have a sw decoder for a format that is one order of magnitude more complex than those codecs and yet, consume less CPU.
No. I understand your vitriol, but please just stop. I'm not demeriting AV1's technical basis.
Depends on SoC. A15? or low end UniSoC SoC? But generally speaking a modern Smartphone should be able to play 1080P 25fps 1mbps AV1 video on their phone without dropping frames. You will just be burning away your battery life.
Eventually computing overcomes the need for optimized assembly and understandable code become more desired.
In modern codebases his has the effect of micro-optimization of the "rewrite in assembly" sort mattering a great deal in heavy lifting library code like codecs, CPU emulation, rasterizers, VMs etc, and not at all in other areas. Emulation is a decent candidate in this regard but it also massively benefits from higher level simplifying assumptions(e.g. not polling for I/O changes at the true device frequency). It can be hard to notice differences in emulation quality where high level techniques are applied well.
On old 8-bit architectures it was rather the opposite: you couldn't be high level about anything because you had to design towards micro-efficiency. Therefore most of your problems were solved with very simple data structures and algorithms that heavily favored either space(by recomputing the intermediate results) or time (by precomputing all answers in a LUT), and then given a very tight cycle-shaving, code-golfing implementation.
It could be that they are waiting to see if AV1 will actually get a marketshare outside Youtube or not.
...to bad that they (in the usual Apple fashion) don't support it in a standard container which everyone else supports and require you to package it in their special snowflake .caf container, so even though in theory everyone supports Opus you still need to use MP3 if you don't want to deal with having to support multiple formats/codecs at the same time.