Rav1e, an AV1 encoder written in Rust and assembly
github.com
github.com
* Open & Royalty Free (no licensing required to use)
* Backed by Mozilla, Google, Microsoft, Cisco, etc. [1]
* Approximately 50% higher compression than x264 [1] and 30% higher than H.265 (HEVC) [2]
* Supported in current versions of Chrome, Firefox, Opera, and Edge [3]
* Slower encoders than HEVC, so not typically used for live streaming [2]
[1]: https://en.wikipedia.org/wiki/AV1
[2]: https://www.theoplayer.com/blog/av1-hevc-comparative-look-vi...
I was unable to find any evidence of this anywhere I looked. For instance this webpage: https://www.webmproject.org/docs/container/ shows only VP8 and VP9 as acceptable video codecs.
Any idea why they flipped the switch on that? Seems irresponsible when .mkv will do in all circumstances.
Is there any examples of major services currently using HEVC?
Given Apple is an AOM governing member it's pretty likely Apple will eventually support AV1 (and AVIF).
VP9 is simply not on their roadmap at all, it's not a "yet" thing.
We can reasonably assume that will change, as basically every major player in the CPU/GPU spaces[1] has signed on to back this. Even Apple is a member.
[1] the exceptions are mostly downstream ARM vendors - notably, Qualcomm
This is a very misleading statement.
AV1 encoders are slow because they're slow, not because they're not hardware accelerated.
x264 and x265 are not hardware accelerated, either.
This is incredibly false.
https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video
https://en.wikipedia.org/wiki/Nvidia_NVENC
https://en.wikipedia.org/wiki/Video_Coding_Engine (this has been widely available for 8 years)
x264/x265 are software encoding libraries for H264/H.265 video formats and don't really use hardware acceleration beyond SIMD instructions on CPUs. x264 is known for its great output quality while still being fast on x86 CPUs.
Huh, TIL.
Though I would also sort of want to know where x264 and x265 were a year in, tbf.
What is the relative improvement between codec generations (e.g. H.264 vs. H.265) compared to implementations of the same codec standard (e.g. early reference implementation of H.264 vs. current x264 with well-chosen parameters)?
Most rust projects have these sort of safety claims to attract companies, devs to promote it in meetups or tech conferences. Maybe they should audit this encoder to see if it lives up to these claims.
https://github.com/xiph/rav1e/search?q=unsafe&unscoped_q=uns...
This is also often true for cryptographic software, particularly fixed-parameter optimized code.
The performance gain usually comes from some mixture of good scheduling or register allocation that the compiler fails at and the use of instructions that the compiler won't (reliably) emit.
For these kinds of straight-line codes, I haven't found mechanically validating them to be significantly more complicated than verifying C code... and as C code these functions tend to be the lowest risk type.
So, without any direct experience, I'd expect there to be more risk in rav1e from the unsafe-rust than the asm.
That said, already the 'attack surface' of an encoder is pretty small. I'd personally expect rav1e's gain from rust's safety would be less security and more in avoiding wasting time on blind alleys caused by memory corruption in the codebase. I've seen more than one poor design decision made in a multimedia codec which was ultimately due to a bug that a better language might have prevented.
By focusing on an encoder (at least initially) you can just support a subset of features that work well for you. Then focus on making a decoder that can at least play back videos from your encoder.
More people will find it useful to have an encoder that works all the time, rather than a decoder that only works on a subset of videos.
> ~70% of the vulnerabilities Microsoft assigns a CVE each year continue to be memory safety issues
https://msrc-blog.microsoft.com/2019/07/16/a-proactive-appro...
There are similar reports from other organizations.
For any reasonable definition of "safer", yes. 100% safe? no.
Might it be possible to compile this to wasm and decode it that way on the client?
I could have sworn they used to though. You’d get an app called “Webkit”, but it functioned just like Safari for all intents and purposes.
[1]: https://www.theverge.com/2018/1/4/16805216/google-chrome-onl...
[2]: https://www.computerworld.com/article/3199425/top-web-browse...
[3]: https://erik.itland.no/chrome-is-the-new-internet-explorer-4...
It would be nice to finally have a codec again that works with every (modern) platform after such a long time. The split of support for HEVC and VP9 always doubles the effort to distribute content effectively.
If Netflix or YouTube want AV1 on Apple TV then they'll just implement it in their applications. Netflix has already implemented AV1 as an option in their Android application:
https://netflixtechblog.com/netflix-now-streaming-av1-on-and...
It's still a bit early to get upset about other vendors, when Google's own Pixel line doesn't have hardware decode yet.
http://jabberfr.org/tmp/ogv.js/
Note this works fine in desktop and mobile Safari.
It also doesn't support VP9, so don't hold your breath.
> Might it be possible to compile this to wasm and decode it that way on the client?
Maybe, but I doubt the performance would be tolerable.
There are some cloud providers that offer this as a service, allowing you to have the end-points of, e.g., video calls to use different audio and video codecs.
This is quite useful, e.g., when some people join the call only using audio via a cell phone in a different country using a different audio standard (or a land line, etc.). Or when somebody joins the video call from laptop tethering from a phone on a train. Or for switching video codecs depending on whether somebody is sharing their desktop or using a webcam to record their face.
The client can picks the codecs that are the best fit for the current situation (content, bandwidth, latency, etc.) and a could server transcodes the video from everyone else in the meeting to their clients format.
However most consumer/server GPUs include hardware IP blocks specifically for doing codec work.
The codecs that NVEnc supports are just a few bunch, there is no real-time audio, no nothing.
All of this is implemented in CUDA, using normal CUDA implementations of audio and video codecs, and running on normal GPUs using normal GPU cores. In real time. Supporting thousands of audio and video channels concurrently.
Also, even for NVEnc and NVDec themselves, in some of the GPUs they do not use any specialized hardware and use normal GPU cores instead (e.g. see the older GM20x GPUs).
I'm not sure they are for sale either (probably for the right price), since they sell these "as a service", which pays better. I also don't think these are on sale for "small" customers.
Note: while I work at Netflix, I do not work on anything related to video encoding, these are just my semi-informed understandings :)
The main enabler of more advanced compression strategies is the higher performance available to the encoder. Decoders are essentially deterministic state machines. Encoders however have to search a large space of transformations to find ones that captures the entropy for a given situation.
In the development of AV1 they called these transformation "tools" and the research encoder experimented with lots of them and only the most profitable ones made it into the standard.
The compression can thus still evolve even with the AV1 standard frozen, just like how x264 and LAME have gotten so much better over their lifetime.
That's quite thought provoking.
What do you see when the current generation codec is heavily bitrate constrained? First thing you'll notice are the ugly blocks with non-matching colors on the edges. Such primitive expression of lack of bitrate is showing quite clearly these codecs have very little understanding.
With a codec of the (probably far) future where both encoder and decoder have good "understanding", you wouldn't see any blocks, lower color depth, smudged details or so. If some detail is missing because of constrained bitrate, decoder will just fill it in intelligently.
It can go crazy far - when constrained with bit rate, the encoder can just send sort of a movie script - "Man stands on the beach and looks into the distance" and the decoder will "understand" it and render it in awesome detail including adding non-mentioned splashing waves, man's thinking expression and perhaps sunset (decoder effectively becomes little automated film director who takes missing information in movie script as artistic license). It will most probably look very different from the original movie, but it will be believable - if you didn't see the original then you won't really have a reason to suspect it's fake.
With more bitrate available you can add more detail to the movie script - the man is black and old, he wears this and that which the decoder would take into account when reconstructing the video. At some point you're always constrained on the bitrate, but "understanding decoder" can always go deeper and fill in more realistic details when needed.
It is true that for lossless compression there are constraints on the size of the input and the size of the compressed artifact+decoder, but I believe for lossy compression, the fourth element is the perceptive capabilities (or preferences) of human minds.
A lot of the constant improvements is just that the amount of compute that is reasonable to use keeps growing, and so new codecs that are balanced around ever slower algorithms keep getting developed. I'm not saying that there is no new inventions -- plenty of useful tricks keep being found, but most of the difference of AV1 and what came before is just that the developers thought it was okay to spend more on search.
Apart from that, I am really glad to see how Rust is slowly making an appearance in such fundamental multimedia libraries.
AVI is already a container so it would just be confusing.
No, AV1 is evolutionary. You'll "just" better quality for the same amount of bits (just like going from MPEG2 -> H.264 -> H.265 etc). The license of AV1 however is revolutionary.
It hasn't made it upstream yet since I can't bump the minimum required version in FFmpeg git until rav1e tags 0.4.
We are using Rayon for multi threading, and also tiling to boost the output
Is it faster than dav1d?
rav1e = rav1e is an AV1 Encoder
dav1d = dav1d is an AV1 Decoder
https://github.com/xiph/rav1e/issues/2378
https://github.com/xiph/rav1e/issues/2310
Does not seems safe at all, definitly not the "safest"encoder out there.
No one uses “unsafe” in C (because it doesn’t exist). Does that mean C libraries are safe?
Is there some safer encoder you have in mind?