Memory-Safe WebP Decoding
halide.cx
halide.cx
Also, wuffs is not bit-identical to libwebp, and animated WebP is completely unsupported, so I'm not sure it is a real option for WebP decoding.
We got a pull request (https://github.com/google/wuffs/pull/168) in February to add lossy WebP support. Lossless WebP (which in some sense is an entirely different format, just reusing the WebP "brand") has had a Wuffs implementation for a couple of years now.
Anyway, the PR included SIMD acceleration and performance was on par with libwebp (C code).
The PR's code was, as far as I could tell, somewhat or mostly AI assisted. While that's great in terms of features, I still have more confidence in hand-crafted code.
I have since been working to manually rewrite the PR. I'd also like to add animation support, and last month I landed some Wuffs tooling changes re animated PNG, to be better able to (as a comparison baseline) decode and test animated WebP.
A lot of that manual rewrite has been committed, but the SIMD parts haven't landed yet. So yes, for what's on the main branch (not the PR), performance is not as good as libwebp yet, but landing the SIMD parts should fix that.
> wuffs is not bit-identical to libwebp
That's news to me. PR 168 says it produces pixel-identical output to libwebp. And what's in the main branch aims to be pixel-identical, e.g. for YUV to RGB conversion, it implements libwebp's formulae, not libjpeg's formulae. Both use BT.601, but libwebp uses studio range and libjpeg uses full range.
Can you link to some example .webp images that are not bit-identical?
Congratulations to the Halide folks on wpd, their new decoder. The performance numbers are impressive. If wpd (a Rust library) works for you, great, you should use it!
But since Wuffs was mentioned, I'll just drop a few selling points for why you'd still consider Wuffs.
1. Wuffs' implementation is transpiled to C code (and that in-C-form is checked into the repository, as well as into the leaner google/wuffs-mirror-release-c Github repo). If your existing project is C/C++, not a Rust one, then it's very easy to add Wuffs as a dependency. It's like adding any other third-party C library. It's just not hand-written .c code. It's hand-written .wuffs code that gets transpiled to a single-file C library, as easy to integrate as the STB libraries but memory-safe.
1.a. Similarly, if you're a Python project, or Java, or whatever, if you can wrap C code, you can wrap the Wuffs library (in its C form) and still get in-process, memory-safe image decoding without having to add a new toolchain to your build process.
2. Wuffs is a zero-capability language. It's a language for writing (safe) libraries, that only compute. It's not a language for writing applications. The Wuffs language cannot open files or write to the network. It can't even dynamically allocate memory. That means that Wuffs code can operate under `SECCOMP_MODE_STRICT` sandbox that prohibits basically everything except reading from stdin and writing to stdout.
2.a. For out-of-process, extremely memory-safe image decoding, the example/convert-to-nia/convert-to-nia.c program in the Wuffs repository reads an image (JPEG, PNG, WebP, etc) on stdin and writes NIA (a trivial image format, similar to Farbfeld) on stdout. Even if you don't want to audit the Wuffs language and toolchain itself, the security review for that convert-to-nia program is also absolutely trivial, because one of the first things that the main function does is to self-impose a `SECCOMP_MODE_STRICT` sandbox.
3. Wuffs uses intrinsics for SIMD, like C/C++, but memory safety is enforced on all loads and stores. And the toolchain also enforces that Wuffs general code can't call into Wuffs AVX2-using code unless your cpuid is AVX2-capable. The memory safety story of all of that is a bit better than just "we do SIMD via assembly in unsafe blocks". Wuffs (the language) doesn't even have an "unsafe" keyword.
The handwritten SIMD kernels sit behind safe Rust wrappers that calculate or validate the exact slice bounds before producing raw pointers, and CPU-specific routines are only installed after runtime feature detection. wpd even has guard-page tests specifically checking that the assembly kernels don't read or write beyond those wrapper-established windows. Handwritten SIMD was valuable enough that we felt this was all worth it (see some of the dav1d talks for justifications, we share these views).
Wuffs does still have a stronger formal property in the fact that its compiler checks the loads & stores inside the SIMD implementation itself, but most of the interesting memory safety attack surface (parsing and decoding attacker-controlled structure) remains entirely in safe Rust in wpd anyway, so I think wpd is still safe and incontrovertibly a significant improvement over libwebp anyway.
P.S. you can disable wpd's assembly if you feel so inclined, and get similar performance to Wuffs with safety.
They said:
> Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
> For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
A programming language specifically geared for file format libraries, including safety and performance
This is great but it’s got caveats:
- assembly code. They may have been careful but it’s an escape hatch
- if it has basically any dependencies then those are likely to transitively pull in more unsafe code
If you can't afford WUFFS then this Rust is the sensible choice. There aren't going to be any real cases where you can afford Fil-C but you can't afford WUFFS.
A WUFFS codec is obviously going to be a lot faster than Fil-C for similar development effort but it's also going to be entirely safe at compile time whereas in Fil-C too bad any mistakes will get caught in production as denials of service.
The Rust is a compromise in that regard, it's not going to catch everything WUFFS would catch during compilation but it will catch a lot more than Fil-C
Anytime you already have the C or C++ code but not the WUFFS.
The main use case for the Fil-C build of webp is my memory safe web browser, which is built entirely using Fil-C (including all dependencies).
> WUFFS codec is obviously going to be a lot faster than Fil-C for similar development effort
For greater development effort. Fil-C is just C, so WUFFS is only comparable effort if that’s the only version of the codec you ever write, and in the unlikely case you are as productive in WUFFS as you would have been in C.
> it will catch a lot more than Fil-C
Like what?
Fil-C catches not just the bugs in your codec but any transitive misuse of it from other C or C++ code. So, while you might be able to pick out things Rust or WUFFS catch that Fil-C doesn’t, I can easily pick out things that Fil-C catches that nothing else can.
But we literally do already have WebP for WUFFS.
> Like what?
Take a recent (September) double-free in the C codec, it's easy to write that in C. We have a Thing, we sometimes make an alias to the Thing, and then later we free both the alias and the original Thing. In Rust when you try to write this the compiler is unhappy. We can either have the same realisation which led to that being fixed (we do not need this alias, it should not be created) -- which I'd say is likely for a competent Rust programmer or we can clone the Thing, and then free both the clone and original, which is not a bug unlike the C library but is a waste of performance.
This comes down to how strong your claim is, and what point you're trying to make.
If you're saying that compiling codecs in Fil-C is not a good use of Fil-C in general, then that's wrong, because not all codecs have a WUFFS version.
If you're saying that compiling WebP with Fil-C is not a good use of Fil-C because there's a WUFFS version, then that's also wrong, because I can't use the WUFFS version in a fully memory-safe web browser, since WUFFS doesn't support Fil-ABI and nobody has written a browser where the whole transitive closure of everything that the browser uses is WUFFS or some other safe language.
On the other hand, that is possible with Fil-C. (I'm posting this comment using a memory safe browser built with Fil-C, which has WebP built with Fil-C as well as everything else down to the libc/libc++ layer built with Fil-C.)
> double-free
Double-frees are deterministic panics in Fil-C. So yeah, Fil-C catches that, and any other memroy safety violation that could be used to achieve weird execution
What does "deterministic panic" mean here? It sounds like that's a runtime panic, but the whole point of the paragraph you were quoting is compile time safety.
WUFFS is able to guarantee safety of the codec at compile time.
Rust is able to offer only some of its guarantees at compile time, and the rest are delivered at runtime.
Fil-C delivers almost all of its safety guarantees at runtime.
You can also run the same wasm on the command line:
npx @qip.dev/qipx qip.dev run \
image/webp/webp-to-ktx2-r8g8b8a8-srgb.wasm \
image/ktx2/ktx2-r8g8b8a8-or-b8g8r8a8-srgb-to-png.wasm \
< input.webp > output.pngIt’s neat that you can compile it without hand-written assembly for more safety, but what does the performance look like in that case? I assume the presented benchmarks don’t show this?