However, I suspect parsers for big objects like images tend to have more vulnerabilities because developers try to avoid copying data, for performance reasons. Many memory-safe languages have their own performance issues, or make it hard to avoid copying bulk data, so those languages aren't a great fit.
This makes Rust a particularly good choice for writing parsers: Rust is pretty strong at supporting complex data sharing patterns while remaining memory-safe.
Don’t load them into memory, parse them as a stream byte-by-byte in accordance with the standard for the codec, check every offset before seeking, and reject images that don’t conform to the standard.
And of course, a ton of fuzzing to accompany it.
Also, maybe I'm wrong, but when I read "image parsing" I think that actually means "image decoding".
In environments where you’re prioritizing performance I’d still argue streams are likely your best bet when the size of the file to be parsed is not a constant. You wouldn’t want to load 50 large files into ram on a server environment let alone a phone.
If your input buffer is a bunch of tiny 10 KB files and you trust them? Sure, load them into memory and access their indices on the stack. Make sure you reuse the buffer to avoid unnecessary allocations.
If you want parallel processing with zero-allocations then streams with an array pool for their backing buffer are the best bet.
Not loading arbitrary files into memory will always be safer than doing so.
As for decoding - I believe the functions for validating if an array of bytes is an image should be far removed from the decoding and presentation of those bytes to the frame buffer. You don’t need to decode a JPG to validate that a file is a JPG. It either conforms to the standard or it doesn’t; the pixel data is irrelevant.
The goal is never just security.
E.g. for a Web browser like Firefox the priority has to be to be as fast or faster than the competition, THEN be secure. That's just the reality of what users care about. If the goal was just security we'd all have been using HotJava for the last 24 years.
The goal for Rust was performance plus safety. That's pretty hard to pull off.
> You wouldn’t want to load 50 large files into ram on a server environment let alone a phone.
mmap() works pretty well here.
> As for decoding - I believe the functions for validating if an array of bytes is an image should be far removed from the decoding and presentation of those bytes to the frame buffer. You don’t need to decode a JPG to validate that a file is a JPG. It either conforms to the standard or it doesn’t; the pixel data is irrelevant.
Yeah but in a browser for example you never want to just "validate" an image file, you want to decode it, and separating validation from decoding is just asking for trouble. That is the meaning of "parse, don't validate".
It’s still probably going to be wrong for a complex grammar, even if it’s not going to literally crash. Bratus et al did a survey of PDF parsers a few years ago and iirc all the popular ones were wrong in ways that don’t necessarily correspond to memory errors (like infinite looping).
> This makes Rust a particularly good choice for writing parsers
In particular, Rust has ADTs, which is the main feature I advocate for this.
Isn't inheritance to create hierarchy enough? why?
Node
SubNode1 SubNode2
SubNode1.1 SubNode2.1