Writing a JPEG Decoder in Rust – Part 2: Implementation I
mht.technology
mht.technology
That's not to say you could just throw this into the back-end of a public-facing website. It's still going to panic (abort) if someone unexpected happens, and it's still going to chew indeterminate amounts of CPU, memory or storage unless sandboxed (maybe even just ulimit). And there's still the danger one of the Rust standard library functions has a flaw in it. But this is the kind of starting point you wouldn't get with the plain C equivalent.
But to detect panics, you can use `afl.rs` to fuzz a parser: https://github.com/frewsxcv/afl.rs#user-content-trophy-case
while i < vec.len() {
encoded_data.push(vec[i]);
if vec[i] == 0xff && vec[i + 1] == 0x00 {
// Skip the 0x00 part here.The syntax "vec[i+1]" means "I expect this index to be in bounds, and if it's not, my program is so hopelessly broken it should be shut down immediately."
If you're really confident in your code and you need to wring out the last bit of performance, you could also use:
unsafe { vec.get_unchecked(i+1) }
...but in this case, your code would produce undefined behavior. If you want to detect this case and recover from it, you could also write: if let Some(value) = vec.get(i+1) {
// Use value.
} else {
// Fail with an error.
}
...or you could immediately exit the current routine with an error if no value is available: let value = try!(vec.get(i+1).ok_or(make_an_error()));
If you're writing a lot of parser code, your best bet might be to define a custom macro named that looks something like "try_get!(vec, i+1)" wrapping the final example above. > and you need to wring out the last bit of performance,
... and if you've already demonstrated that LLVM hasn't removed the bounds checks itself already.I find myself wanting to walk an iterator pairwise regularly for stuff like this and I'm using:
for (cur, next) in (& vec).iter().zip(vec.iter().skip(1)) {
encoded_data.push(cur);
if cur == 0xff && next == 0x00 {
// ...
}
}
Is there a better way to write this? What I'd really like is the equivalent to Clojure's partition[1]. I know that the stdlib has windows `(partition n (dec n) seq)` and chunks `(partition n n seq)` implemented on slices but I usually want this on iterators for the lookahead-like behavior shown here. let codes: Vec<HuffmanCode> = data_table.iter()
.zip(code_lengths.iter())
.zip(code_table.iter())
.map(|((&value, &length), &code)| {
HuffmanCode {
length: length,
code: code,
value: value,
}
})
.collect();
Nice! (mapv HuffmanCode. code_lengths code_table data_table)
I don't know if even dependent types would allow for that, given the (variable number of) arguments are all different types. val codes = (data_table, code_lengths, code_table).zipped map HuffmanCodeIt may be possible to implement `.zipped` on tuples though.
> (except the obvious `:Vec<_>` type that is not required)
It is required, otherwise, `collect()` doesn't know what type of collection to collect into. > It may be possible to implement `.zipped` on tuples though.
This is impossible without varargs, no?Nah, you'd just implement it as an impl over and over on tuples of reasonable size (up to 16 or so).
It wouldn't be elegant, but it'd work in practice.
Yes, unless there is some other mention of Vec, i.e. it is returned from the function.
> This is impossible without varargs, no?
Varargs would definitely help, and would allow Rust to borrow more ideas without workarounds. But one can do it manually for each tuple variant.
Not elegant, but it works.
My question is then: is there something fundamentally preventing Rust from achieving similar levels of expressiveness to this Scala example without incurring in unnecessary runtime overhead? Given that both languages are statically typed, my inclination would be to say no: the information is there, the compiler should be able to figure it out. But, alas, i know very little about these things, hence my question :)
Update: pcwalton gave some nice insight on why Rust needs some of these constructs on an uncle comment.
The `map` is uglier in Rust because
* the comparison was against a constructor with positional, rather than named, arguments
* the Scala code didn't need to dereference any arguments, and
* and Scala's `Zipped` is a special type with a special `map` function that takes three arguments, unlike a normal iterator.
The first and last points could be easily copied in Rust: you'd build a constructor for HuffmanCode and augment iterators of tuples with with a starmap method (that can be done in a library). The middle point can be done before the zipping. The result would be
let codes =
izip!(data_table.iter().cloned(), code_lengths, code_table)
.starmap(HuffmanCode::new)
.collect();
Rust's collect is never implicit, like Scala's CanBuildFrom. This prevents accidental collects, which helps writing fast code, but in principle I don't see why it couldn't be implicit - it would just require the whole standard library to be overhauled. let codes: Vec<_> =
izip!(&data_table, code_lengths, code_table)
.map(|(&value, length, code)| {
HuffmanCode {
length: length,
code: code,
value: value,
}
})
.collect();
(If HuffmanCode was a tuple type, this could even be #![feature(fn_traits)]
let codes: Vec<_> =
izip!(data_table.iter().cloned(), code_lengths, code_table)
.map(|args| (&HuffmanCode).call(args))
.collect();
but now I'm just playing around.)Core language features: both have algebraic datatypes like ML ("case classes" in Scala, enums in Rust), and a 'match' statement that makes working with them easy. Both have pervasive destructuring that works with match arms, 'let' bindings, etc.
Data structures and mutation: both encourage a pragmatic version of immutability -- use values and provide functional interfaces primarily, but support mutation where needed. In Rust there are 'Cell' and 'RefCell' types, just like ML's cell-based interior mutability.
Library idioms: all three have the standard set of functional-style higher-order transformation functions ('map', 'filter', etc. -- Rust in particular has a very rich Iterator trait API), and it's idiomatic to use these.
FWIW, Rust's first/bootstrap compiler was written in OCaml and there's a lot of obvious influence.
I don't want this to devolve into a No True Scotsman debate, I just feel that the parent post -- which decried a lack of similarity between Rust and Scala because they were both in the ML family -- was founded on a weak premise.
For me, type inference, strong typing, and ADTs with pattern matching puts a language into the ML family quite easily.
In effect, the "elegance" of Scala here is really the product of its heavy GC dependence. So the comparison isn't as meaningful as it might seem.
let codes: Vec<HuffmanCode> = data_table.iter()
.zip(&code_lengths)
.zip(&code_table)
.map(|((&value, &length), &code)| {
HuffmanCode {
length: length,
code: code,
value: value,
}
})
.collect();In the "Why" section he claims that Rust is a performant language with a low-level feel.
In this part, he claims that decoding a tiny JPG image took 2 seconds with his code.
How is that "performant" by any definition?
One common pitfall is not turning the optimiser on (with --release for cargo) i.e. benchmarking a debug build rather than a release one, and from a quick glance this code looks like it may be doing a system call to read each and every byte of the file due to the use of .bytes() on an unbuffered file, whereas it would be better to use, say, read_to_end to read many bytes of the file at once.
After fixing it, it apparently runs in 100 ms.
I am having some trouble reproducing the 2 seconds number, but with optimizations I'm getting around 1.5 seconds (without, I'm up to 5!), and as balducien mentions, /u/DroidLogician pointed out a _huge_ performance flaw in the file reading code.
Writing bad code is easy in any language (as I just showed :) - writing _super_ performant code is only possible in some.
EDIT: I made a pretty bad mistake while writing this comment, claiming that fixing the file reading code improved the speed of the decoder _a lot_, which was not true at all. In fact, the difference is barely (if at all) noticeable when decoding lena.jpeg. However, when reading a larger file, the improvement is definitely there. So sorry! Replying to HN comments while drinking and watching a movie turns out not to be a great idea after all :)
From the article it sounds like the 2s result was for the 90Kb file, which would have been rather odd.
I'd change "Rust" to "safe programming" here, but OK.
There have been lots of fun security vulnerabilities coming from programs trying to handle OOM gracefully. e.g. SpiderMonkey: https://bugzilla.mozilla.org/show_bug.cgi?id=982957, https://bugzilla.mozilla.org/show_bug.cgi?id=730415
https://doc.rust-lang.org/std/panic/fn.catch_unwind.html
Does anybody know if the standard library allocator aborts the process or throws a catchable exception? When they added catch_unwind they officially split panics into two types: abort panics and unwind panics. Code can choose one or the other when panic'ing. Also, programs can be built in such a way that all panics become aborts. That kind of sucks but it's probably one of the compromises needed to assuage those Rust developers who don't believe it's practical to handle OOM conditions.The default allocator crate always does an abort, both before and after this change.
You can of course write code that does anything, including use some other mechanism than liballoc.
There's motion on several fronts here:
The first, and currently unstable, is swapping out the implementation of liballoc for a different one. (we ship jemalloc and a pass-through to the system allocator with Rust)
The second, and currently in the "working on an RFC" phase, is per-object allocators.
Both of these are unstable because we're not 100% sure of the interface we want to stabilize yet, it's still a work in progress.
> when boxing requires dynamic allocation?
Boxing always implies dynamic allocation.Example: Basically any highly concurrent network daemon that multiplexes many clients on the same thread. In that case, you want much more control over where to put your recovery point. Even if the process doesn't abort, if the recovery point is beyond the scope of the kernel thread (i.e. a controller thread, which was the recommended solution before catch_unwind), that can be really inconvenient, and also requires a lot of unnecessary passing of mutable state between threads, which is usually something you try to avoid.
[1] Lua has robust support for OOM recovery, which is noteworthy because there can be situations where you both want to handle OOM but where a scripting language is more preferable. Example: An image manipulation program with scriptable filters, where you don't want an operation that can't complete to take down your process or thread.
Most software is riddled with buffer overflows and other exploits, and yet it's rare that you come across an intruder while he's installing his rootkit. That doesn't mean it's not happening, just that people are ignorant about it; and that things can appear normal even with rootkits installed.
Like buffer bloat, people can be experiencing a problem without even realizing it's a problem. When software crashes under load they just think that it's _normal_ to crash under load.
Or when it crawls to a snails pace under load because it's swapping like mad, they think that's normal, even though QoS would have been much better if the software failed the requests it couldn't serve rather than slowing everybody down until they _all_ timeout, sometimes even preventing administrators from diagnosing and fixing the problem.
OOM provides back pressure. Back pressure is much more reliable and responsive than, e.g., relying on magic constants for what kind of load you _think_ can be handled.
It's not nearly as simple as that. A general approach we try to use is the following.
- Small allocations are infallible (i.e. abort on failure) because if a small allocation fails you're probably in a bad situation w.r.t. memory.
- Large allocations, especially those controlled by web content, are fallible.
The distinction between small and large isn't clear cut, but anything over 1 MiB or so might be considered large. You certainly don't want to abort if a 100 MiB allocation fails, and allocations of that size aren't unusual in a web browser.
Aborting all 10,000 requests provides really poor QoS. You can get away with that if you're Google, because 10,000 requests is still miniscule relative to your total load across "the cloud". For almost everybody else it will hit hard, and among other things makes you more susceptible to DoS.
[1]: http://stackoverflow.com/questions/1692230/is-it-possible-to...
[0] - http://www.briancbecker.com/blog/2010/analysis-of-jpeg-decod...