Everything Is Broken: Shipping Rust-Minidump at Mozilla
hacks.mozilla.org
hacks.mozilla.org
They have to deal with basically random apps doing whatever they want and it sounds like hell.
I've been involved with minidumps in one way or another since around 2010. Was at a startup at the time that had a browser based on Chromium and we needed crash reporting for our own app. So I wrote a pretty simply backend that received minidumps, ran them through the breakpad processor and shoved the output into Splunk. That was our crash-reporting system.
Circa 2013 the company gets acquired by Yahoo which at the time was using Crittercism for its mobile apps but Yahoo wasn't happy with it. Somehow I was now the mobile app crash reporting expert at the company though so I built a whole new in-house crash reporting solution.
For iOS I wrote an SDK around PLCrashReporter because unwinding stacks on the client works out way better on iOS than dealing with a minidump.
For Android I had to deal with both JVM (er, Dalvik, er ART) stack traces, easy enough, but also native code crashes. For the latter I used breakpad's crash handler and minidumps. But it turns out that minidumps from Android devices are almost useless for two reasons:
1) If the crashes originate in managed code or calls into managed code you can't trace back through the managed code frames from a minidump. Especially if you don't have frame pointers.
2) You basically cannot get the symbols for all the different flavors of Android. Without symbols any stack trace that breakpad reconstructs is pretty useless.
Eventually I abandoned minidumps on Android and instead unwinding on the phone using corkscrew, wait no, libbacktrace, wait no, libunwind. But that still doesn't give useful stack traces very often. In the end, I ended up capturing logcat output when restarting after a crash which actually tends to have the most useful stack traces.
Which is all to say, both Apple and Google make it really hard for a mobile app to find out why it crashed. Both Android and iOS create a crash report for any app which crashes, but the app can't access those. So we're all shipping apps with third-party crash handlers built-in that try to capture a stack or minidump in-process and make sense of it later.
https://github.com/ivanarh/libcorkscrew-ndk
https://github.com/ivanarh/libbacktrace-ndk
https://github.com/ivanarh/libunwind-ndk
https://github.com/ivanarh/libunwindstack-ndk
Along with a wrapper that can use all of them:
https://github.com/ivanarh/ndcrash
It's been a while since I dealt with this and I don't recall why Android keeps changing unwinders. Looks like Sentry is using libunwindstack-ndk (well a fork of it anyway) on Android:
https://github.com/getsentry/sentry-native/tree/master/exter...
One thing I appreciate about writing Rust is that ADT support implies writing parsers is simpler under the "parse don't validate" mindset (which was clarified for me I think in this [0] article).
[0]: https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
What ADT does is give you enough flexibility so that a strict typing system doesn't suck.
Just to double check my understanding: are you talking about raw pointers (i.e. void*) being common in C++ and not in Rust? You're right that I was using ADT a bit loosely; to be honest the main value add for me has been the first class data-holding enums/sum types. C++ has std::variant, but the syntax support in Rust feels nicer.
You can't trust pointers have a value, or that the value is valid, you can't trust that your enums have a value inside their interval, or in fact you can't trust that any value from any type is inside its interval at all.
You also can't really trust that your values have the correct size.
We choose some of those to ignore, otherwise we wouldn't be able to program at all, but C++ gives you no guarantees at all about anything. The point is that if you do a parsing run in C++ and encode your value, you will still get many of the above problems because of bugs in your code.
If you don't set the underlying type, assigning a value that doesn't match an enumerator via `static_cast` is undefined behavior. See https://en.cppreference.com/w/cpp/language/enum . (Doing weird pointer casting things is also undefined behavior per the strict aliasing rule, though, come to think of it, I'm not sure whether memcpying an out-of-range value into an enum through the "reinterpret_cast to `char*`" loophole is undefined behavior.)
> If the underlying type is not fixed and the source value is out of range, the behavior is undefined.
Note the fine print about the meaning of ”out of range”:
> (The source value, as converted to the enumeration's underlying type if floating-point, is in range if it would fit in the smallest bit field large enough to hold all enumerators of the target enumeration.)
So this is not undefined:
enum E { A = 0, B = 1, C = 2 };
E valid = static_cast<E>(3);Choosing one just makes a difference because code has problems.
In particular, I always recommend "Text Rendering Hates You."
And the answer is always well, things are more complicated than they look. Even something as trivial as rendering text on a screen.
The great thing about Rust is that it reduces the knowledge and skill a developer needs in order to write a stable and (usually) performant program. This is a good thing.
Programs on the other hand can be written in entirely safe Rust, and the vast majority of Rust developers never need to use unsafe Rust to do this.
In Rust, we try to isolate unsafe usage and test it excessively. In C++ you have no such option.
And non-tree-shaped object graphs are not a library concern, but permeate entire applications; if you want to rewrite an app using shared mutability in Rust, you must either write code awkwardly with Cell/RefCell (and RefCell has runtime overhead), restructure the whole program in one Big Rewrite, or fallback to unsafe accesses (at this point, outside of multithreading C++ is a better Unsafe Rust than Unsafe Rust).
Please respond with a factual rebuttal before downvoting.
Repo: https://github.com/bluss/indexmap
IndexMap type itself: https://docs.rs/indexmap/latest/indexmap/map/struct.IndexMap...
It has amortized O(1) reads and writes, and it supports removal without maintaining order in O(1), or while maintaining order in O(n).
Sticking with Java, Rust also offers a better story around thread safety, such that sharing mutable state across threads requires types that allow that to work.
Finally, Java’s exception handling and null usage generally allows a lot of low hanging fruit bugs to slip through where in Rust these types of errors are far less likely to happen.
I never said Rust is “bug free”, I’ve been working with software far to long to make such a provably wrong statement.
All that being said, I would choose Rust over C++ for any new development anywhere. If my boss told me we had to use C++ for a new project I would actually quit. I've worked on plenty of C++ codebases (including Firefox) in my career. Sure, you can write bad code in any language, but C and C++ are just bad languages.
(I also ported Mozilla's sccache tool from the original Python implementation to Rust, which was a fun exercise. The Rust version is in production use in a wide variety of places, and AFAIK is still the only ccache-like tool that can cache Rust compilation.)
Let's give them credit for what we've achieved using them. I would definitely pick Rust over C++ any time, I respect C++ for all the cool things it gave us.
I once had some MS C code that called some 3rd-party library function. This was a long time ago, I don't remember all the details! But roughly, the code did a comparison, called this function, then made a jump conditional on the result of the comparison. But the condition wasn't being evaluated correctly.
It turned out that the library was built using a Borland compiler; my code was Microsoft C. The calling conventions differed; Borland expected the caller to save the flag register on the stack before calling, and restore it on returning; Microsoft expected the called function to do this. As a result, the flags that had been set before the function call were junk after it returned.
The way that Borland did it was the "convention" - as far as I'm aware, Microsoft stood alone on this. I formed the belief that Microsoft was doing things the way it did in order to deliberately make Microsoft C code incompatible with libraries built with 3rd-party compilers.
I was interested in the topic before reading, but it could have easily been a slog of technical minutia. I'm glad that wasn't the case!
Edit: the comment I referenced was deleted in the time I took to post this. It's probably for the best
For example, Rust deliberately doesn't have the tertiary operator, and random other types don't get silently coerced as booleans - so you can't write a = x ? 1 : -1; however you can write a = if x != 0 { 1 } else { -1 }; with the same effect. But Mara isn't satisfied with this verbose yet sensible answer, and proposes you could instead, for example:
a = x.count_ones().count_ones().count_ones().count_ones() as i32 * 2 - 1;
Hilarious? Or maybe terrifying? Entertaining certainly. https://twitter.com/m_ou_se/status/1404034056405368833?lang=...
Aria is more informative but I'm not going to end up choking and spilling my beverage all over the desk.
I think a lot of us can recognize this feeling when you have deployed big changes. Everything seems to be working fine, but you just don't trust it.
> You are reading part 1, wherein we build up our hubris.
Props to anyone willing to own their faults this readily:)
Now I know there's nice Rusty work happening with this stuff that I could maybe make use of for my next employer or a personal project. Neat.
The article is very vague about the details, which could be of interest to people in as similar situation.
Minidumps designed so well the initial Windows/x86 impl could be easily extended to multiple platforms like Google breakpad did. Symbols on demand over HTTP. For all the crap they get Microsoft got many things right.
I did not check lately, but is rust syntax still sane compared to the abomination which is the c++ syntax?
If the dump is corrupt then just stop trying to parse/make sense of it; it's garbage.
You can't expect "thing that runs when a process may have just experienced memory corruption" and "all builds of your application for all eternity" and "every toolchain you ever built your program with for all eternity" to be even vaguely reliable, because those things are in the past and we're trying to figure out how to fix the bugs people are experiencing in production today.
It is a horribly miserable answer to tell your coworkers "yeah sorry I know users are getting thousands of crashes this morning but the crash-dumper didn't sign its name in cursive so I'm gonna refuse to let you read the letter it sent at all".
And just an incoherent answer to say "yeah I know this is a stack overflow but it left the stack in a mildly corrupt state so I absolutely refuse to try to even look at the stack and figure anything out about it". Like, that is the entire purpose of a crashreporter, to investigate a program in an invalid state!
So the entire reason for being for things like rust-minidump are to make enough sense out of files that are known to be corrupt garbage to be able to find bugs.