Debugging a Futex Crash
rustylife.github.io
rustylife.github.io
There are also rule-based PII removal methods as used by Sentry minidump processing and described at [3].
[1]: https://news.ycombinator.com/item?id=37063459
[2]: https://crash-stats.mozilla.org/documentation/protected_data...
[3]: https://docs.sentry.io/product/data-management-settings/scru...
It is an issue with e.g medical data: Doctors using mswindows or any browser based tool can at any moment export random medical data to god knows which country. In the EU, it is probably illegal to even look at it if you receive it, and you don't know which dreports even have this issue.
Afaik, the only reason things aren't crashing down is total non-enforcement by all authorities. Which raises interesting question about their will and capability to enforce e.g. the medical secret at all.
On Linux and Android, you can similarly do unwinding using something like libunwind on the device. However, the traditional solution is that a "core" file is dumped and then is unwound later.
Years ago Microsoft invented a format called the "minidump" which is basically the stack for each thread, CPU registers, and other data. This format was adopted by Google Breakpad, and which ended up being used by both Firefox and Chrome for their crash reporting. The minidump format is often used when for whatever reason, unwinding can't be done on the device. It's a fairly flexible binary format.
Breakpad can perform sanitation before sending the minidump off the device, but this can be hit or miss.
Sentry (and probably other services) support minidumps among other formats, but minidumps are a PITA and I'll bet they hate supporting them. It's a large file despite its name, and it's expensive, complicated, and risky to unwind on the backend. The only open-source minidump unwinder used to be fairly esoteric C++ code that was part of Breakpad, but there's now a rust-based unwinder too.
Breakpad came up with something called a microdump, which is an even smaller version of a minidump that's text based instead of binary.
Both Apple and Google make mobile crash reporting way harder than it need be. Android and iOS have their own crash handlers that can perform on-device out-of-process unwinding, but the reports are not available to the app itself. You only get whatever reports Apple and Google feel like showing you after they get uploaded to their systems.
So apps that want more detail have to embed their own crash handlers. Because those handlers are operating in the same memory space as the app which just crashed, their operation is somewhat fraught with peril. At least on iOS, unwinding is straightforward due to mandatory frame pointes, and symbolication is easy because Apple makes iOS symbols available.
Android is a cluster fuck when it comes to native code crashes. You may or may not have frame pointers. So unwinding on the device is hit or miss. You also have to deal with both native code and ART (Android Run Time) code. Tracing back through the runtime is non-trivial to impossible. And you will almost definitely not have debugging symbols. So good luck with symbolication. The lack of debugging symbols also make off-device unwinding basically useless, so a minidump often isn't that helpful.
I designed and built the mobile crash reporting solution for my company about a decade ago, though we're currently migrating to a SAAS solution. On Android I evolved the solution from minidumps to microdumps to microdumps + on-device unwinding. It also turns out that there's a lot of helpful information sent to logcat and including that with native crash reports can often be more helpful than whatever you can capture from the stack during the crash.
Unwinding then occurs later by walking through the stack memory, which is difficult to impossible unless either: a) the stack contains frame pointers; or b) you have unwind data. Without either of those the unwinder uses heuristics to guess at return addresses on the stack, but that's very unreliable and often leads the unwinder into a loop till it gives up.
https://support.backtrace.io/hc/en-us/articles/360040517131-...
So there are at least some efforts to do it. IMO it can only ever be best effort, but it doesn't mean that it's not worth doing.
>>> frame 5
#5 0x00007f7430627733 in pthread_mutex_lock () from /home/rusty/futex/sysroot/lib/x86_64-linux-gnu/libpthread.so.0
>>> disassemble pthread_mutex_lock
...
0x00007f7430627723 <+211>: mov %rdi,0x8(%rsp)
0x00007f7430627728 <+216>: and $0x80,%esi
0x00007f743062772e <+222>: call 0x7f743062efc0 <__lll_lock_wait>
=> 0x00007f7430627733 <+227>: mov 0x8(%rsp),%rdi
0x00007f7430627738 <+232>: jmp 0x7f7430627689 <pthread_mutex_lock+57>
...
>>> x/w ($rsp + 8)
0x7f7371ff1ce8: 0x08001ea1
I love reading these articles, but sometimes they skip over things that the reader might not understand.In my case I don't understand where x and w come from, and how the result of 0x08001ea1 is invalid (I can see it's not aligned to a 4 byte boundary).