Mozilla finds CPU bug (bad store forwarding) in Samsung Galaxy S20
bugzilla.mozilla.org
bugzilla.mozilla.org
I had taken over maintenance of a library that allowed predicting satellite orbits from TLE data, which was used by the European Space Agency (and others) for the occasional mobile app.
Sporadically, we were getting reports of strange situations where the altitude of a satellite was way off, by hundreds or thousands of kilometers. The bug was really difficult to track down and reproduce, and by chance one of our Samsung tablets finally started showing the behaviour.
It turned out that Dalvik (the now obsolete Android VM) was figuring out that the TLE calculation was in a hot path and would JIT it to run more efficiently, and the optimized code used some native arithmetic (or was it trigonometic? can't recall) call that had a bug. We saw that the incorrect values started showing up right around the time when the JIT started kicking in. Fortunately, we could turn off JIT for specific devices and we rolled out an update that fixed the issue.
Fun times!
This gave me a hunch that it could be related to JIT, and then I tried completely turning it off, which caused the issue to disappear.
Or hope for microcode updates to fix the bug at the CPU level.
Certain MIPS CPUs also have incorrectly implemented instructions and the -mfix-* compiler parameters could be used to enable workarounds.
JIT compilers can do the same.
[1] https://bugzilla.mozilla.org/show_bug.cgi?id=1833315#c13
[2] https://hg.mozilla.org/mozilla-central/rev/c6a272593c07#l7.1...
> Tip #33, pg. 95:
>
> “select” Isn't Broken
>
> It is rare to find a bug in the OS or the compiler, or even a third-party
> product or library. The bug is most likely in the application.Documentation / wikis etc. advised against that not just for performance-of-generated-code reasons, but also because of compiler bugs using -Os was considered unreliable.
Note this was with ancient GCC (v2.9.x, early 3.x releases iirc)
If you think select is broken, read the docs and make sure you're using it correctly. If you still think select isn't doing what the docs say, try to reduce your actual use into a minimal test case and see. Go look at the source for select on your OS. Search the internet and see if anyone else ran into the same thing (but be careful, because there's lots of bad leads out there).
Then, when you've done that, send it to samsung so they patch it in their CPU design. They may even be able to patch it in their devices in the field (it's common for CPU designs to have special hardware whose purpose is runtime detecting certain trigger conditions and running replacement code to avoid hardware bugs discovered after the CPU was produced).
They can also issue an errata which might let compiler authors modify their code generators to avoid any pattern that triggers the bug, therefore avoiding other projects being impacted.
I never loved Samsung devices but since then I avoid them even more.
https://www.theregister.com/2012/03/07/amd_opteron_bug_drago...
I wonder if it happens to them too?
The Mozilla messages fail to mention this information that is essential for any CPU bug, which CPU model was used in the phones with crashes.
It is more likely that the bug is present in the Samsung Exynos 990 CPUs and not in Snapdragon, because Cortex-A77 was used much more widely, so it is improbable for such a bug to not also be noticed elsewhere.
The standard S20 used Snapdragon in North America, and Exynos in Europe. For the S20FE the Snapdragon version was sold in Europe - which was a selling point by itself.
I did not know that the S20FE had Exynos at all until I just looked it up... Apparently version 1 does, while version 2 uses Snapdragon.
I refuse to buy any phone with an Exynos SOC ever since I got screwed with a poorly tuned Exynos 9810 in my S9. Andrei Frumusanu did a great two part article on turning the 9810 when he used to write for Anandtech:
https://www.anandtech.com/show/12615/improving-exynos-9810-g...
https://www.anandtech.com/show/12620/improving-the-exynos-98...
* Octa-core (2x2.73 GHz Mongoose M5 & 2x2.50 GHz Cortex-A76 & 4x2.0 GHz Cortex-A55) - Global
* Octa-core (1x2.84 GHz Cortex-A77 & 3x2.42 GHz Cortex-A77 & 4x1.80 GHz Cortex-A55) - USA
(from https://www.gsmarena.com/samsung_galaxy_s20-10081.php )
> but only if you collect the data.
If the problem is impacting the user and they want it fixed they'll happily tell you. This bug (which causes crashes) would've been investigated just as well by asking users "Firefox has crashed - send bug report?", while giving the user the choice to decline if they had sensitive data in the browser's state.
Obviously nobody in their right mind should opt-in to generic telemetry because all the "product improvements" excuses never panned out - in the last decade software has only become worse, more annoying, and less feature-full, so people are making the rational decision and are not opting into something that's not actually benefiting them. But opt-in telemetry in specific cases where the user actually has a problem like this occurrence can work just fine.
Even when not abused to violate our privacy, the end result is often highly questionable. In place of good design, common sense, and user studies, companies try to extrapolate trends/patterns from data where there are none, resulting in mediocre decisions at best, and self-feeding cycles of devolution. ("Users are clicking this button a lot, let's make it larger" - users are actually clicking the button by accident.)
Telemetry doesn't have to be 100% anonymized - it can (and in some cases needs) to contain PII, just needs to be collected respectfully and transparently. Ask the user before sending something, and let the user review it. The user can then make their own decision.
I under The sentiment, but say you were sending a memory dump as part of your crash reporting. How is a user going to review that?
It's not feasible for a smaller developer, but it wouldn't be unreasonable for Mozilla to have a big lab bench with like 100 different devices on it that are just constantly running a script to start Firefox and load a suite of test sites.
Without substantially increasing the cost and work to maintain the test pool, you'll be testing exactly one combination of OS version/patch level/settings, one internet connection medium/carrier, one set of browser config options, etc for each device in the pool.
The investigation into any fault isolated to a single device in the test pool has to start with suspicion of the health of the specific device. The test S20 crashing doesn't tell you that every S20 will crash in the same way, it could be an isolated local fault. Yes you can throw more devices but that's further increasing the cost and effort of maintaining the test pool.
Android has an immense variety of devices. Without a representative sample, you wouldn't know which devices to buy. Regional variants sometimes have a huge difference and sometimes not; this bug seems to be tied to non-US Samsung flagships, which you might not buy if you're US based and don't know what people are using.
I'm not saying bad actors don't exist, just that telemetry is often used for good, with no nefarious goals.
Telemetry is probably just one more excuse to ship crap in the first place.
You spelled 'send crash dumps with the user's permission' wrong.
Useful for what? We've had an explosion of telemetry and "analytics" the past decade and software quality and functionality has only gone downhill. In some fields we've outright regressed compared to early 2000s tooling.
A simple example is that despite all the telemetry, manpower and tech innovation, every single mainstream communications product out there is inferior to pre-Microsoft Skype (despite the latter being built with much less manpower and running on much more primitive hardware & bandwidth).
Of course, not all of this devolution is to blame on telemetry alone, but I disagree that telemetry is some sort of requirement or game-changer when it comes to software quality. We've successfully built good quality software before it and within reasonable budgets.
I do agree that it seems that despite increased capabilities across the board and piles and piles of newer and better ways to do things, the quality of consumer software has dropped somewhat.