A photo is crashing some Android phones
bbc.com
bbc.com
I was fortunate enough to have developer mode & USB debugging enabled, which allowed me to install some code to clear the wallpaper using the WallpaperManager [0].
Sadly it wasn't as simple as pushing the code to the phone and running a command to execute it. When the device boots, it would crash _almost_ instantly and reboot into recovery mode after attempting to show the wallpaper a few times. This meant I couldn't put my passcode in, which meant Android refused to run the application I made to clear the wallpaper.
The ADB (Android Debug Bridge) tool [1] lets you replicate user input for the connected device, so you can basically swipe the screen / enter keys from the command line. I was able to use ADB to swipe up on the lock screen and enter my passcode. This was a start but wasn't enough, as the crash was happening before the wallpaper deleting app could be started.
I figured the unlock process was taking too much time, and found out after Android 8 (Oreo) you can actually use ADB to change or remove your passcode [2], provided you know your existing passcode. After I removed my passcode I was able to reboot the device, and run a single ADB command to swipe up on the screen and delete that cursed wallpaper.
And no it wasn't worth it, but I enjoyed the challenge
[0]. https://developer.android.com/reference/android/app/Wallpape...
EDIT: Some phones also require you to unlock it in settings too don’t they?
No need for dev mode that way :)
Edit: for this to work, the app must act as a service/receiver so it can start automatically.
Edit 2: now that I think about this, Google could easily fix affected phones remotely via their play store services.
Good work Google, I guess.
https://twitter.com/topjohnwu/status/1245956080779198464?s=1...
Another simple way to fix this if you have something like TWRP installed, is to go into data/system/users/0/ and rm -rf wallpaper
Sure, there's some work happening in the TWRP repos, but I can't figure out does that mean TWRP works on Android 10 or not.
https://gerrit.twrp.me/plugins/gitiles/android_bootable_reco...
I would consider doing write up. Another less skilled person may have tried doing the same.
There's no "Skia color profile" per se. When Skia's asked to write a JPEG from a pixel buffer tagged with a color profile (SkColorSpace), we auto-generate an ICC profile programatically from the parameters that describe the color gamut and transfer functions we use to represent those color profiles. That's what the code in this file is up to: https://source.chromium.org/chromium/chromium/src/+/master:t...
So this profile is one of infinite possible profiles that'll have "Google/Skia/..big long MD5.." somewhere in them.
This profile in particular looks like it's probably describing the ProPhoto color space, which is very, very large and uses imaginary colors outside visible light to increase its gamut. While the JPEG itself is always storing values in range in the ProPhoto gamut, when you convert that to another gamut (say, to match the display, or to sRGB) it's very easy to end up with logical color channel values less than zero or greater than one, which do make a kind of sense, but not to code that expects colors are always sRGB bytes.
It's unclear to me what gamut the colors being read in that Java code are in, and that'd very much affect what the appropriate values for that LUMINOSITY_MATRIX should be. Those are the correct values to get the luminosity for sRGB colors (though I think what's being calculated here is subtly-different gamma-encoded "luma" instead), but they're not likely a sound way to get luminosity if the image is in any different gamut. A typical strategy is to convert any image you load to the display gamut early, since you'll need to do that eventually to display it. And it's more and more common that phone display gamuts are not sRGB, usually something wider now like Display P3.
> out of bounds write in Java
Isn't Java supposed to be a language safe from these kinds of bugs? At the very worst it throws an exception, which should be caught and dealt with accordingly… How can that crash the whole phone?
Yes, from what I have read that's exactly what was happening, it was throwing an IndexOutOfBoundsException. Apparently, nothing was catching that exception, so the whole process was terminated.
> How can that crash the whole phone?
The main problem is that the process which was crashing is the Android equivalent of the window manager or the X server. It recovered from that crash either by restarting the process (which immediately crashed again), or by rebooting the phone (and immediately crashing again).
Apposite moment perhaps to refer to the mind-boggling finding in "Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems" (Yuan et al, Usenix 2014 https://www.usenix.org/conference/osdi14/technical-sessions/...):
"almost all (92%) of the catastrophic system failures [in the real-world study] are the result of incorrect handling of non-fatal errors explicitly signaled in software"
It absolutely results in the same kinds of failures.
Which is easy when you have an Either<L, R> construct for your returns or when you have checked Exceptions, which Java devs dislike.
So you get errors like this, because devs are often not aware or misremember all the failure modes of calls they make. Because they're only humans and they're using tools and practices that like to pretend everything that matters is the happy path.
Like what?
That said, I think the decision to do lambdas the way they did tacitly killed checked exceptions.
https://docs.oracle.com/javase/7/docs/api/java/io/ByteArrayO...
If an implementation of an interface can throw an exception that its interface doesn't declare, that breaks the abstraction. How can you safely use the interface? You either have to catch Exception (which is effectively the same as wrapping every exception in some InterfaceException class), or give up on handling any Exceptions thrown by the interface.
That's just how I phrased it 18 hours ago. The thing I find bad about it is not that exceptions have to be declared in the interface, that's good, it's that every implementations will have to declare them also, even in the case that a particular implementation has no failure mode.
But in general, of course an implementation can't just randomly add checked exceptions, because that would violate the Liskov substitution principle.
Unfortunately, because of the way Java handles checked exceptions, I can't feed the supplier with a lambda or any other method reference that I'm aware of that throws a checked exception and let it be passed up to the original caller directly. So I need to catch my checked exceptions and wrap them in an unchecked exception to catch. Not pretty.
There's probably something I'm missing but the language certainly doesn't go out of its way to help with this sort of thing.
Further, I'd say that IndexOutOfBoundsException can happen pretty much everywhere, so it does not make much sense to make it a checked exception anyway.
Indeed. I don't know of any programming languages in which out-of-bounds array access is an error whose handling is enforced by the type system (checked exception, result, etc). Ditto for numeric overflow and division by zero. I believe even Ada doesn't force error handling of these.
Perhaps there is no catch at all, so the systemUI process crashes. Since that process is critical, and responsible for things like the status bar, some supervisory process panics, and reboots the phone.
Alternatively perhaps a higher level catch clause does catch this error, but it is at too high a level to be able to recover cleanly by say displaying a black background instead of the wallpaper.
Hard to say without being more familar with the codebase.
Obviously, when one thing goes wrong in a program, frequently that leads to other things going wrong, but IMO thats better than crashing the whole process.
Sure, you might get unexpected output, but for many usecases, unexpected output is better than no output at all.
I think the post above explains why this isn't the case in mission critical software. I'd argue the lesson from software engineering is that it's not advisable in most situations, not just for nuclear reactors. The error was raised by a reason, after all. I'd say it's the opposite: there are some limited cases where you absolutely know the exception is not relevant, and in those cases you can knowingly swallow the exception.
It was an example of blaming a tool for a hypothetical misuse that is not the only way you can use it.
I'll go an extra mile and say I'm familiar with this construct in Basic, have seen it used, and nobody ever checked any error codes when using it. It's a footgun that's used almost exclusively to shoot at feet.
Just a single top level catch block would prevent other code after the error occurred from running, which might still be able to work fine without the results of the faulty code.
It's not. Crashing the process (if the error remains unhandled at the top level) is the right answer. When the program raises an exception, you want to either handle it or pass it up, not ignore it. It was raised for a reason. Once an error was raised, if the program ignores it and happily continues all bets are off. This is like reasoning out of false premises: everything goes, which is undesirable. Data is probably corrupted, maybe preconditions are unmet.
For this reason, most automated code checkers consider ignoring exceptions a serious antipattern.
I didn't always have this habit, but picked it up when I switched from working on servers to working on safety-critical hardware. It was a lot easier a habit to pick up than I thought it would be, and I've found it makes you write a more robust, well-understood system.
It's quite a responsibility on the human side of software. Handling unexpected errors is a pretty easy thing to miss, for programmers who are inexperienced or have become too comfortable (not paying attention to every detail).
I can't help but think that the language or compiler (or tests, I suppose) could enforce this better, so that it forces the programmer to explicitly handle recovery from failures - at certain layers/thresholds - so that the error doesn't bubble up to the top and bring the whole thing down.
As you pointed out, an error in the UI layer shouldn't have been allowed to trigger a system crash loop. But, that's kind of the nature of bugs like these, they creep in through unexpected logic paths. In which case, the development environment needs to better enforce the (explicit) behavior of all possible paths that an error could take. Well, I'm sure that's constantly improving, as a perennial challenge of reliable software, and there will always be a need for discipline and attentiveness by the humans involved.
Sure, agreed, this isn't a post-hoc "what an obvious bug, they should have just <done everything perfectly from the start>". I was just addressing the dichotomy laid out by the GP comment, which only mentioned top-level generic exception handling instead of case-specific, intelligent exception handling.
It's basically the same thing here, but without some physical keys bound to 'switch to new session'. And maybe you can do that over adb anyway, which is about as (not) normal user friendly.
The trouble is that it's really easy to say "access the 10th position in a 5-element array". It's possible to make a programming language that doesn't allow you to say this, but it either makes it really awkward to do anything with arrays, or simply pushes the error somewhere else -- or both.
Java doesn't claim it's magic. It's objectively safer by this measure. (More specifically, Java claims to be "memory safe" [1])
So, lesson learned: in safe mode, do not use a background image that the user set.
I love both (I'm one of the few who wrestled through the entire WoT twice and enjoyed that) but your citation shows very pointedly how Jordan was often just copying LoTR.
So maybe Jordan is actually the one who got ripped off. Still though, pretty cliché idea, but the formulation is too close to not be influenced one way or another.
https://support.microsoft.com/en-us/help/12376/windows-10-st...
Microsoft knows what its doing.
It's an unnecessary rigmarole to boot into safe mode without OS access, but is still possible.
Their prejudice toward...?
> "After setting the image in question as a wallpaper, the phone immediately crashed. It attempted to reboot, but the screen would constantly turn on and off, making it impossible to pass the security screen," he noted.
> Restarting the device in safe mode (by holding down the volume button during boot-up) did not fix the issue."
If restarting doesn't fix it, and not even safe mode will allow you to change the wallpaper to get around it, what method of recovery is left for the affected user? Factory reset from recovery mode?
It advertises that support, but I've never really seen it work. It seems that modern FDE on Android is such that you can only really "decrypt" from the system environment itself, not from different code - and it's not clear how to fix this.
works for me. it really depends on whether your TWRP distribution implemented it properly. AFAIK android phones don't have a mechanism to bind encryption keys to a system state (similar to sealing keys to PCRs for TPMs on PCs), so I don't think your theory is correct.
* @return A float value with a range defined by the specified color's color space public static int blue (int color)
Anyway this thing has range [0, 255] and adding three of them together as an index of an int[256] doesn't seem like it would ever work regardless of the colorspace.I think what's happening here is that the image is in a format that's holding those original red, green, and blue values in a format that can hold values outside logical [0,1].
Update: sorry, the comment below me doesn't seem to have a 'respond' link so I'll just edit one in here. I totally agree with you it's good practice to document your invariants in code, but in practice it would result in the same thing... an unhandled failure with nothing better to do than crash the process. In a way (if you squint) indexing into an array of size 256 is itself documenting the invariant that the index is less than 256. It's just that the invariant itself is wrong.
CHECK(Color.red(pixel) < 55);
Or whatever.Truth in advertising. If this had been named "getHistogramOfSRGBBitmapElseOOBE" nobody would have stamped it because that's obviously dumb.
For future reference: I'll often see that for a post in the context of a larger thread, but the 'reply' link has always appeared when I navigate to the post itself (by clicking on the timestamp). Maybe you already tried that, though.
This is done by using the formula: ".2126f * r + .7152f * g + .0722f * b" Apparently this will not yield out of bounds values if the colors are in normal range.
Yes, I've got some alarm bells going off in the back of my head about the float to integer rounding (or float multiplication rounding itself), possibly causing a value slightly too high, but it seems like that might not actually happen. Or perhaps it can only happen for some in range values that can only occur after a color profile correction. (i.e. the float versions of 8-bit sRGB never cause bad rounding, but coming from very specific other profiles might create such values). Lastly there is the possibility that the color values started out of range, so after the multiplication they can still sum to more than 256.
In any case, the sane way to do this, would be to create a true grayscale image with the matrix, and pick an arbitrary color component to look at. (I.E. using the matrix for both multiplication and addition, rather than only for multiplication, and then doing addition afterwards.) I'm guessing the `blue` et al static methods clamp their outputs, so this bug would have been avoided.
Quick edit to add that, despite this, I do not believe those changes have made their way into most devices. It seems the error stemmed from the possibility of returning a value over 255 when a histogram was calculated from the addition of color values. As stated in the article, this seemed to result from the use of the Skia color profile in particular. I do not know about other color profiles. The code mentioned by gruez was what I got when the emulator was crashing.
It's not just that the image had a non-sRGB colorspace, it's also that the pixel values in this specific image were out of the expected range.
Testing combinations of edges is not so common/obvious.
GIMP sees it as a JFIF (apparently a predecessor to EXIF) with
"Google/Skia/E3CADAB7BD3DE5E3436874D2A9DEE126 Copyright: Google Inc. 2016"
For everyone who tried it and wants to recover phone without clean wipe. remove these files (I did it with twrp)
/data/system/users/USER_ID/wallpaper_info.xml
/data/system/users/USER_ID/wallpaper_orig
/data/system/users/USER_ID/wallpaper
it resets your wallpaper to default oneApparently not, and it's not too surprising either. The only way such a photo would get on is either through the on-device camera, or downloaded from the internet. The former seems most likely, considering that most picture sharing sites probably normalize/sanitize/recompress whatever their users upload. Are there any android phones that produce HDR (ie. 10 bit color, not the effect that's commonly available) pictures?
They have literally billions of beta testers, who are even paying for the privilege.
What exactly is the issue?! Been using Android for about a decade now and can rotate the device just fine.
I actually think between the HAL layer, OEMs, chipsets and 3P ecosystem it is very difficult to dig out of the hole that Android has built. Too much was put into short term thinking on annual releases and gap closing with Apple. While Apple planned APIs years out for end to end products, Android had to blindly copy APIs and everything is far less elegant. Android isn't working. It isn't capturing the smartwatch ecosystem well (Fitbit) and it isn't capturing the TV ecosystem well (Roku).
The net impact is harmful to developers. We don't get the reach and compatibility we were promised. We have very fragmented ecosystem to develop for. We have specific device SDKs we have to use to tap into novel hardware. I can't complain enough.
Android had a huge opportunity to define the next evolution of compute platforms on top of the linux kernel and it has kind of blown it.
I think there are even signs Google leadership sees this and is investing in a portfolio of strategies such as CameraX, Pixel, Flutter/Fuschia.
With the Huawei situation too, I think the general health of the Android ecosystem is some of the worst it has been. Globally I think we are going to see even more fragmentation.
We need something that ties this all together. It's not going to be a full blown OS layer, but something to manage the madness for 3P devs.
Jetpack may be a good temporary solution, but there is no reason why the core APIs, the ones that sit between the hardware and 3P apps, can't just be better.
Same applies to rebooting frameworks every year, or half baked support for Java.
Latest examples being the c, naturally created as answer to Flutter, while they are still "selling" latest improvements on GUI tooling for Constraint and Motion Layout, which are going to be legacy when Jetpack Composer happens to be finally released as stable.
Speaking of which, all stable releases have regressions. There is hardly a stable release that isn't followed up by regression complaints on Android developer channels.
Oboe and Vulkan are very good examples how clunky everything is.
Here we have two API that are supposed to be relevant for the platform (real time audio C++ framework) and next generation 3D API support, for which we get told to clone github repos, compile everything from scratch and do the integration with the applications ourselves.
Comparing this with iOS SDK and XCode templates is just feels like having to nuke it all.
I have been moving away from NDK to WebGL 2.0 / WebAssembly, as it is good enough for my interactive demos and Chrome team at least seems to be a bit more sane.
Or if you need more low level access to the host OS, as TWA,
https://developers.google.com/web/android/trusted-web-activi...
As sidenote, the TWA idea started originally at Microsoft, where packaged PWAs get access to Windows APIs without you having to write your own wrappers.
https://developer.microsoft.com/en-us/windows/pwa/
I expect the ongoing Edge Chrome efforts to eventually improve the TWA experience.
Sceneform seems to be dead, and Filament is nothing more than a tech demo.
To be fair, I have had the same experience while doing iOS development.
At least Symbian had a proper C++ API, not a bunch of C wrappers to Java and C++ code, expecting everyone to write their own wrappers and cloning half maintained github repos released by someone on the Android team.
histogram = getHistogram();
should be wrapped automatically in: try { histogram = getHistogram(); }
catch(Exception e) { console.log(e.toString()); histogram = DEFAULT_HISTOGRAM; }
Something like: y = a / b;
should be wrapped in: try { y = a/b; }
catch(Exception e) { console.log(e.toString()); y = 1e999; }
Wrap every function call like this, and just don't let Java exceptions bring down an entire system in production especially when a lot of the time it's just a petty UI issue.Obviously never do it in development though. But odd UI behavior instead of crashes might make a bunch more users happy which is ultimately what matters.
A divide by zero error on a button's width shouldn't cause a spacecraft to give up and abort its mission to the moon. That wouldn't happen, but it's just a figurative description of what I would like. It's ridiculous for an OS to go into a boot loop because of a color histogram. A better behavior would be for the background to be set to an array of NaNs, and the background renderer to fail gracefully to a plain black or white background.
Really we need some kind of NaN value for all java objects.
https://techxplore.com/news/2020-06-wallpaper-image-android....
Basically everything where we can pass some specially encoded date (read as format or protocol) can be made with bugs and edge cases not covered by test or by intent look (how to say it properly in English? :)
So yeah, fuzzing.
[1] http://www.aerasec.de/security/advisories/decompression-bomb...
Also curious to know if perhaps there are flaws in the image format's spec that are being exploited?
To be clear, this was on my old phone not my main phone, so I didn’t mind testing it.
People are scared of AGI. This sort of shit scares me more.
Samsung fixed this with an extremely eraly June security update. I think it took less than a week.
IIRC, last time they had a really bad bug they updated some 4-5 year old Notes.
I wouldn't click from a phone just in case, though.
Seems like the crash is in some wallpaper-related code: https://news.ycombinator.com/item?id=23404772
It is making some of them unusable, requiring a hard reset.
Well that's still a "soft brick" right?
It would be interesting to train a generative NN on human EEG response to create images that produce novel or unusual effects on the viewer.
Other Windows included apps, like Paint, Paint3D and Photos open it just fine.
If notepad doesn't open it, nothing will.
site:openwrt.org "soft-brick" - 213 results
It's almost impossible to brick anything with software, and even in hardware unless you damage the PCB itself you can usually recover if you tried hard enough (e.g. Even if you fry a $1500 FPGA you could technically reflow a new one and pray the bitstream is OK if you need it fixed now - although I'd rather you than me)
I've always understood it to mean 'broken beyond software or high-level firmware repair'. i.e. if not actually a blown component, something needs to be reflashed which has an inaccessible JTAG header or something, rather than something 'higher level' like recovery images over ADB & USB.
https://www.merriam-webster.com/dictionary/brick
Under the 'verb' usage:
2 : to render (an electronic device, such as a smartphone) nonfunctional (as by accidental damage, malicious hacking, or software changes) // … those who dared hack the phone to add features … risked having it "bricked"—completely and permanently disabled—on the next automatic update …
I'm curious to see the post-mortem on this one, i'd guess a bug in the libjpg used in android, and a malformat in the encodibg of the picture ?
> it soft-bricked their phones (ie. they had to factory reset their device)
But the whole point of coming up with the word "brick" was to describe the set of scenarios where even things like factory reset won't work. This is "literally" all over again.
I like soft vs hard brick, because it connects to software and hardware. With a soft brick, the device is basically a brick until you reset the software. With a hard brick, it requires potentially a hardware fix.