Issue 62938 – Barometer driver hangs and kills accelerometer on its way
code.google.com
code.google.com
If someone wants to look into it and is willing to tinker with the guts of the phone, the first thing is to see the state of the bus when the devices stop responding. I2C's "idle" level is high thanks to pull ups, the devices only drive the bus through an open drain.
Sometimes a device's state machine will go fubar (either because of a hardware bug or programing error) and will lock the bus down, basically making it impossible for anybody else to use it.
If this happens the next step is to disconnect/reset all other devices on the bus to make sure which one is screwing up (that's the difficulty with I2C, since the two lines can be driven by any master or slave you can never really know who's doing what when things go wrong).
An other thing to look for is the level of the line. Since there are many devices and pull ups on the wire it's not common to have messed up levels (0 is really 0.2 or 1 is really 0.8, if the pull up is too strong or too weak respectively). Depending on temperature and other factors that can lead certain interfaces to sample bad values.
And then well... you have to capture the transaction that causes the lock up and try to understand what goes wrong...
As a quick fix it might just be possible to force a reset of the bogus chip when a lockup is detected, that would prevent having to restart everything. There's usually a GPIO for that (if they wired it...).
I hope you have a good oscilloscope!
Quite frankly I can empathize with the dev not wanting to look into this bug, by the looks of it that's the kind of minor bug that'll take several days to track down and fix.
> it might just be possible to force a
> reset of the bogus chip when a lockup is
> detected [..] There's usually a GPIO for
> that (if they wired it...).
Yes, if they wired it. In my experience, actual design with this best practice is frustratingly rare. It's as if the hardware designers think, "Oh, it's just I2C, what could go wrong."If they didn't do it, the chips are probably only connected to a master board level reset, and you're basically screwed.
To find out whether it's the case on the affected devices, absent schematics or a scope, I'd grep around the kernel sources and look for definition of such a pin. (If sources aren't available, try symbols.)
Keep in mind how precious board space and GPIOs are in a modern phone. The boards are tiny, and even with 8 or 10 layers, they are still packed with traces on each layer.
Add to that, that the processor itself doesn't have a lot of GPIOs left over for a design like this. There are soooo many peripherals these days. Sure, you could use a port extender, but that is (a) another chip, (b) extra board space, (c) extra cost. So that's not going to happen unless something really important needs it, like the audio subsystem.
Look at the datasheet for the BMP280 barometer chip, the AK8963 compass, and the MPU6500 accelerometer (close enough):
http://datasheet.octopart.com/BMP280-Bosch-datasheet-1369120... http://www.akm.com/akm/en/file/datasheet/AK8963.pdf http://dlnmh9ip6v2uc.cloudfront.net/datasheets/Components/Ge...
The compass (AK8963) has a reset line, the other two have no reset line at all. Your best bet is to drop VCC, but what are the chances the hardware guy just tied them to the power bus and left it at that?
That is, if you also have even more circuitry to disable the bus pullups; otherwise they will continue to power the device via clamping diodes, potentially calling further confusion. The hardware guys don't implement I2C slave power control for a reason: it gets bloody expensive real fast - both in terms of BOM cost and PCB estate.
Having said that, the lack of a RESET line on many I2C slaves is just ridiculous. I have a sticker on my monitor that says "fix the hanging MMA8451Q bug"; it's been there for at least three months. The very thought of debugging I2C transactions makes me stop even trying.
As it is there's strong presumption that the bug is not in AOSP, so that's missleading. As for the "impeding science" part it's clickbait more than anything else IMHO.
There's very little context on what the bug is and how to reproduce it besides some vague references to PressureNET (which means very little to someone who hasn't used that before).
This bug looks to be between the kernel driver and sensor HAL to me. It might be fixable in code we can see, but none of that is part of "AOSP".
but the whole embedded market seems like this now - everything just the wrong side of reliable and abstractions an impossibility. maybe I am just too new to the area.