Linux may have been causing USB disconnects
plus.google.com
plus.google.com
Props to Intel for hiring leading Linux developers and turning them loose.
As an EE turned "software engineer" this bothers me, a lot.
I like the EE part of it, but I prefer thing that change more easily and are more "playful" (not to mention today hardware is at the mercy of software, so you take the reference design and go with it)
But I've come into situations where I uncovered a HW bug (in the chip reference board implementation, no less) that only manifested itself because of something specific in software (in the HDMI standard - or better, things from the standard inherited from things like VESA)
The Software Engineer see ports/memory to be written to and doesn't know what happens behind that
The Hardware engineer sees the "chip" and its connections but doesn't realise the rabbit hole goes deeper "ah this is a simple USB device, only 8 pins" now try communicating with it
Some software engineer working with drivers are distant from the hardware developers (especially in Linux) and even inside corporations there's a wall somewhere.
And of course, sometimes there's an abstraction between hardware and driver (usually through a firmware). Commonly relating to a standard, like USB storage, ATAPI, etc
" You can't make a piece of hardware without thinking about how the driver will work."
Unfortunately I've had to work with some devices that had very hard requirements on the software (basically, response time) (or you would add extra hardware to deal with it). In the second revision this problem was "fixed" by increasing a certain buffer size.
So yeah, sometimes hardware engineers don't think about that comprehensively enough.
Maybe there's so much complexity in the software stack that we can't start CS majors starting from the hardware level any more, but I can't help thinking we've lost something as a result. These days, there are Java programmers who get that "deer in headlights" look when confronted with terms such as "cache line miss".
The name of the course is "Computation Structures", or, as MIT students would know it (since nearly everything at MIT is numbered, including buildings and departments), 6.004:
lol, I see what you mean there. Now as long as the person you're talking to has a true engineering mind he/she will be happy to learn about the subject. But there's unfortunately also those that start looking you with eyes begging you to go stop the hardware mumbojumbo talk and go back to oftware only. I don't really consider them true engineers.
Forget that, most of these folks can't reason about a program that doesn't have automatic garbage collection. Even if they have direct experience with C or similar, I have asked recent grads how they imagine reference counting or malloc/free works, and they very often start pulling out GC-influenced magical thinking about "the system" reclaiming things under the covers.
Strongly disagree with that statement, though I sincerely wish it were true. My company manufactures hardware and does not provide a reference driver for any OS. We provide binary blobs and textual "guidelines".
For our hardware, driver authors operate without knowing any details beyond the interface.
I don't know who you think you're disagreeing with, but it clearly isn't me.
Based on the way this article is worded, it seems like there is no way to check when TRSMRCY is over. Imagine if you were waiting for a database query and if the database wasn't done thinking yet, simply accessing the socket would make the query abort.
I would rather prefer kernel hackers to do something useful, instead of wasting time and money to make every LKML archive look aesthetically beautiful.
hackers != designers
> 1337hax0ll
> hackers != designers
I guess not, so there's no reason.
(I bet all of them had their terminal font set to Courier New Bold while doing so)
Is there something wrong with me?
There is no "maximum" for a reason. Because it should be evaluated as "hey hardware developer, you will have guaranteed 10 ms from System Software to resume". If you don't wake up in 10ms, you are clearly violating the spec.
9.2.6.2 states: After a port is reset or resumed, the USB System Software is expected to provide a “recovery” interval of 10 ms before the device attached to the port is expected to respond to data transfers. The device may ignore any data transfers during the recovery interval. After the end of the recovery interval (measured from the end of the reset or the end of the EOP at the end of the resume signaling), the device must accept data transfers at any time.
In that table, TRSMRCY has a minimum value (of 10ms) but no maximum.
USB System Software is expected to provide a “recovery” interval of 10 ms before the device attached to the port is expected to respond to data transfers
A better spec might recommend or mandate values for fallback quanta and repetitions, or a maximum bound on the delay, rather than leaving it vendor-gets-to-choose.
I didn't look too hard, but a decently marked timing diagram would be nice, and might make it easier to spot the unbounded nature of it, rather than having to cross-reference the inline '10ms' value with the minmax table elsewhere.
[1] At least, if I understand it correctly. If accessing the status info uses the same mechanism as general traffic, it's obviously subject to the flaw described, and this interpretation is wrong.
[2] c.f. https://en.wikipedia.org/wiki/Don%27t-care_term and perhaps https://en.wikipedia.org/wiki/Virtual_particle
> 7.1.7.7 Resume
> The USB System Software must provide a [minimum] 10 ms resume recovery time (TRSMRCY) during which it will not attempt to access any device connected to the affected (just-activated) bus segment.
> 9.2.6.2 Reset/Resume Recovery Time
> After a port is reset or resumed, the USB System Software is expected to provide a “recovery” interval of [at least] 10 ms before the device attached to the port is expected to respond to data transfers.
... I would say that thinking the hardware can safely take more than 10 ms seems like a naive interpretation. You may note that system calls like usleep(10) sleep for "at least" 10 ms; there's no upper bound. The spec simply reflects this fairly typical aspect of software.
The people who really care about and study the spec, are those who have to support fixed devices i.e. USB devices internal to an appliance. They physically cannot be removed by the user. So suspend/resume has to work.
Embedded programmers have to deal with totally-broken drivers/specs all the time. There are probably 100s of folks who knew about this and dealt with it (bumped the timeout in their embedded kernel to match the devices they support) and never said anything to anybody.
> This bug has been reproduced under ChromeOS, which is very aggressive about USB power management. It enables auto-suspend for all internal USB devices (wifi and bluetooth), and the disconnects wreck havoc on those devices, causing the ChromeOS GUIs to repeatedly flash the USB wifi setup screen on user login.
This is an amazing fix if it is the root of the sorts of problems I've seen on Linux (which've kept me crawling back to Mac for hardware support)
I kid, I kid...
The reason is that it is incredibly difficult to link the disconnect to the cause as the 10ms is likely sufficient in 99% of cases - until it suddenly isn't. This means that you could be running test cases on a certain device for a year, and suddenly the test will fail the day after. When the test case mysteriously fails randomly like that on only a subset of devices, the assumption is that the hardware is faulty. These kind of failures would likely be higher on lower quality, less optimized hardware as well, furthering the perception.
As far as I can tell, the reason this is fixed now is because known good hardware from Intel started exhibiting the same error which got people at Intel to track it down directly, as they knew it wasn't their hardware at fault.
from TFA:
> the time is above 10 ms in about 8% of the remote wakeup events I've tested.
So 10ms is sufficient in about 92% of cases, barely more than 9 in 10.
>Out of 227 remote wakeup events from a USB mouse and keyboard: > - 163 transitions from RExit to U0 were immediate ( < 1 microsecond) > - 47 transitions from RExit to U0 took under 10 ms > - 17 transitions were over 10ms
So, 10 ms might indeed be sufficient for 99% of devices. But some devices (i.e. this mouse/kb combo) needs between 10 ms and 12 ms in 8% of all wakeups.
My brain hurts, but at least I'm never bored. :/
The spec says 10 too. It's the "at least 10" part that was missed. That's very subtle, does not stand out and is easily over-looked unless someone is really auditing code and reading specs carefully.
Also, somebody uses Google+ ?
That said, G+ is not half bad. The app beats Facebook hands down. Live Hangouts are also a neat way to engage with your audience.
I have a Das Keyboard that sporadically become unresponsive until I unplug and plug it back in. How do I know if my problem is caused by the issue described in the article?
Hopefully that helps narrow down your issue.
If you are going to make a statement like that, make it "64ms ought to be enough for anyone". Either way, it won't help. If you write kernel code or interface with unknown hardware, you must be paranoid to the bone to get robust code. Double so if your kernel code talks with hardware you do not control.
"more than any proper device should ask for."
If devices asked for time, things would be easy; you either reply 'no', or you give them the time they ask for. The problem is that they take time without telling you.
Embedded programmers know this; you can't ship working appliances without dealing with these issues.
You're forgetting the feeling of smug virtuousness you get when you end up being incompatible because you're more technically correct (the best kind) than the other components you're interacting with.
You get to say the other guys are all wrong, wage wars against them, blacklists, all the usual religious crap etc.
There was a mention about this in the OP. There were no interrupts for this state transition in USB prior to USB3. "The Intel xHCI host, unlike the EHCI host, actually gives an interrupt when the port fully transitions to the active state."
In addition, a lot of hardware initialization is based on delays and polling by design.
Your second comment about a lot of hw initialization being based on polling gives place to an "a fortiori": Why cascade this necessary evil to software. It's true, in the beginning, writing code for microcontrollers, you avoid working with interrupts. You work with assembly language, and get a bit lazy and hard-code delays (and even then, I'd choose a longer delay than the one specified as "max" in the datasheet : it's not like a chip not performing as well as is written in a datasheet never happened, so I take my precautions).
But then battery life reminds you of bad code practice: This continuous polling is draining. (And you put the uc to sleep :) )
Even more, if the application's involving sensors there's no way around using interrupts. Unless the thing is powered by a wall-socket, but even then, you get your consciousness preventing you from sleeping at night thinking about that horrible, barbarian code you put in there.
But then again, I don't know the USB spec and this may change so it's cool. And they have come a long way.
It occurs under all distros I've tried and it's been there for years. Even on different computers with different hardware...
Have you eliminated all other variables E.G. Common utility/setting, cabling, switch/router etc..? I've had paths open for much longer than that without access issues.
I don't even know what this means. What's the actual, reproducible scenario?
The router provided by my ISP (Virgin Media) is ruthless at closing idle TCP connections after only a few minutes. I'd see this with idle SSH logins being closed all the time.
The solution (for me at least) was to ensure connections used TCP keepalives, and vastly decrease the keepalive times (various sysctl calls, I don't have the details to hand).
"man nfs" and "man mount.nfs" and search for "soft", "intr", "tcp", and "timeo". It sounds like the solution to your problem is some combination of those options.
this terminology doesn't really make sense in the context of a TCP connection.
but, anyway, check:
net.ipv4.tcp_keepalive_time
net.ipv4.tcp_keepalive_probes
net.ipv4.tcp_keepalive_intvl
docs for the values here:
https://www.kernel.org/doc/Documentation/networking/ip-sysct...