Google fixes nearly decade-old Linux kernel TCP bug
bitsup.blogspot.com
bitsup.blogspot.com
(Google are actually in a good position to provide these data; a number of Google's own services will doubtless benefit. I suspect the scale of Google's operations means that even a barely-measurable decrease in latency or queue depth compounds to a big gain pretty quickly.)
> The whole web, including Firefox users, will benefit.
Why is Firefox thrown in here? And the post is even tagged with "firefox". But the contents of the post seems to have absolutely nothing to do with Firefox, except inasmuch as Firefox is an application that interacts with the web.
I often tag non-Mozilla posts as "mozilla" if they are still relevant to Planet's scope on my blog too. I've submitted a category feed instead of the full RSS blog feed so that I can exclude things which are physicsy or otherwise irrelevant.
The bug fix is in the CUBIC congestion control algorithm, which has been the default in Linux TCP since kernel 2.6.19 (released in November 2007). So impact in distros has been ~ 2008-2015.
That to me sounds like this is not really a bug, but a performance enhancement; a bug would be something like corrupted/missing/extra data at the application level. Tuning the TCP algorithms beyond the basics is really more of an art.
I would call this a bug fix since someone has found what is probably an unintended behavior in the cubic algorithm. The cubic algorithm itself though would be a feature improving on the original tcp algorithm (nagle, if I remember right)
> that an endpoint that is not moving any traffic cannot use the lack of errors as information in its feedback loop.
I think the bug is that it was counting something when it should have not been, which would have violated both the design and intent even at the time it was written.
But I agree that code is only a bug if it does not do what was originally intended, if the original design was bad, and the code correctly implemented that bad design, there is no bug.
The frame of reference for determining if something is a bug or not is the requirement specification, not design. If the design incorrectly or inadequately addresses the requirement, that is a bug as well. There isn't much point in rejoicing over the code that correctly implements an incorrect design!
if (x = 42) {
Then that is most certainly a bug.If a bit of text was supposed to be blue but came out green, maybe because the person writing the requirements doc got it wrong, it is hardly a bug in the code. The code is doing what the programmer intended, even if that is not what the user wanted. That's why the world has moved on from "bug trackers" to "issue trackers".
> There isn't much point in rejoicing over the code that correctly implements an incorrect design!
And there isn't much point blaming a coder for correctly implementing an incorrect design, especially if the design document is all they have to go on.
Sure. However, that is secondary. The sort of mistakes that you mention can be caught during code review or unit testing, completely independently of what the actual user requirement is. This post, evidently, does not involve such a case.
> The code is doing what the programmer intended, even if that is not what the user wanted.
Your perspective appears to be inside-outward. It may be helpful during a performance appraisal, but does the business no good!
Edit: inside-out --> inside-outward.
If the design has a flaw, whatever the reason: it is a bug.
A big part of development is "discovering the requirements." It is typically the hardest part for real world projects, and when that fails, it is perhaps not a bug in the code, but it still a development error that needs to be fixed.
If a developer got a piece of paper that said "blue" but the real world requirement was "green", and it was not possible for the developer to find out together with any stakeholder, then there is a development process error, that needs to be fixed.
In any event, it needs to be fixed, so pointing to someone else wont really help.
Imagine a car turning off traction control in a pouring rain every time you stop at the lights because 'Hey,no tire spin in the last 5 seconds!!1'
[1] http://www.linuxfoundation.org/publications/linux-foundation...
Red Hat is an amazing company.
Their CEO has written a management book (or perhaps anti-management) which explains their take on these things in some detail.
I'm pretty sure that if Red Hat had been running Android, it wouldn't be a big fork.
cough systemd cough
There are of course others working in these areas, but I think Google has been going after them: Van Jacobson is a high-profile example. Dave Taht is still subsiding on ramen though (https://www.patreon.com/dtaht).
wow - thanks for posting about David Taht. someone like Netflix or Google should get him on their payroll.
I was trying to setup ECMP on Ubuntu 14 and found I was constantly getting TCP RST. Turns out after digging into the code I found the path selection algorithm is not a consistent hash of a packet header tuple but rather pseudo-randomly chosen.
In older kernels there was a route cache so a path would be chosen and cached and so you would get a stable ECMP route per source. The route cache was removed so now the random selection runs on every packet leading to unstable routing.
It seems a simple fix of choosing the right hash function and maybe adding some configuration flags to determine which fields are hashes on (typically 5-tuple SrcIP, DstIP, Ip Protocol, TCP Src Port, TCP Dst Port).
But what would be incredibly cool is if my patch could one day possibly even merge to mainline. Is it reasonable to just blast a 'here's what I'm trying' out to netdev mailing list to get some feedback? I must admit I'm a bit intimidated to post...
It sounds like you have a good understanding of the problem. If anything, I'd look into why the route cache was removed; perhaps there was a good reason for this? If nothing else, you can ask and learn something.
Which, depending on your viewpoint, can be seen as: the only reason this bug fix is good is because some company can still sell crappy hardware and get away with it.
Then we have the user perspective of course, which is "hey, my device doesn't crash as often" - which is awesome. But given the context the hardware is still buggy and it will still crash, albeit not as often. Which really isn't much comfort in the long run. Rare hiccups are worse than daily hiccups. Because daily hiccups you learn how to handle, a rare hiccup can really bite you. So you really haven't gained that much anyway, and in the long run the rare hiccup is rare enough for you not knowing who to attribute it to so you might buy the exact same thing next time which of course is very, very, unfortunate.