Uncovering a 24-year-old bug in the Linux Kernel (2021)
engineering.skroutz.gr
engineering.skroutz.gr
Earlier in the article, the author mentions that they recently upgraded some network hardware, and the problem seemed to become more frequent after that.
Packet loss or other network issues would force the stack to fall out of fast-path and update the counter, avoiding the bug.
Running over ssh would avoid the bug. The only time you'd run rsync not over ssh would be within your own network.
So it sounds like (this is my conjecture here) this would only appear to someone running rsync internally, over a high-performance network with no packet loss, and upgrading the switches might've finally gotten the network good enough to expose the bug?
I wrote an article this past year that talks about silent bugs that slowly eat resources and collectively can be very expensive in terms of wasted time and energy: https://didgets.substack.com/p/finding-and-fixing-a-billion-...
My Didgets tool lets you create pivot tables against relational database tables, even very large ones. For the pivot values, you can choose to just count the occurrence of each value or if it is a number type you can add them up. You can also add up the values in a separate number column. Here is a quick demo video: https://www.youtube.com/watch?v=2ScBd-71OLQ
When adding up numbers in a separate column, I had just a few lines of unnecessary code that ended up being called exponentially. For smaller tables it was barely noticeable, but for tables with 30 million+ rows it really bogged down.
A simple fix to the affected lines caused a certain test against a large table to go from over 10 minutes down to under 20 seconds. The effects of just a few lines of code when applied to a big enough data set can really impact performance. It is the old Einstein equation E=mc2 in effect which is discussed here: https://didgets.substack.com/p/musings-from-an-old-programme...
I think the idea here is to write code quickly that's inefficient, and re-write it to be efficient if the performance is required down the line. For companies where there's bigger fish to fry, i.e. customer acquisition, it's more useful to pump out more features (even at the expense of bugs) because that draws customers.
But in places where performance is important, you do see developers squeeze out more cycles/memory. I.e. kernel/OS development, database servers, video games. It's just that most developers aren't in those areas of specialty anymore.
Btw, have you heard of https://handmade.network/ and https://en.wikipedia.org/wiki/Demoscene ? Wondering what your thoughts are in those areas. There are probably more communities like the ones I mentioned, where developers are interested in writing the kind of code that you are talking about.
Reminds me of the GTA Online quadratic time JSON parsing bug
As someone that is cursed to inevitably find some obscure bug the second I start using some piece of software I'm happy I'm not the only one
> I wrote an article this past year that talks about silent bugs that slowly eat resources and collectively can be very expensive in terms of wasted time and energy
"Using JS for backend is ecoterrorism" lmao
For instance, a small team of 40 people found no issue sending 4MB of json english to chinese string localisation to each website visitor for angular to translate. 1 million visitor a month in Hong Kong alone, 4 million MB + a few second of mapping per user per month completely pointless...
Imagine if this bug were somewhere in closed source software. You'd have to reach out to the software's customer support team. Every time I reach out to customer support I expect to have an unpleasant experience. It is rarely otherwise.
But if the bug is obscure and has little impact, bad luck!
It is an interesting thought experiment to consider what kind of tool or automated detection could have found this. Some type of dependency linking between variables might have shed some light, but I'm not sure that would have really highlighted this kind of issue.
Great description of both the bug and the path to the solution!
All of this is incredibly difficult, and an open area of research. Probably the biggest example of this approach is the Sel4 microkernel. To put the difficulty in perspective, I checkout out some of the sel4 repositories did a quick line count.
The repository for the microkernel itself [0] has 276,541
The testsuite [1] has 26,397
The formal verification repo [2] has 1,583,410, over 5 times as much as the source code.
That is not to say that formal verification takes 5x the work. You also have to write your source-code in such a way that it is ammenable to being formally verified, which makes it more difficult to write, and limits what you can reasonably do.
Having said that, this approach can be done in a less severe way. For instance, type systems are essentially a simple form of formal verification. There are entire classes of bugs that are simply impossible in a properly typed programs; and more advanced type systems can eliminate a larger class of bugs. Although, to get the full benefit, you still need to go out of your way to encode some invariant into the type system. You also find that mainstream languages that try to go in this direction always contain some sort of escape hatch to let the programmer assert a portion of code is correct without needing to convince the verifier.
[0] https://github.com/seL4/seL4
Also hire significantly more skilled people. Write formal verification on job requirement and the pool of candidates will shrink massively.
Explains why it is so rare really. "Spend 5-10x on developers to have some bugs not happen" is not a great sell.
At the time this bug was introduced it would probably have been cost prohibitive to create a test case. We were proud of 100mbit networks, had flaky nics the vendors didn't help maintain much of the time (and which were often broken in hardware) and the filesystem max file size was something like 2tb, and most drives wee're in the handful of gbs. Conceiving of testing for something like this would have been expensive. And none of the big system vendors took Linux seriously then.
Though perhaps flooding zeros across a TCP socket could work, I really think that a kernel hacker would have found a lot of other hardware and driver issues before ever being able to trigger this.
Unconstrained zero sending is too fast; you would tend to flood the connection, causing packet loss, breaking you out of the fast path long before counter loop around. You would need to avoid saturating the network while still causing the recipient to fall behind.
=> https://news.ycombinator.com/item?id=26102241 Previous Discussion (497 points - 41 comments)
I absolutely disagree. Most capable engineers I know have this urge to go down rabbit holes and fix any issue, this is nothing special.
Everyone wants to be the hero that found a bug deep in the stack, make a glorious pull request, and be celebrated in the community.
I much more value people who have enough self-control to pick meaningful battles, and follow the right priorities.
Oh what is that you say, security vulnerabilities are also just bugs that get exploited? Oh well...
That is a perfect example of how things works and should work. They contributed to the community. I think it was a great prioritization.
I'm certain there were lots of other people hitting this bug and killing processes or rebooting to get around it. The troubleshooting and reporting done here, silently saved a lot of of other people a lot of efforts - now and in the future. I don't think they were after it to be heroes; they just shared their story, which I'm sure will encourage others to maybe do the same one day.
In the end I believe we struck a good balance between time spent and result achieved: we gathered enough information for someone more familiar with the code to identify and fix the root cause without the need for a reproducer. We could have spent more time trying to patch it ourselves (and to be honest I would probably have gone down that route 10 years ago), but it would be higher risk in terms of both, time invested and patch quality.
Finally, I'm always encouraging our teams to contribute upstream whenever possible, for three reasons:
a) minimizing the delta vs upstream pays off the moment you upgrade without having to rebase a bunch of local patches
b) doing a good write-up and getting feedback on a fix/patch/PR from people who are more familiar with the code will help you understand the problem at hand better (and usually makes you a better engineer)
c) everyone gets to benefit from it, the same way we benefit from patches submitted by others
In my opinion fixing upstream whenever possible even if not the best short-term solution should be considered the price to pay for using OSS.
A good rule of thumb regarding meaningful battles is to ignore everything promoted by companies like Google or Facebook - everything they do is either going to be abandoned in five years, or makes sense only in the context of solving problems nobody else have.
If I'd let every fucking team member go on an exploratory bug hunt whenever they feel like it (hint: that would be always) we would never get anything done.
What if they don't find anything? Is this issue really worth 2 weeks of dev time? That's 15k down the drain for a senior engineer, if not more.
As a user of software, though, I want someone to fix the bug. I want software that doesn't have bugs. So let me repeat my original statement. We need more like this that are willing to spend engineer time fixing bugs, even upstream bugs in open source projects. Instead of prioritizing shoving half-baked features out the door for next week's press release.
- Work around it, most likely creating technical debt inside your organization in the process
- Invest the time to fix it yourself
- Pay someone else to fix it for you (e.g. the original authors via a support contract)
None of these options is for free, and which one is the most cost-effective depends largely on the complexity of the issue at hand, the skillset and availability of the people involved and the criticality of the impacted system.
It is not unheard of to have 4-day weeks and developer-first mindset at that place.
https://wiki.archlinux.org/title/Create_root_filesystem_snap...
0) make sure the the database data volume is on lvm or zfs
in a sql prompt:
1) BACKUP STAGE START; BACKUP STAGE BLOCK_COMMIT;
2) \! the shell command to take the snapshot
3) BACKUP STAGE END;
you can now mount your snapshot, copy it offsite and delete it. The restore procedure is left as an exercise!This requires the cooperation of the dbms software to get the on-disk data quiesced. Then your snapshot has to go fast enough that the dbms doesn't end up with too many spinning plates before you let it start writing normally.
I don’t know what sort of data these people process, but most datasets about people are not anonymized by simply removing the PII.
Once all the PII is removed, by definition the dataset is anonymized.
Take for example a database over all mobile phone positions over time, this can be 'anonymized' by removing all connections from the phones to information on who owns the phones.
But it can still be trivially deanonymized by analyzing where the phones are at night and during office hours, not very many persons work in the same building and sleep in the same house.
Maybe TCP stacks are one of the few cases where that make sense, but I'd suspect if it was "worth the cost" it would have already been done
A formal specification is less ambiguous than a prose specification. Formalizing the TCP specification will, if anything, expose aspects where the specification is unclear, or corner cases where the specification actually leads to unwanted behavior and doesn’t provide the desired guarantees.
So, while you can’t prove that the formal specification matches the prose specification a 100%, you can prove that it provides all the guarantees the original prose specification was aiming for (once you’ve formalized those desired guarantees), which is something you can’t do for the prose specification.
That said, my quick search shows some academic efforts to formally verify QUIC, both in whole and in parts.
I would hope that bespoke (boutique?) TCP replacements, like Homa (specifically for datacenters), are verified as part of the design process. From a quick scan, I gleaned that Homa, and other aspirants, are simulated, compared, and benchmarked against each other. Maybe that's sufficient.
https://homa-transport.atlassian.net/wiki/spaces/HOMA/overvi...
UC Berkeley was a hardware company? TIL.
Companies like Sun had thousands of engineers working on their operating systems, and they very much could and did find and fix obscure bugs, both on their own as well as based on customer bug reports. Some customers did have access to source code -- I know because I've seen that myself. And customers that didn't have access to source code could still do clever things to diagnose problems. In this case, for example, a packet trace should be enough to diagnose the nature of the bug, though one would indeed need source code to write a fix.
Of course it's much better for customers facing obscure bugs to have access to the source code.
I have no idea how to "hot-patch" a C++ application though, are there libraries for this?
Speeding Up Our Build Pipelines - https://news.ycombinator.com/item?id=20775297 - Aug 2019 (24 comments)
The infrastructure behind one of the most popular sites in Greece - https://news.ycombinator.com/item?id=9982361 - July 2015 (5 comments)
Working with the ELK stack - https://news.ycombinator.com/item?id=9008119 - Feb 2015 (35 comments)
There have been some sporadic posts from Skroutz in the past, but nothing that gained so much attention.
For those that don't know it, Skroutz is the biggest Greek online price aggregator/e-commerce market/price comparison site.
5.15.32 was around the kernel where i noticed the issues start. If i'm just connected via SSH and streaming video from a LAN server, everything is great. if i go on youtube.com (or whatever), i'll get "network unreachable" on ping within a minute. I swapped NICs to make sure my NIC wasn't the issue; now youtube doesn't cause this issue, but i tested rsync oddly enough and the NIC goes AWOL after a few gigabytes of transfer. I have to physically unplug and replug the NIC (or a reboot if it was PCI).
I haven't had time to track down why, but it has stayed with newer kernels, too: 5.15.41, 5.15.59 also have this issue. I compiled 5.15.72 last night but i haven't rebooted yet.
Any idea how we might check that?
>I haven't had time to track down why, but it has stayed with newer kernels, too: 5.15.41, 5.15.59 also have this issue. I compiled 5.15.72 last night but i haven't rebooted yet.
I'm fairly certain we've seen it on `Linux 5.15.0-1017-aws x86_64`.
It's because the fools responsible never rewrite their code, use a broken language, and don't even try to prove half of the broken garbage they write. Then, when it turns out to have been broken for decades, they chuckle and shove another finger into another crack, never understanding how they misuse computers.
The C language makes it unreasonably difficult to write anything, even before proving it to be correct.
LOL
You needn't use your real name, of course, but for HN to be a community, users need some identity for other users to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...
Also, could you please stop posting unsubstantive and/or snarky and/or flamebait comments? It's not what this site is for, and it destroys what it is for. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.