Linux's strcmp() for the m68k has always been broken
phoronix.com
phoronix.com
We had an issue for years where occasionally our backend database would fall over and we were just seeing strange corruption in the stored data sometimes.
Turns out that memcpy() didn’t handle data that was not aligned to a 4 byte boundary correctly in the 32 on 64 case. It took one of our sysadmins a very very long time to find it.
What blew me away is the idea that there can be a bug in something that is used literally everywhere all the time.
In addition, there was a pull request with a fix that someone had written that got ignored for years that we found. We just applied the diff in the pull request and it fixed our issues. The fix was eventually applied upstream a few years later.
We've become so accustomed to computers being generally unreliable that I guess people see something weird and just reboot. We checksum, run reductant servers, and just generally tolerate and work around way more failures and weirdness than we should.
It's funny because we're theoretically an industry that could be full of maths and proofs but ask anybody outside of it and you'll get "yeah computers just don't work sometimes I guess that's how it is"
And once you start designing around that unreliability, it is harder to justify chasing esoteric bugs rather than folding that alongside the probability model of random hardware failure.
It’s hard for any system dealing with the real world to be completely “maths and proofs”. There is plenty of it that’s applied even in the real world systems we operate today. At the end of the day, it’s Engineering and not a pure science.
And I guess the other problem is that quite often you don’t need to quantify the failure modes and do “proper engineering”, because a lot of the systems we build and use just don’t need the level of reliability that e.g. a bridge needs (yet). There are definitely industries where software reliability is considered paramount, but that’s not the majority of the market.
Sure. As the saying goes, "Anybody can build a bridge that never falls down. But it takes an engineer to build a bridge that _almost_ doesn't fall down" (referring to being able to build it in a cost-conscious way). There's a spectrum there for sure, but as an industry we're on the other side of it from where I personally wish we were.
We should be embarrassed that "it must have been a computer glitch" can be used to explain away nearly any problem. Because it's usually right, and it's usually our fault rather than the cosmic rays.
memmove was broken on 32-bit executables on systems with SSE2 whenever the move occurred over the 2GB memory boundary due to an issue with signed vs unsigned integers.
That is surprising outside of the kernel. There are a large number of programs that would fail very hard very fast with that sort of artificial restriction.
That sounds like an interesting detective story. Did they write it up anywhere?
Core dumps didn’t show anything useful because the crashes were just that a pointer was mysteriously pointing into a bad place. But how did it get corrupted?
We were doing increasingly insane things to reproduce the problem locally.
An interesting idea I had to reproduce it was using a database replica, which did crash sometimes.
We used TCPDump to capture the replication stream, then waited a week for the replica to crash. We could start up a copy of the replica in the state it was at the beginning and replay the week of replication data into the socket to reproduce the problem.
My theory was that since a DB replica doesn’t do anything other than the sequence of operations that come to it via the replication stream, that this should deterministicly reproduce the problem.
Turns out that it didn’t.
Anyway, if my memory is correct, the issue was that memcpy was copying an extra 1-3 bytes at the end of the range, if the start of the range was unaligned. If there happened to be a pointer there then that would lead to a crash down the line. My memory could be off though.
Ahhh I could see how that could be! Fascinating. That's exactly the kind of heisenbug I used to love. I always liked the ones that seemed literally impossible but had a nice explanation.
The stacktrace said the crash came from SSE memcpy. I did not bother investigating it, since I have no users anyways.
I had an really old CPU (like i5-520m). Perhaps the code required some SSE feature, it did not have
That was libc. I also use Pascal, which has its own memcpy implementation. A while ago, they discovered it does not handle alignment correctly. (like it would adjust by 4-i when it should have adjusted by i) I thought that is what I get when using Pascal rather than a more commonly used language, but apparently nothing works
https://www.tuhs.org/cgi-bin/utree.pl?file=V7/usr/src/libc/g...
No. As the article notes, there's also a bug when char is signed (the subtraction may overflow, which is UB for signed char, and additionally produces the wrong result even assuming defined overflow). It's just more obvious when char is unsigned.
-2 > 127 > 64 > -2
This can make sorting algorithms loop infinitely or crash.Signed overflow is undefined in C, but this was implemented in ASM. I'm not familiar with the m68k isa, but I'd be very surprised if it doesn't define signed values to wraparound as is natural for a two's complement implementation.
C has no arithmetic operations on types narrower than int and unsigned int. This bug (which involves subtracting 8-bit values and yielding an 8-bit result) could only happen in assembly language (or in an exotic C implementation where CHAR_BIT > 8 and int is 1 byte, or in C code that uses casts to deliberately create the bug).
The doubly-linked `list` type that's used ~everywhere in the kernel didn't have any in-repo unit tests until 2019 ( https://github.com/torvalds/linux/commit/ea2dd7c0875ed31955c... ), after kunit, a library for userspace testing of kernel code, was added.
(Overall, I actually like this feature. You WOULDN'T BELIEVE the CRAZY Headlines We'd See Otherwise!)
Also I can't even remember the last time I saw a real article with all caps words in the title (on the original page), let alone one linked on HN.
Avoiding all caps for emphasis means you sometimes have to change "Faa" back to "FAA" after submitting a story about the Federal Aviation Administration. And avoiding grammatically-incorrect or all-lowercase headlines means you sometimes have to change "Strcmp()" back to "strcmp()".
HN's software is no longer open source, but at one time, this is how it processed titles on initial submission: https://github.com/wting/hackernews/blob/master/news.arc#L15... (This function does add capitals as well as remove them.)
I know. I don't think you understood what I meant.
I'm saying "What otherwise? If you removed that code, I still don't think those headlines would show up."
But removing caps is easier to justify, so I'm also asking why, even if we decide we need caps-removing code, we would also need the code to add caps.
You have more faith in the average HN submitter than I do. :)
(There also publications that capitalize everything indiscriminately, I think.)
[1] E.g. https://apastyle.apa.org/style-grammar-guidelines/capitaliza...
Then they promptly forgot all about it.
Just like it is hard to proof read your own papers, it is hard to write complete tests for your own code.
There are many, many reasons for maintaining support for alternate architectures. How many problems would be baked in if we always had this attitude?
Back in the day when "all the world is an i386", people complained when told about how their code has bugs when compiled on 64 bit processors. Imagine if we had baked in all of those 32-bit-centric bugs that broke compiling on 64 bit.
People do the same thing now when told their code doesn't compile cleanly on 32 bit, on PowerPC, on ARM, on big endian, et cetera. Imagine if we now acted like "all the world is an amd64", and lots of code was simply broken on aarch64?
If you were alive and paid attention through other transitions, you'd understand how every example of programmers being "forced" to care about the correctness and portability of their code has paid dividends years later.
Though I wasn't making a definitive statement on which category m68k belongs to in the first place.
If anyone wants a further back in an even earlier day example, see http://www.catb.org/jargon/html/V/vaxocentrism.html
I would very much like to learn how to get one 1) running on an FPGA 2) running Linux and then 3) set it up to automatically run new builds.
[1] https://silvaco.com/design-ip/embedded-processors/ [2] https://www.farnell.com/datasheets/1703092.pdf
As far as I'm aware, there are no longer any "true" 680x0 parts in production. NXP stopped production of the 68SEC000 in 2014; anything still in stock is at least that old.
I also read somewhere that the 2011 earthquake+tsunami destroyed the only fab that had still been producing the original (well, CMOS) 68K-family chips, and Freescale had been planning to close that fab anyway.