“Unexplainable” core dump (2011)
stackoverflow.com
stackoverflow.com
It does not encourage me how much this sounds like the short story "Coding Machines". The original post even happened right about 2 years after the short story was posted, then in that comment reoccured after another 2 years.
Though in reality it definitely would've been some guy on the compiler dev team adding it before publishing binaries. I wonder if you could set it up to inject a halt into compiled code that runs if some conditions are met, crashing most of the word's infrastructure on a predetermined date.
(Or…maybe not except?)
Another one, the debugging was aided by the fact the developers ensured that everything was accessed through const pointers, so it wasn't their code corrupting their memory.
Took a full year to figure it out. Good times.
Actually, I remember reporting the bug to the compiler authors and being stunned when they told me that they were not going to issue a new version with the bug fix because the project was no longer being funded. (This was the T dialect of Lisp in case you're wondering.)
What was done in the time period between discovering the bug and its cause?
Fortunately, it only happened when the robot's arm was moving, and we were mostly doing mobility research so we were able to be productive simply by not using the arm.
Also minor point: a const pointer is a pointer which always points at the same address. You can still change what is pointed at. You probably meant "a pointer to const"
People have been using the term "const pointer" to refer to "a pointer to const" for 20+ years (as long as I've been doing C++), although that's probably more out of laziness than incorrectness. Certainly the language definition didn't do anybody favors.
The delta has to be small, because you don't want big adders in the front end and because the delta has to be sent down with the push/pop memory uops. That means it can overflow or underflow, at which point it has to be reset by sending a synchronize operation to the back-end to update the original stack register (agner has a better description).
So the delta register is probably 10-12 bits on Barcelona, and this bug is probably a corner case where the stack register update is happening, hence 1024 bytes off. Perhaps there is a window where a uop can get the old delta value, but the new base value (or vice versa) when a sync is operation is concurrent.
Setting that MSR value possibly disables that stack engine feature entirely, or it could be it disables some aggressive and complicated detail where the bug is, e.g., allowing stack operations to run concurrently while flush operations are in progress.
It's not a coincidence there just happens to exist a way to disable this at runtime. The way processors are designed means that everything must be able to be observed, debugged, and fixed in the field. That means everything has to have fine-grained ability to control, disable, enable safer fallback paths, and even engage additional logic to reduce the state space in some cases (e.g., serialize pipeline while a particular operation occurs).
Usually these bug fixes decrease performance (except in cases where a performance bug is found and the fix actually increases performance), so you want the switches to be very fine-grained. So it's possible they fixed the stack engine bug without disabling it entirely.
It would be like shipping software and providing support and bug fixes for it for the next 5-10 years without patching the software, only updating the config file. It's quite amazing. For every one of these issues that hits the field, there will be many found internally during the internal hardware bring up and verification (which will be ongoing for at least part of the life of the CPU).
During college, we were arranged in groups of 2 or 3 to do some pair programming for the more complicated exercises.
I had attached the debugger and had set a variable to 99, so the loop would execute one more time and we could test if our changes would work.
Went through the instructions step by step, and suddenly we got a segmentation fault. Between setting it to 99 and it accessing some data later on, the value changed to 107.
Quite a bit of confusion ensued. I made a backup of the executable before recompiling. Running them again, both the new version and the backup worked perfectly. The files matched bit for bit. To this day we have no idea what caused that bit flip.
https://devblogs.microsoft.com/oldnewthing/20050412-47/?p=35...
In our case, we built programs that ran enormous numbers of semi-random programs on the accelerator and compared it to reliable results computed offline. About 1 in 1000 chips would - reproducibly - fail certain operations. Identifying this helped solve problems many of our researchers reported on specific accelerator clusters- they would get a Nan in their gradients which would kill training, and it was almost always explainable by a single processor (out of ~thousands) occasionally corrupting a float.
At a previous company we had to run a dll in our web servers (provided by the payment processing company) for PCI compliance reasons, and we later discovered it was messing with our FP flags and as a result serialization code was producing invalid floats. That was a fun one.
Realized that it was only happening when a particular motor was running, then further isolated to when it was only running at higher duty cycles. I had tested the motor in isolation but at low speeds, so I didn't see issues until I tried to run the full application. Turns out the EE had royally screwed up the current sense circuit, and the MCU on the ADC was seeing voltages lower than -1V, well beyond the absolute max ratings for the chip.
I guess ST doesn't make guarantees about what happens when you violate those ratings, but a corrupt stack pointer is not what I would have expected.
If you were on a platform that did this wrong, speculative execution could have been dereferencing the NULL pointer.
I found lots of bugs porting code to it. It was my first 64-bit platform.
If a supplier sent you out-of-spec parts because they didn’t think your spec was actually important, you’d call them fraudulent, not clever.
It's not merely "pretty sure" - the language is specifically defined this way. C++ in particular requires as a minimum standard from practitioners a perfect knowledge of the vast, complex language standard. Anything less and you'll write a program which is ill-formed or has undefined behaviour, ie your program is nonsense.
GCC can be passed `-fno-delete-null-pointer-checks` precisely to prevent this. Linux uses it, see e.g. https://lkml.org/lkml/2018/4/4/601 where it's discussed.
Yeah based on https://joelaro.wordpress.com/2015/09/30/gcc-optimization-fd... the assumption is that if the pointer was previously dereferenced, it can be assumed to be non-null, and hence all subsequent null-checks can be elided.
Based on https://news.ycombinator.com/item?id=17360772, the case in which this actually makes a difference is an address offset like
int *n = ¶m->length
The compiler knows the structure so it can just inline the offset, but the optimizer considers this as being non-null.And address 0 isn't too special on a machine level either. Embedded systems have valid data to read and write there.
E.g. if the architecture has pointers with non-address bits (modes or segments or whatever) and those bits were set yet the rest of the address was 'null', and the check was for 'all bits zero' then you could conceivably get that situation.
See more here: https://stackoverflow.com/users/50617/employed-russian
We were asking git to do something utterly trivial, like clone, and it was segfaulting. We installed the debug symbols or something (as I recall you can install these separately in Debian) and started trying to see where the function was crashing. The crash was in a parser, and the code was doing,
char *foo = strstr(input, "something constant");
and segfaulting later (but not chasing null!), during the first use of foo. I figured that input must, therefore, be a bad pointer, since the string literal is by definition fine, and the only way to get a bad pointer out of strstr was to give it bad input in the first place.So, in gdb, I print the input point. It's valid. "Damn" I think to myself "we've gotten unlucky, and whatever triggers the cascade of UB hasn't happened this run. Freakin' heisenbug." So, I told gdb to just continue running the program: it crashed. Same backtrace: dereferencing foo triggered a segfault, and not with a null pointer.
If the input pointer to strstr is valid, how can the output be a bad pointer? strstr is documented as,
These functions return a pointer to the beginning of the located
substring, or NULL if the substring is not found.
So, what gives?I attempted to debug strstr … that was almost a mistake. strstr is, unfortunately heavily optimized. It might have been this one: https://sourceware.org/git/?p=glibc.git;a=blob;f=sysdeps/x86...
I'm okay with a debugger, but I'm not going the be able to follow where vectorized code is going wrong. Almost stupidly, I told strstr to (f)inish the function's execution, at which point gdb printed the return value: it was a valid pointer!
"Odd. Is this run going to succeed?" I continue execution: it crashes.
I restart the debugging, and run to the completion of strstr: the pointer is correct, but again, crash. How. The segfaulting address isn't the pointer strstr is returning, either, and nothing modifies foo after the strstr.
I single step out of strstr, and immediately print foo. The resulting pointer is wrong. strstr is working fine … but what, the assignment operator is broken? At this point, I'm sure I'm crazy. I disassemble the source, and lo and behold, the disassembly is trivially wrong. Forgive me as I can't amd64 unless I'm looking at it, but the disassembly is something like,
call strstr
cltq
<store foo>
We look the odd instruction up; it sign extends eax into rax. That's … not a valid operation on a supposed char pointer … there's not even technically a value in eax at this point. strstr returned a pointer in rax. eax is just the lower 32 bits of that pointer, and sign extending that back into rax makes no sense. It's like the compiler thought the return from strstr was a signed int … and that's where it hits me: C is shooting us in the foot again.> and if no declaration is visible for this identifier, the identifier is implicitly declared exactly as if. in the innermost block containing the function call. the declaration
> extern int identifier();
> appeared.
The default return for an undeclared function, in C89, is int. (I'm actually not sure what C11 says about this. The C89 rule appears to be gone, but there doesn't seem to be anything in its place, which is bizarre. "Undeclared" does not appear, except for an unrelated footnote.)
Slap in the proper include for strstr, recompile. No segfault, disassembly is correct.
Upgrade to latest version, disassemble: disassembly is correct. Someone else fixed the bug, in the meantime. Should have just tried upgrading in the first place…
(After typing this up: had to dig up the debugging session. It was strchr, not strstr. Potato, potato. Someday … this should be a blog post…)
If you're really lucky, you'll never see one at all!
It lets you build your own turing complete processor, and define a simple assembly language, starting from NAND gates and you can create your own arbitrarily wild edge cases for specific opcode combinations.