Debugging memory corruption: who the hell writes “2” into my stack? (2016)
blog.unity.com
blog.unity.com
Still an amazing debugging story, though :)
I now looked up "mystical" in Merriam-Webster: "having a spiritual meaning [..] that is neither apparent to the senses nor obvious to the intelligence"
That sounds like a pretty good description of my relationship with Unity's support engineers. I hope they exist, but to me it's more of a religious thing because I have never met one.
"mythical" says: "existing only in the imagination: fictitious, imaginary, having qualities suitable to myth: legendary".
That also fits well, because so far they exist only in my imagination but reading this story, I am impressed by their legendary problem-solving power.
Thanks to you, today I learned the difference between two similar words :) And now I can even more clearly articulate someone's progress on being less useful ^_^
When I first met some of the Unity people at the Nordic Game Jam in 2012, they seemed like an awesome team making game development accessible to everyone. At that time, UE was still closed source and out of every indie's reach. And Unity was Mac-only. Since then, Unity has become the de-factor platform for cheap outsourced mobile games, reskin spam, and app store gambling scams. They now have millions of developers using their game engine, but it's an army of people building 1€ apps. Obviously, you can't charge them much because their revenue is so low. And so they had to make support cheap to compensate.
But the part which I don't get is why Unity doesn't offer paid support for medium-sized indie studios. And why they don't publicly sell source code access. I mean their biggest competitor UE4/5 already has all source code on GitHub.
The company I worked for had Unity source access, but we were required to submit any source changes back to Unity (even if they chose not to integrate them), and announce our game at UnityCon 2018.
They also wanted us to help write some of the cutscene timeline tools. I don't know if they used any of the code we wrote, but we definitely did a bunch of work on those timeline tools and sent them back to them.
https://en.wikipedia.org/wiki/De_facto
Post left in the spirit that the OP was a typo, but that someone else might not know and be interested in the Latin roots of this phrase.
The timeline is a bit off here. Unity was initially Mac-only, but only for a short time. It ran on Windows from version 2.5 in 2009.[1]
[1]https://blog.unity.com/technology/unity-25-for-mac-and-windo...
Not taking anything away from the great debugging, but that is just incredibly bad design and reeks of bugs from a mile away. System calls block for a reason. If you can't handle the blocking, how about a non-blocking architecture? But in this case, the engineer thought "hey, I don't want to block on my precious main thread, so let's just give this blocking thing to another thread, that'll surely solve it!". Except it doesn't when you use that same thread as a worker to execute all other thread's blocking crap ... but then write the queue push() function to be blocking as well!
Keep it simple folks. Bad design always breeds more bad design.
Which, admittedly, sometimes folks just don't know, but usually I more often run into complete cluelesness that select()/poll() and friends exist.
"Non-blocking" sounds very special but for anything except batch programs it should be used by default because whenever there is more than one possible input (inputs can be trivial as a cancellation signal) one can not afford to be blocked.
Then again, that doesn't mean that one shouldn't use select(). Unless one goes ultra-high frequency where the CPU is busy almost 100% of the time, select() saves CPU cycles, and decreases response time. Without select(), finding an appropriate polling frequency is a tradeoff between wasted compute cycles and increased response time.
That's exactly what the author of the article did.
This design is pretty much the same that you can e.g. find in the go runtime (netpoll), the .NET core runtime, Rust multithreaded async libraries, Java nio2, etc.
The only thing that was special here is that they used an APC for waking up the reactor in case new events had to be added. Its non-standard, and kind of comparable to using a signal on unix to interrupt `select` instead of a message on a preregistered socket.
Microsoft (at least used to) take developers very seriously. The OS actually plays along when you're trying to debug things, and Windows error reporting actually sends you data that you can work with.
Say what you want about their business practices, but I'd deal with Microsoft's developer support over Apple's anytime. And mandatory Youtube: https://www.youtube.com/watch?v=Vhh_GeBPOhs
You have Event Viewer, Windbg built-in.
Windbg Preview blows away other debugging tools.
And then there's freeware "API Monitor". This thing is incredible, it's like a GUI for looking at all syscalls and COM events on every program without needing to explicitly debug.
The closest thing I could approximate it to would be something like BPF
I’m not sure about that. I think you downloading the GUI, not the debugger. The symbolic debugger engine is in DbgEng.dll and is included in Windows.
I would opine that a debugger in the usual sense is a software that interacts with both the end user (frontend) and the debugee program (debugging/tracing engine), allowing the user to study, analyze, and manipulate the dynamic execution of said program.
The debugger stacks on e.g. Linux and Windows are as follows:
- Linux: GDB, LLDB (frontend + debugging framework); libbfd, libdwarf (object file/symbol engine); ptrace (debugging/tracing engine)
- Windows: WinDbg, cdb, ntsd, kd (frontend); DbgEng.dll (debugging framework); DbgHelp.dll (object file/symbol engine); DbgUi*, Nt* syscalls, kd stubs (debugging/tracing engine)
I have a question: would you call gdbserver a debugger?
Note that in Windows, the GUI is the only piece missing from the default OS installation. The rest of the stack (DbgEng.dll, DbgHelp.dll, kernel calls, tracing, and more) is shipped with the OS.
> would you call gdbserver a debugger?
I’m not an expert on Linux. But based on this page https://www.tutorialspoint.com/unix_commands/gdbserver.htm probably not. I think it says that thing’s an IPC server which doesn’t do debugging on its own, doesn’t even need symbols, instead relies on the host to implement the actual debugger.
I have a Windows gaming rig that I was getting BSODs on. It took about ten minutes to reconfigure Windows to write full crash dumps, get WinDbg set up, and determine where and what was crashing the system (bad display driver install - DDU and reinstall fixed it). There were plenty of guides as to how to get it set up and how to get it to pull Windows debugging symbols from the Internet. It was completely painless.
I have to keep a well-worn grimoire covered with notes and scribbles on how to get kdump working correctly to do the same sort of work on the Linux machines that I support - assuming that devops even set them up in the right configuration for kdump to work.
Seems like the Windows PCs are more likely to show you a blue screen/crash than a Linux one.
Would that match your experience too?
A few years ago the biggest source of BSODs were the performance-optimised video drivers. That has thankfully improved a lot…
My Windows PC actually seems to have video card problems less often, though I don’t really blame Linux here because they both have nVidia cards.
But I have to hand it to Windows in this realm, despite how I might feel about other parts of it.
Linux has strace, valgrind/kcachegrind, eBPF, etc
And those tools are incredibly powerful, but have a far steeper learning curve and (personal opinion) UX than similar Windows tools.
I've used Windbg Preview to debug code written in niche languages like D and it's been able to sync source code from .d files and let me set breakpoints.
That really blew me away. I know that debug info is standardized in PE/COFF and in Linux with DWARF, but still.
Neat (at least I think so) screenshot of Windbg Preview debugging a D .dll with source synced:
https://media.discordapp.net/attachments/625407836473524246/...
The JS scripting is pretty wild too.
Repo full of them here. Example for generating a mermaid diagram image of callgraph, given a function name:
https://github.com/hugsy/windbg_js_scripts/blob/main/scripts...
(Material from a Defcon workshop by same author):
Pros: no PatchGuard sh*t, complete source available
Cons: your out-of-box experience may vary depending on your distro
Fedora/CentOS has -debuginfo for every binary package, while Debian seems to vary for each package
The other thing is, how can a system call block a thread but somehow that same thread runs other stuff in between that blocking call? Even if that is possible that sounds like a dangerous game to play, and this class of bugs seems natural.
I can imagine the scenario where a thread is cancelled mid system call, but it will be the kernel that will cancel the thread and responsible to know that half way syscalls need to be cleaned up. This way the application does not get system call call leaks? Also if i remember one debugging session correctly, when you have a blocking syscall interrupted you will get a sigabort on linux.
I mean, signal handlers exist on Linux, too ...
> when you have a blocking syscall interrupted you will get a sigabort on linux
Doesn't the syscall just return EINTR?
https://docs.microsoft.com/en-us/windows/win32/api/processth...
Imagine how much he could have achieved in those five days if he had just unfucked obvious bug magnets proactively without first painstakingly proving that they were really manifesting as bugs, and how.
But the code that was supposed to wait until completion was interrupted by an exception, which unrolled the stack beyond the function owning the relevant memory. So the completing async operation writes into memory owned by somebody else now.
This is a general complication with completion based async IO. You could get the same problems without any interrupts, if you just don't wait for async IO to complete in some circumstances.
I know this is not about Rust. Forgive me that I mentally connect this story with async Rust.
* Readiness based: You use `select` (or preferably its successors) to figure out which file descriptor has data available. Then you synchronously read the data, which completes instantly because there is data available. This is the traditional way on Linux, and also the way Rust's async works. This approach doesn't run into any lifetime issues, because you only need to provide the memory (`&mut [u8]`) to a sync call once data is available. This approach works well from sequential streams, like sockets or pipes, but isn't a natural match for random access file IO.
* Completion based: You trigger an IO operation specifying the target memory. The system notifies you when it completes. This approach needs you to keep the memory available until the operation finishes. This is the traditional way on Windows (IO completion ports) and also used by the more recent io_uring on Linux. This one runs into the lifetime issues that caused the bug in the linked article. Thus in Rust it can't safely be used with a `&mut [u8]`, you have to transfer ownership of the buffer to the IO library, so it can keep it alive until the operation completes. The ringbahn crate is an example of this approach in Rust.
I find this article helpful in clearing up confusion of asynchronous vs non-blocking I/O: http://blog.omega-prime.co.uk/?p=155
Anything that obscures the flow of control is a violation of literate programming principles, and exceptions are pretty high on that particular shit list.
Yep, it's called Application Verifier and it will check all kinds of system calls and make the memory allocator incredibly aggro towards checking validity
> The other thing is, how can a system call block a thread but somehow that same thread runs other stuff in between that blocking call?
When you call select(), your thread is in a special state called Alertable Wait; one way to get out of it is to have an APC queued to your thread, which has a similar vibe to a UNIX signal (with all of the similar caveats about how easy it is to fuck things up in it). Basically, OP's scenario is similar to corruption caused by a badly written signal handler
Ha, I have always been using this trick to break from a blocking select() call. I feel less evil now.
On Linux and macOS you could also use a pipe instead since select() doesn't only work with sockets.
On Windows you can alternatively use WSAEventSelect (https://docs.microsoft.com/en-us/windows/win32/api/winsock2/...) on all sockets and have another Event just for signalling, but the downside is that it forces your sockets to be non-blocking.
Windows overlapped (=async) I/O to a stack allocated struct is pretty certain to blow up just like that, sooner or later.
If you guard for it, worse, since canceling is asynchronous as well, at least in theory you could end up in a scenario you just can't reliably ever let the execution to continue again. Luckily this should be extremely rare... I hope.
Much better idea: allocate async data structures on the heap. In the worst case you leak some memory. Just never deallocate until the I/O has actually completed.
Variation of the same can happen on the hardware level as well. Don't DMA into stack, unless you can really be sure it'll complete fast enough or you can afford to wait for it.
The underlying problem is that you 1) queue an asynchronous callback and reserve some memory for it to use 2) wait (or so you think) for the function that calls the callback to be done and return 3) something interrupts the wait, but the callback is still queued for execution sometime in the future. 4) since the wait is over, you assume the callback was called, and the memory it was going to use can be safely freed 5) if the memory was on the stack, it gets freed purely by returning. if it was in the heap, you'd deallocate it. 6) the callback that was still queued fires up and writes to memory it no longer should be able to use.
To really fix this situation, you'd need to free the memory in the callback. This is not possible in this case unless you can rewrite the library select() function significantly, perhaps even rewriting other parts (not familiar enough with those APIs to be really sure).
Even in similar situations where it is possible (e.g. if you have garbage collection), it will often turn a memory corruption into a logical corruption, because you have a callback fired related to something that is not relevant anymore, and any side effects of that callback will happen without the proper context. (in this case the only side effect is probably the memory access, so in this particular instance the problem would be "fixed")
[Edit: if the heap allocation was exception-aware (e.g. RAII that gets called during exception unwinding, then you would get heap corruption. If it was manual deallocation and thus skipped during exception unwind, then indeed you would just get a memory leak exactly as you said]
If there's no other option, I'd rather leak that memory.
The actual problem was that the rug was being whipped from under it by an exception being injected into a place where it ought to have been impossible. Your suggestion of using heap allocation instead wouldn't have fixed anything; it would've turned the crash into a memory leak, which would've just masked the the problem.
The solution they went for in the end addressed this real issue, and the async IO (happening under the hood of that Windows API function) remains unchanged.
Of course exception in the queued APC was the reason why select's WFSO was prematurely interrupted and stack rewound, so the bug was of course there. It's pretty dangerous to execute code in the context of a waiting thread.
While the solution to fix it by using heap allocation is a bit ugly in some sense, if your choices are to abort execution completely or to infrequently leak some bytes of memory, I know I'd choose to leak and perhaps log what just happened.
There are rare cases when the I/O never completes and CancelIo doesn't complete either. Probably hardware faults or kernel driver bugs. Stuff that should absolutely never happen, but can still be defensively programmed around.
<rant>
Let's just ignore not using overlapped I/O and completion ports for a moment. How has using APCs even come as a good idea? Like, perhaps a dedicated thread pool for async jobs would work a little better?
If you need to interrupt that specific thread though, sorry for that. Perhaps uhh CancelSynchronousIo() would work since there's IO_STATUS_BLOCK? Or maybe call CancelIo() on the socket in that same APC. However it goes, throwing exceptions in an APC isn't something I'd like to see in production code.
Also, even in Unix world, the select(2) system call is infamously known as a buffer overflow landmine [1]. Just FD_SET any fd >= FD_SETSIZE and BOOM. If you're on Unix, at least use poll() or better. If you're on Windows (and can't afford switching to proper async I/O), maybe WSAEventSelect() would work better?
</rant>
[1]: This does not apply to Windows Sockets though, where fd_set is not a bit array but internally just an FD array in disguise. You get FD_SETSIZE = 64 though, which is conspiciously close to MAXIMUM_WAIT_OBJECTS in WaitForMultupleObjects.
Good diagnosis though. I think I would have given up at the kernel debugger and just refactored that entire module until it went away. I had no idea kernel debugging was even possible.
That part wasn't done by them though, but by select(). Still a good lesson why exceptions should only be used if you know exactly how everything that allocates any resources in between works. And sometimes you can't know until it bites your like this.
On another note, getting old isn't that bad. When I saw the title I knew I read this post before but couldn't remember how it went, so the second read was just as exciting. ;)
They probably picked the select based solution since its decently portable between windows and unix, and good enough for their requirements.
In the same way you don't buy Unreal Engine and then expect to see GLFW being used underneath instead of the native platform APIs or something like that. You expect any big game engine to be thoroughly natively integrated, and so-be-it if they need to maintain multiple native backends. That's their value sell for an engine.
And I think there would be plenty of 32-odd player games using P2P sockets.
https://www.tvfanatic.com/quotes/bender-what-is-it-whoa-what...
Anyway fast forward a bit, I'm working at Convergent which was a nascent Intel's biggest single consumer of x86 chips. Intel was coming out with a shiny new version, the 80286, showed our engineers the spec. They complained "Hey! We're your biggest chip buyer but you didn't consult us on features!"
Intel guy said "Ok, what features do you want?" Our folks said "We'll get back to you..." Then they circulated a questionnaire among us software types for ideas.
I suggested "Let's get rid of that Intel Blue Box, have a feature in the new CPU that traps on bus condition mask and value registers." I even drew out what I wanted.
Well, what do you know, Intel did it. And it remains to this day, and is called "Data breakpoints". Probably the only thing I ever did that will last any length of time, since software has a really short sell-by date as a rule.
When a new socket is added, epoll() gets triggered (by the loopback epoll file descriptor that was used to add the new socket, since it's also actively being polled by epoll). epoll() immediately returns to userspace, the event is handled (to do necessary socket bookkeeping) then go back to polling again with the updated list of sockets.
As a personal example, my team sees me as the consult for all the arcane fuckery that you find as a result of unusual C++ behaviour or build system issues. This doesn't make me a good dev but it does mean that I've been through my paces there and bled my share. But at the same time I'm practically a blind, helpless child wrt most of the JVM or web tech stacks. IMHO my colleagues are wizards for being able to make heads or tails of issues in those spaces.
Point being, this isn't a matter of "Oh I could solve that in seconds", it's a matter of "Oh god I remember dealing with this type of issue. Here's the cause and how to fix & avoid it".
This. Exactly right.
There are thousands of geeks reading an interesting debug report. It is not very surprising to find one who'd think of the right idea as the first thing.
I remember the time that I "solved" a slow database query that was annoying everybody for weeks just by creating an index. Everybody was avoiding the issue like the plague, probably out of fear that it was something serious, but after I created the index (something all of them could do in their sleep), one-uppers showed up to 1) claim that "it was obviously a bad index" and 2) that the index I created could be done better.
My point is that claiming "I knew it" adds absolutely nothing. Claiming that it could be solved in 5 minutes or in a better way, adds even less, turning a moment of deserved proudness into fuel for your imposter syndrome. What really adds is fixing the bug and write a beautiful post explaining the process.
Maybe unpopular opinion, but don't throw exceptions at all, anywhere. It isn't worth it.
Reading low level debugging war stories like this it always amazes me that computers work at all let alone as well as they do. It’s surely due to people like this carrying the torch for the rest of us.
[0] https://blog.didierstevens.com/2010/10/17/setdllcharacterist...
A non-C++ programmer might find this weird.
It's not like you don't have cores/threads for blocking when you just have one server, how many servers are they going to connect to?
Or is this for the Unity server part?
It fits in nicely with how everything else works in your game loop, and means you don't need to deal with marshalling data to/from a dedicated thread.
The age is relative, sockets are from the 70s, non-blocking is 2000+ and in most languages it's only stable since 2010+.
No, the client would have one thread for everything, not one thread for networking. You shouldn't do blocking IO if you only have one thread that also needs to do other things.
Select returns the list of non-blocking sockets that have data read/write pending.
BTW, the sockets themselves don't even have to be non-blocking.
Also, I'd like to stress that the concepts or blocking/non-blocking I/O and synchronous/asynchronous I/O are really orthogonal. You can do synchronous networking with non-blocking sockets and vice versa.
There is no practical difference between blocking on select() for multiple sockets or blocking on recv() for a single socket - in both cases the thread can't do anything else.
Some platforms don't have these luxuries.
I believe this comes from the fact that I am presented with a new _kind_ of bug: async syscall corrupting memory because the stack was unrolled (by an exception). And I integrate this is my mental debug checklist.
We were good about doing code reviews, stacks weren't overflowing, etc. So it was puzzling. Finally, just like the article said, I figured the only way to find it was to catch it "red handed", in the act.
The good news is that memory locations getting corrupted were always the same.
Long story short, I set up a FIQ [1] -- some of you the FIQ -- which would check the location each interrup. I forget if it checked "for" a value or that it "wasn't" an expected value, ugh, sorry... If the FIQ detected corruption, it did a while (1) that would trigger a breakpoint in the emulator. Then I'd be able to look at the task ID -- we were running Micrium u/C OS-II as I recall -- the call stack, etc.
Originally I set up a timer at 1 MHz to trigger the FIQ, but the overhead of going in & out of the ISR 1 million times per second, at essentially a 10 MHz rate, brought the processor to its knees.
So I slowed the timer interrupt down to 100 kHz (!!), which still soaked up a lot of the CPU slack that we'd been running with. And time after time I'd hit the breakpoint in the FIQ, but the damage had been done usecs earlier and the breadcrumbs didn't finger a victim.
Then it happened. Remember, the hardware timer is running completely asynchronously with respect to the application. Finally, the FIQ timer ISR had interrupted some task's code in exactly the function, at exactly the place (maybe a couple instructions later) where the corruption had occurred.
Took about a day start to finish, I'd never seen or heard of using a high speed timer to try to "catch memory corruption in the act", but as they say, necessity is mother of invention.
And to non-embedded developers, this is an embedded CPU. No MMU or MPU, etc. just a flat, wild-west open memory map. Read or write whatever you want. Literally every part of the code was suspect.
Good times.
[1] On ARM 7/9, maybe 11, I think also Cortex R -- the Fast Interrupt Request, or FIQ, uses banked registers and doesn't stack anything on entry -- so it's the lowest-latency, lowest overhead ISR you can have. But you can only have one FIQ I believe, so you have to use it judiciously.
> the socket polling thread then dequeues these requests one by one, calls select()
The select() API is not good regardless on the platform.
If you like the semantics and developing for Linux, use poll() instead. It does exactly the same thing, but the API is good.
Windows API supports thread pools, it's usually a good idea to use StartThreadpoolIo instead.
Thread pool API doesn't have this class of bugs. One doesn't need to wake up any threads to change the set of handles being polled.
Windows only writes to the location because the program gave it that pointer. It's the program's fault that they gave a pointer to an async API when they didn't first guarantee that said pointer would be valid in the future. Really there should be tooling for catching this kind of thing (there is on Linux but less so on Windows).
They are great though, I used use a Lauterbach occasionally back when I work for a mobile OS company.