Finding Windows HANDLE leaks, in Chromium and others
randomascii.wordpress.com
randomascii.wordpress.com
Also nice how the issues are handled by team members, that know more about stuff in their field. Like, maybe other people wouldn't be skilled enough to find the issue, but he wasn't skilled enough to fix it.
Learning along the way. When tackling issues, there are always some knowns and unknowns and the experience is gained. Thanks to the way randomascii is written, I'v had some success using WPA.
If you want to see some real offenders... have a quick look at Asus's "LightingService.exe" (the daemon that controls their rgb LED coloring suite). Gets up to 2m+ handles after a day or two of running on my system.
The next biggest offender on my system? Their ez updater "EzUpdt.exe" at ~60,000 handles.
Haven't bothered looking at it with process explorer/procmon yet, but I'm sure it's a quick fix.
I honestly didn't even want any lighting control software but all the RAM that had the speed and latency that I wanted were decorated with them so I needed software to set them to a static color. Thankfully I learned about OpenRGB.
Especially for Linux-users where official support may not be available for your RGB HW (mobo, RAM, whatever), OpenRGB is a godsend.
And scriptable too, so I can disable all that crap at boot-time.
There is OpenRGB (among others) that should be able to control lighting on multiple peripherals. And FanControl (among others) for fan curves.
Also the manufacturers probably love it because it pushes people towards not mixing products from different companies.
https://github.com/Microsoft/wil/wiki/RAII-resource-wrappers
That said, I don't know that they would have helped in the article's case. It looks like there was a higher-level resource getting leaked, which would also most likely leak these.
Click the details tab in task manager, right click on any column, select "add columns".
I'm not sure that Electron apps have the same local storage characteristics as the full browser, but it seems feasible.
There's also an Asus Lighting Service that still has 10K handles open.
Now? You watch a couple YouTube videos, or perhaps a terrible Apress book. You get something to appear on the page, and viola!, you're a programmer.
It’s remarkably hard in 2021 for a newbie to start learning Windows programming “properly”.
I still can't name a single UWP app that isnt a crude mockery of some win32 app that runs circles around it in both performance and UI design.
My favorite example of how this new generation of MS devs handle performance: https://github.com/microsoft/terminal/issues/10362
After sorting by handles, I see similar results. Logitech Hub Agent has 13,000 handles. It seems these gaming companies hire the worst coders without care for actual performance. Yet...they advertise their products as cutting edge performance. Quite the irony.
Unfortunately gaming companies routinely hire inexperienced coders because they're cheap. By the time those coders have skilled up, they've been burned by the gaming industry in one of the many ways that the gaming industry burns coders, and so they no longer work in gamedev.
I might change that stance when I get home tonight and check for leaks like this, though.
Handle type summary:
<Unknown type> : 4
<Unknown type> : 2
<Unknown type> : 332
ALPC Port : 19
Desktop : 2
Directory : 4
Event : 395
File : 131
IoCompletion : 55
IRTimer : 11
Job : 3
Key : 104
Mutant : 21
Process : 66
Section : 45
Semaphore : 18109
Thread : 154
TpWorkerFactory : 5
WaitCompletionPacket: 70
WindowStation : 3
Total handles: 19535And, like many commenters below constantly surprised that so many apps still ship with these problems, since it's so easy to spot them. Windows perfmon also gives you a nice graph, so you can correlate e.g. GUI behaviour with a jump in handles being created that then never get freed.
If you leak 6GB of RAM, there's 6GB of evidence about what was leaked exactly. If it's full of terrible love poetry you can rule out "FootBallScores" and focus on the "TeenagePoems" data structure and related code. But if you leak 50 000 HANDLEs then er... oops?
The post shows you can narrow it down to say, Event HANDLES but after that it gets increasingly sticky. Hopefully somewhere there's a C++ object that owns the handle and that has leaked which you can trace, but as I understand it, the handles themselves might be all that leaked, leaving you to instrument software so that you can find out which code made all the handles and then trace back those that seem leaked.
many start using 'auto managed' style languages (c++, java, etc) where the life cycle is not as clear. The life cycle is the same though but you own it in an indirect way. Who makes it. Who uses it. Who destroys it. In some languages that is easier to do, others you own it front to back. When doing this I try to start with create/destroy then usage. It is a style that helps remove leaks before they happen. They can still slip in there...
I have used the common string thing a few times to help narrow leaks (does not work in all cases :( ). You also can use tools like valgrind, boundchecker, purify, etc. Think MS has a couple that I can not remember off the top of my head.
Modern operating systems mostly rule out the latter type of leak, you could leak files I guess, and of course Cloud users could leak things like S3 objects, even whole instances, but many resources are now automatically cleaned up when you exit.
As a result of that though, for long-lived processes, such as Chrome but also most background tasks and server software, just "I definitely clean up the mess eventually" doesn't get the job done, the OS was going to do that too. The user doesn't care whether the resources would have been returned half a second after they closed your program when "clean_up_everything()" is called by the main thread, or, a second after that when the OS cleans everything left behind.
So this is a real problem, about the actual meaning of our programs, and (though they are still a good idea) can't be helped by good programming techniques, garbage collection, Rust's Drop trait, the C++ RAII way of thinking, deferred clean-up in languages like Zig or Python, or anything else I'm aware of.
We need to actually express in our programs the intent to hang on to only what's actually needed and clean everything else up as we go. And it can be sorely tempting to consider that "It gets cleaned up eventually" is good enough, that's where leaks get in.
Mostly these days most machines have a decent amount memory so leaks are not as noticeable, unless you look. If in the early days if I leaked 50MB of memory and my machine had 16MB. I had a real issue and the machine would be borked. If I do the same today you would not notice it.
It is why I stressed watching life cycle of an object. You made this thing, who is cleaning this mess up, and when. 'When' could be anywhere from never 'I need this all the time' to bunch it up when idle/reuse (garbage collection style), or 'right now' I need this memory back right now. There are trade offs and you need to watch for that too. The write it down while you are thinking of it has served me very well over the years. I personally got bit by not following my own rule a few weeks ago. I allocated something and I had not cleaned up correctly. I got 'lucky' and that was actually the right thing to do. But in the code review I rightfully got dinged on it.
One thing I wish many more docs would do is 'this makes object xyz use remove_xyz to clean it up'. Or 'this looks like it is creating an object it is not, this is returning some global'. Right there in the doc. It would help so much.
That temptation of 'eventually' is one that some languages push hard. But I find many leaks that I have chased over the years were just a poor understanding of the calls being used (bad docs, not reading them, or a combo). You may have had a different experience.
Arena allocation. Every allocation must be attributed to some arena (eg current tab, current network request, current frame being rendered, etc); when the arena goes away, so do all its allocations.
There are some patterns to help, like using().
Such non trivial bugs and debug methods.
On the one hand, most distros default to 1024 for the (soft) limit on open fds, which would be easier to hit if you were leaking fds and thus easier to detect.
On the other hand, a lot of things that HANDLEs get used for might end up being userspace pointers from userspace libraries instead, so the leak would be of memory instead of fds and hard to detect again.
More details: as I explain in the first link in the blog post, if a process leaks process handles then roughly 64 KB of memory will be used for each leaked handle (each zombie process) but that memory will not be attributed to any process. If you have the Handles column open you will see a large count in one process, but that process will not have a large memory footprint. If you kill that process you will reclaim the memory from all of the zombies.
Good times.