32 MiB Working Sets on a 64 GiB machine
randomascii.wordpress.com
randomascii.wordpress.com
The Windows BACKGROUND mode is useful for stuff like virus scanning or database log compaction, where you want to make some progress, but only if you can be certain not to hurt any foreground workloads. Back when we had hard drives, this wasn't actually trivial to achieve, because it is very easy to interfere with the foreground workload just by doing a few disk seeks every now and then.
I agree the 32MiB working setlimit is somewhat arbitrary and should be documented, but Windows is full of these arbitrary constants, like the 32kiB paging chunk for instance.
My recommendation for Chrome would be to stick with the background mode, and fix whatever problem is causing the working set to exceed 32MiB.
Working left to right means you have to force the compiler to always align the code in the same order, which sounds difficult and like a waste of time.
If you only make the output go left to right, you don't actually save much memory. You're still randomly accessing a huge buffer.
But you are hurting the overall system performance, or battery life if this is a laptop and you keep thrashing the working set for no good reason. I don't see how this proactive approach has any advantages over just swapping out the process first if you run into actual memory pressure, or maybe start paging out pages of that process if they've been idle for a few minutes if you insist on some kind if proactive measure.
Had the background process been allowed to keep growing its working set then foreground applications would be forced to page out, defeating the goal of having a background mode in the first place.
Without background mode setup.exe would have a 48 MiB working set, for a short period of time.
With background mode setup.exe has a 32 MiB working set, plus 16 MiB of its memory is bouncing back and forth between the working set and the standby list.
Thus, none of its memory will get paged out and foreground apps will get paged out, and this will continue for 250 times as long.
If all you care about is getting your updater to complete quicker, you should not be using the background mode, but ideally the Chrome updater would run in the background without affecting the performance of the foreground applications. Allowing for working set peaks, no matter how short, will fail to achieve that on systems experiencing memory pressure.
Counting on CPU starvation to save memory seems like a very fragile solution.
I agree that commit peaks can affect foreground performance. I've filed bugs for an ephemeral 480 MiB bug in Chrome (caused by an errant product image), and I blogged about an errant 4 GiB allocation that wiped out the disk cache on my 8 GiB machine (https://randomascii.wordpress.com/2012/09/04/windows-slowdow...).
But treating working-set peaks as synonymous with commit peaks is problematic. I still believe that trimming of working sets will rarely save memory, and a working-set cap will virtually never be better than occasional trimming. The cap fails to address the higher-than-the-cap memory consumption in many cases, it just makes it more expensive.
And, as said, draining your battery for no good reason.
> Had the background process been allowed to keep growing its working set then foreground applications would be forced to page out
Why would they? As said, you could still page out the processes running in background mode first if you actually start running out of memory. And if after trimming all the background processes to 32mb you still don't have enough memory, you would be in the same situation either way.
The current implementation has absolutely no advantage.
This isn't even about the exact limit, this implementation just makes no sense. Add a hysteresis at least because running this on every page fault is guaranteed to cause a lot of CPU load while not returning very much to the system.
Not sure if that's why the Chrome installer uses 32MB+.
For example it does pretty interesting diffing - Chrome downloads a patch and uses the previously left-behind copy of the installer to re-build new binaries from the small diff it pulled. See more:
- https://chromium.googlesource.com/chromium/src/+/HEAD/compon...
- https://www.chromium.org/developers/design-documents/softwar...
Just a callstack can be 1MB easily, granted it's RESERVED most of it, just some parts commited, but it can grow there (like directory recursion). Then you have 3 or 4 ThreadPool threads created just like that, etc. etc.
I assume you’re referring to the 64KiB allocation granularity? That’s actually not at all arbitrary: https://devblogs.microsoft.com/oldnewthing/20031008-00/?p=42...
> So if allocation granularity were finer than 64KB, a DLL that got relocated in memory would require two fixups per relocatable address: one to the upper 16 bits and one to the lower 16 bits.
…
> Forcing memory allocations at 64KB granularity solves all these problems.
It does, at the cost of silly limitations. A much nicer solution IMO would have been a way to ask for a 64kB aligned allocation for DLLs.
With 64 bits, my real objection is that it’s unnecessarily complicated. There’s some data structure that maps addresses to whatever is mapped at those addresses, and on Windows it also needs to track reserved memory, and it needs to do so at a different granularity. (For huge pages, additional granularities are needed.)
And I would argue that the 32 MiB is worse than arbitrary - it is pointless. I have been thinking about this for a while and I cannot think of a situation where it makes things better. It wastes (a lot of) CPU, and I claim that it doesn't actually save memory compared to letting the balance-set manager do the trimming.
Low CPU priority and low IO priority are important for being a polite background process. Not using too much memory is important. But a working-set cap fails to achieve that last goal.
In what scenario would be 32 MiB working-set limit, enforced when pages are swapped in, be more effective at saving memory than a similar limit enforced by the once-per-second balance-set-manager?
Obviously trimming the working set at swap time does a better job of reducing the working set, but that's not what matters. What matters is saving memory. If you are trying to save memory then trimming once a second is just as good, and far more efficient, than trimming on every page fault.
But, trimming at all on a 64 GiB machine with 47 GiB free is just silly. It doesn't make sense to spend lots of an expensive resource (CPU time == electricity == battery life) in order to save a resource which you aren't even fully using.
Something that only runs once per second is not going to be able to keep up with with a process that is dirtying memory at full tilt, which is probably why they chose to enforce a hard limit by evicting the least-used pages to the standby list, from where one can usually get them back quickly without encountering a hard fault to disk (a soft fault was probably like 10,000x faster than a hard fault back when they did this work). They must have viewed thrashing as a pathological case that programmers would diagnose and repair, just like you have. After all they do provide some pretty good tools for diagnosing memory use.
You may argue that the 32MiB limits needs adjusting to follow Moore's law, but Moore's law stopped working for laptop DRAM sizes many years ago.
But does the working-set cap _work_? In that scenario where you've got heavy memory pressure I still think it doesn't.
If the background process is typically touching less than 32 MiB in a second then a per-second trimming could reduce the working-set effectively.
If the background process is typically touching more than 32 MiB in a second then the fault-time trimming doesn't work because while it trims the memory from the working set, there is no time for the memory to be paged out, and if the memory did get paged out it would make it even worse because it would need to be paged in. So, CPU (and perhaps disk) overhead is increased, but memory pressure remains the same.
The problem (a process touching too much memory) is real. However the solution does not work. A working-set cap doesn't actually reduce memory pressure on foreground applications any more than a per-second trim.
You are overlooking that the read-only part of the working set can be released immediately once another process needs those pages -- just zero them and add them to the free list. Only the dirty pages need to get written out to disk.
Anyway, it would actually be simple to test the effectiveness of the working set cap by writing a program that allocs and dirties memory aggressively, and then running it either in normal low priority or background modes, to see how it affects overall system behavior when running in each mode.
As for saving memory, I see your point, but I still don't see how the cap would be better than per-second trimming. If the clean-then-zeroed page is not touched again by the background process then either method would make it available. If the clean-then-zeroed page is touched again by the background process then the whole process of removing it from the working set, zeroing it, then reading it back in and faulting back in is a waste of time. So, again, I'm struggling to find a scenario where the cap is more effective than once-per-second trimming.
> Trimming the working set of a process doesn’t actually save memory. It just moves the memory from the working set of the process to the standby list. Then, if the system is under memory pressure the pages in the standby list are eligible to be compressed, or discarded (if unmodified and backed by a file), or written to the page file. But “eligible” is doing a lot of heavy lifting in that sentence. The OS doesn’t immediately do anything with the page, generally speaking. And, if the system has gobs of free and available memory then it may never do anything with the page, making the trimming pointless. The memory isn’t “saved”, it’s just moved from one list to another. It’s the digital equivalent of paper shuffling.
I'd always been under the impression that as soon as memory was trimmed from the working set. Perhaps this was the case at some point, and was a reason for the PROCESS_MODE_BACKGROUND_BEGIN priority? As the blog mentions, the SetPriorityClass call has had this behavior since at least 2015, though I wouldn't be surprised if this behavior has existed for much longer.
As for why this "bug" hasn't been fixed, my guess is that it's due to a couple of factors:
- Windows has become fairly good over the years at keeping the core UI responsive even when the system is under heavy load.
- There are plenty of ways to reduce memory/CPU usage that don't involve a call to SetPriorityClass. I'd wager that setting a process's priority class is not the first thing that would come to mind.
- As a result of the previous two points, the actual number of programs using that call is quite small. I'd actually be interested in knowing what, if any, parts of Windows use it.
(As a side note, if there was a bug in a Windows API function, how would you even report it?
I would say that's debatable. There are fewer complete freezes that require a restart, but things like the task manager (can't get much more core than that), which used to be instantly available and responsive if the machine was recoverable at all, now can take tens of seconds to show up and and respond to interactions under heavy load.
I’d love to learn more about ETW and your videos seem like a good place to start. If anybody else has other recommendations, please share!
And the videos I’m talking about are linked from here:
https://randomascii.wordpress.com/2014/08/19/etw-training-vi...
(x86 has had “accessed” tracking for a long long time. What would be wrong with keeping over-the-working-set pages present but not “accessed” and tracking accesses without causing page faults? Or at least only trying to expire pages when there’s some degree of memory pressure.)
This sounds like it should be basically free in a memory-unconstrained system - is the system really just spending all its time managing these lists? Why is it so expensive?
Even though, then you end up with confused reports about "Linux eating all the memory" and "an idle system using all the memory (for caching)".
AFAIK, newer kernels also have a preemptive swap-out of idle pages so they can evict the pages quicker, so an idle system might even be swapping (out). This results in even more confusion.
I find these reports bewildering (alongside CPU usage). You would hope that something would be exploiting your expensive machine to the fullest (so long as the work being done is useful/intentional/desired - obviously not wasteful page thrashing).
Outside cases where you want to optimize for sleep performance, you probably don’t want to do this though because NUMA systems have conflicting requirements for storage (but thankfully for now the systems with NUMA and the systems that benefit from turning off RAM refresh when sleeping has 0 overlap). Consumer desktops are probably not worth doing this on and laptops probably isn’t a huge difference because of battery size / usage patterns.
If, say, you have 256GB RAM and a 100GB folder of ~1GB files, you will only ever have a few GB used actively. A first pass of processing over the folder will take a long time (reading from disk). Subsequent passes will be much faster, though, because reading is done from the RAM-cached versions of the files (the output from the previous run, was my understanding).
But, (undocumented) rules gotta be followed, so trim the working set it is.
If you're accessing a lot of pages, mostly randomly, it's going to be a mess. If you're decompressing files (as an installer might do), and the compression window is large relative to the limit (as it might be to get the highest compression ratio), that could be a lot of mostly random access to pages, causing a lot of page faults, and high cpu.
"This issue has been known for eight years, on many versions of Windows, and it still hasn’t been corrected or even documented. I hope that changes now."
But ultimately as a non-paid tester of Windows I'm under no obligation to report non-security bugs in any particular way. I like reporting them through twitter and blog posts.
Maybe it wasn't a bug (hah!) but they could have come back to me and at least told me they'd looked at it and it wasn't. I'd have appreciated that. But no.
But the feedback hub is an abyss of random complaints to filter through. Sometimes we'd see items like "Windows wouldn't save my word document and now I hate windows" and being on the networking team were like "uhhh thanks"
An example is the MediaFoundation AAC encoder in Windows 10. It's unusable, due to a bug that randomly introduces oink artifacts into the output. I submitted it to Feedback Hub, but of course it was too niche to get upvotes. Looked around, OBS Studio ran into the same issue and had to switch to a different AAC encoder, so it wasn't a rare problem. Someone even tried posting on Microsoft Answers reproducing it with the Windows SDK sample, and got only a generic response. Finally someone got through, and the product team mentioned that this was the first they'd heard of it... but at least it's actually fixed for Windows 11.
The design of Feedback Hub partly encourages the current behavior. It used to have a very visible categorization on the left side, but now they're easily missable drop-downs, and the auto-suggestions are completely bonkers. You can type up a bug about graphics errors and it suggests the accessibility and first-time install categories. Lack of curation also doesn't help; there's something to be said for having a light touch, but there are posts consisting of just random letters that have been there for years.
(It's on my list to do more for submissions like that which accrue a bunch of upvotes slowly but never break the front page...)
(1) if we didn't do that, commenters would start asking "why are there a bunch of comments here older than the OP they're commenting on?" - and our experience is that this leads to more confusion than the other way around; plus
(2) it seems unfair to have all the comments on the earlier thread plummet to the bottom of the new thread because they're so much older.
This topic comes up regularly in a slightly different context, which is re-upping stories that make it into the second-chance pool (https://news.ycombinator.com/item?id=26998308). There are a bunch of past explanations here in case helpful: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que....
Another improvement from Microsoft without really understanding their own OS.
The Background mode has its merits too, but I guess they never imagined that a background process would need to spend huge amounts of memory just to do some background work.
https://support.microsoft.com/en-us/windows/windows-7-system...
It's not like I have a good sense of what it does anyway, in addition to copying files for a fraction of a second.
So, the only time the bug happens is for background updates that may be happening when the user isn't even present. And, since the CPU priority is low they won't harm responsiveness of foreground applications. But they will waste CPU time and electricity and battery life (if on battery). So, a silent waste.
I have a Ryzen 9 5950X CPU and it has a crazy fan ramp. Anytime a process maxes out the CPU it will be audible to me when I'm working. Now, to put this in perspective. It doesn't happen when I'm gaming. I can play Diablo 4 and Cyberpunk 2077 without this noticable fan noise but when setup.exe loses it, the fan noise is how I notice it. It will max out one core and I'm going to assume "never" complete. The longest I waited was 78 CPU minutes before killing the process. This would happen now and then but it would not prevent Chrome from successfully updating. So, it was bizarre to begin with.
It's the other way around, most computers have ram in gibibytes but list it as gigabytes. It's mostly hard drive and other storage manufacturers that treat gigabytes as actual gigabytes as a way to skimp out
My RAM sticks each have 8589934592 = 2^33 bytes.
If the memory isn't owned by the OS, that's probably the wrong actor to manage it.
They say that, instead, they used an undocumented API they obviously didn't understand. Why? If you are concerned about using up too much memory, then do something about it
Is the fourth line of the article linked.
When people talk about using undocumented APIs in windows, that means secret APIs, not badly documented APIs. That is the only meaning that fits the accusatory tone of the comment.
Isn't the thing you being fair to the last line of mrguyorama's comment? The one that was blaming the user? The thing you replied to was an argument over the appropriateness of that specific line.
Trying to strengthen an argument by finding a way that a word could technically work, but not with the original implications, isn't very helpful. I'd even argue it hurts clear communication in a situation like this.
But in addition to this we got a problematic working-set cap which can absolutely destroy performance. That detail needs to be documented (and the rest of the behavior should be as well).