Negative dentries, 20 years later
lwn.net
lwn.net
A deploy that was normally very fast would sometimes hang for a few minutes during a phase where all it had to do was delete the old application directory and move the new one into place.
Turned out that the application was writing a bunch of tempfiles into the cwd and then immediately deleting them. Nothing ever touched that directory while the negative dentries accumulated for weeks or months. When someone finally deployed, the first rmdir that came along bore the cost of deleting all those negative dentries. It hung for seconds or minutes while the kernel essentially cleared out the entire dcache, deleting linked list elements one by one. It showed up in perf as being stuck inside shrink_dcache_parent.
This is actually easy to reproduce:
$ mkdir /tmp/foo
$ touch /tmp/nodelete
# create and delete 100k files
$ for i in $(seq 1 10); do bash -c 'for i in $(seq 1 10000); do rm $(mktemp /tmp/foo/XXXXXX); done' &; done; wait
...
$ time rmdir /tmp/foo
rmdir: failed to remove '/tmp/foo': Directory not empty
rmdir /tmp/foo 0.00s user 0.02s system 91% cpu 0.024 total
$ time rmdir /tmp/foo
rmdir: failed to remove '/tmp/foo': Directory not empty
rmdir /tmp/foo 0.00s user 0.00s system 81% cpu 0.003 total
Both rmdirs fail, but the first one takes 24ms. If you create and delete more files, it takes longer and longer.At some point we probably would've noticed the memory leak as well (I found an 18 GB slab on one host while this was happening) but the machines in question have huge amounts of ram.
I worked around the issue by making the application reuse tempfile names.
> I worked around the issue by making the application reuse tempfile names.
Knowing nothing about the issue beyond what you've written here...
why not make the application create a directory for its tempfiles, and then remove that directory along with the tempfiles?
Considering it again now, I do think it's essentially a bug, but it seems to be a known thing at this point. What I described is the same issue addressed by this unmerged patch: https://lkml.org/lkml/2017/9/18/739 (see discussion here: https://lwn.net/Articles/814535/). And it's mentioned in the article in this HN link:
> Those dentries still take up valuable memory, and they can create other problems (such as soft lockups) as well.
So, it was over some threshold that the kernel decided it was time to start reclaiming all those entries and locked up the entire server while it was doing so. Madness.
Or is there some reason those aren't acceptable in-kernel, while apparently a lack of cleanup is fine? Maybe CPU use is too high...?
Should that size expand and contract overtime somehow?
The article gets into a little bit of that stuff, but they’re all real problems.
Not me I’m watching Netflix :)
find / -type d -exec sh -c 'for i in $(seq 5); do stat {}/"$(head -c 15 /dev/urandom | base32)" 2> /dev/null; done' \;Has an AIMD algorithms[1] been considered to dynamically adjust the limit?
[1] https://en.wikipedia.org/wiki/Additive_increase/multiplicati...
Lol, <3 LWN
Expose practically anything on the internet that leads to a piece of user input causing a filesystem operation (which is going to be extremely common and often unavoidable), and "hundreds" isn't the concern. Millions to billions is.
In theory, git could do its own caching, or it could be refactored so as to move the .gitignore scanning inline with the main directory traversal. In practice, it doesn't do so (at least as of 2.30.2, which is the version I just tested). And I don't think it's reasonable to expect every program to contort itself to minimize the number of times it looks for nonexistent files, when that's something that can reasonably be delegated to the OS.
Without a truly reliable filesystem watching technique, git cannot cache this information. Any cached knowledge needs to be verified because it could have changed, which defeats the whole purpose of a cache here. It could change its strategy, to e.g. only check at repo root or require a list of "enabled" .gitignores, but it doesn't do that currently.
As it stands, when you run "git status", the git executable goes through each directory in your working copy, lists all of its directory entries, and then immediately afterward calls open() to see if a .gitignore file exists in that directory -- even if there wasn't one in the list it just read.
That's what I meant by saying that in principle, there's no reason git couldn't just cache the file's presence/absence on its own, just for those few microseconds (maybe "cache" was a poor choice of words). In practice, I can understand that the implementation complexity might not be worth it.
What happens when the .gitignore above you changes while you are in the midst of scanning a subdirectory? Is that not the same problem?
The problem is that git operations aren't transactional with respect to the filesystem, no? Sure, the window of uncertainty changes, but it's never zero.
If you remove .gitignore 1 microsecond before git open()s it, it's gone.
If you remove .gitignore 1 microsecond after git open()s it (even before it has a chance to read it), unix file system semantics mean you still get the contents.
There is always a race condition, you can play with the specifics, but cannot avoid it when multiple files referring to each other are edited independently.