Mtime Comparison Considered Harmful (2018)
apenwarr.ca
apenwarr.ca
Note that the link is outdated and refers to HFS+. HFS+ originated in classic Mac OS and originally didn't support hard links at all, so they had to be added after the fact [1] using a hacky format based on a hidden directory. APFS has a cleaner design, but it still tracks the list of hard links to an inode. [2]
[1] http://www.wsanchez.net/papers/USENIX_2000/
[2] https://developer.apple.com/support/downloads/Apple-File-Sys..., search for "Siblings"
It might be partly due to hardlinks, but I suspect it is mostly an efficiency issue. Modifying a file could mean updating quite a few directories every time if you have a deep tree. For some uses this could be quite a performance hit, and for some media would potentially reduce the life of your hardware too.
IIRC even when updating atime for every access was the usual default for many filesystems, only the current object had its time updated not all its chain of parents.
Production builds should be a full build.
It is customary to have various optimizations in modify/save/rebuild/execute/verify loop.
I, for one, don't even restart the app for most changes and use REPL to modify application at run time. I only rebuild it from scratch and restart from time to time to verify it all still works as intended.
Here is Shaun Mahood explaining this while building an application to do the presentation without even having to reload the page: https://www.youtube.com/watch?v=cDzjlx6otCU
Depending on how long it takes to do a full build versus an incremental build, doing incremental builds for releases can make a massive improvement in developer experience. Just speaking from my own experiences here. Large C++ codebases in particular suffer from long build times.
Methods that can help you achieve that goal faster are more valuable than ones that are completely correct but require more time to execute.
Most massive improvement in developer experience is if you can get feedback immediately without lapse in your concentration.
Production build for me is something that happens in the background after I have pushed my code and moved to work on something else and so it has very little effect on my productivity. It is nice having it complete in seconds but it does not hurt my productivity significantly even if it takes a whole day, as long as I can continue getting feedback on my current coding task (ie. I have working local environment where I can fully test with little overhead0.
At the most extreme example, I have worked on a credit card terminal where you needed a whole week to get the application built and deployed on a real device because it had to be built by a special compiler on another host and then once it builds it had to get approvals and get signed by external company.
It of course did not affect my productivity in the long run as I found ways to quickly compile and run the functionality outside the real device.
Isn't POSIX fun?
I remember discovering this little fact a lot of years ago. It really shook my confidence in the existence of a rational world. Then I learned about noatime and never looked back.
But your SQL server only has two modes, the whole database is read-only or the whole database is read-write. You really wish there was a soft read-only where only the auditing data is allowed to be updated but it doesn't exit. So now you're stuck with read-only replicas of your database being noncomplaint.
It's a problem that really has no good solution outside of the filesystem driver supporting this mode.
* https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
Turning every read into a write......there are very few cases that would be preferrable for.
So it still requires keeping a database. But because all of that would come out of inodes, it avoids all the I/O of reading the files' contents.
Compared to mtime, it's harder to implement, but more resilient and at a similar cost. Compared to hashes, it's almost as resilient but much faster.
All those fancy checks plus database could be added to make too if you really wanted. It still wouldn't support all the fancy features of redo, but improve your experience with all those existing projects based on make.
Even if Bazel is the ultimate, final, perfect build system, I don't trust everyone, including myself, to use it properly.
I have successfully written my own toolchain rules in Bazel. Yes, it is a pain. But it’s usually only necessary if you have some kind of unusual target—usually, someone else has written the toolchain rules for you.
I've even witnessed a Bazel build system that made working DOS binaries. The entire set of rules was smaller than you might think.
The correct response for a compiler upgrade or a modification to a bespoke tool that generates a product is almost certainly a full rebuild unless you have very high confidence that you've identified ALL the dependencies. I believe this is not easily discoverable (e.g., you might have a dependency on "/usr/bin/cc", but not know enough about 'cc's internals to realize that you also have to have a dependency on "/usr/local/lib/cc1-arm-bfhsh-37.so" for some specific compiler option).
I'd like to believe you, but I have scars from doing build systems and discovering all the wacky and wonderful things that tools do.
Getting complicated cross-platform toolchains working on Bazel is going to be hard, sure. But knowing whether you have hidden dependencies is not hard, because of the sandboxing.
Like https://github.com/bazelbuild/bazel/issues/3360 .
It used to be the case that file hashing wasn't coordinated well with a concurrently edited file, so changing source while building corrupts your cache.
A point that neither this article nor the recent ninja retrospective (http://neugierig.org/software/blog/2020/05/ninja.html) touch on though is that hashing of files is an important step if you'd like to keep caches locally, but especially if you'd like to keep caches remotely.
A key reason that Bazel-alikes (including http://www.pantsbuild.org, which I work on) hash things is for use in cache keys. And many of those systems use file watching facilities (notify, fsevents) to determine whether files need to be rehashed. If done properly (i.e.: no false negatives), this is significantly faster than checking file mtimes, and it allows for "early cutoff", where if after hashing the digest of a file has not actually changed, you can skip rebuilding its dependents.
Meaning with the command:
foo.o: foo.c
cc -o foo.o $(CFLAGS) foo.c
make would save away "cc -o foo.o -g -DFOO foo.c" as part of the rebuild check.Also, I was under the (mistaken?) impression that some make actions came from a dependency tree, not a timestamp. Otherwise everything would be breaking during any invocation.
One other point about ATIME -- turn it off if you value your SSD.
1. The target doesn't exist
2. Any children are dirty
3. The mtime of any children is more recent than the mtime of the target.
All recursively propagated from leaves up.
Does anyone have an idea what's going on there?
Right click the offending text, select "Inspect Element". The dev tools should popup. On the right hand side there should be a "Fonts" tab. If you click this it'll show you the font being used.
I don’t think it’s just about purism. For some it might be. But to me, ever since I learned some of the details about how for example the FreeBSD VFS is implemented, from reading the book The Design and Implementation of the FreeBSD Operating System by M. K. McKusick, it makes sense. Prior to that my own mental model of how file systems work on for example FreeBSD was very fuzzy. But reading about the technical details, like what kinds of structs are being used in the VFS and what fields those structs have, I realized why a lot of things are the way that they are.
In the same paragraph that I quoted above, in the part I omitted for relevance, he also said:
> All sorts of very convenient tree traversals would be possible if the directory mtime were updated (recursively to the root) when contained files changed, but no. This is probably because of hardlinks: since the kernel doesn't, in general, know all the filenames of an open file, it literally cannot update all the containing directories because it doesn't know what they are either.
On this I have two thoughts.
1. Doing this means adding more overhead every time a file is written to. The overhead might be negligible, or it might be significant. We’d have to run a lot of benchmarks to really know. (For me the answer is that with the mental model of the file system that I have as mentioned above, I don’t desire this anyways, so I am not personally even going to do that type of benchmarking.)
2. Consider that one or more directories in the path between the root and your current working directory may be mounted from a file system different from the root file system. They might not even be local to the system! For example, perhaps your home directory is mounted over NFS or SMB. That’s certainly common in a lot of deployments at universities and big companies. From a philosophical point of view, does it even make sense to modify the mtime of every directory from the root down to your current working directory when you modified a file that is being written to a storage medium that isn’t even local to your system? And from a practical point of view, what good does it do that the desktop computer on desk A in floor 2 had all of its mtimes in its path modified when you head on up to floor 3 and sit down at the computer on desk K? Any tools you have that were relying on being able to start at the root file system and scan mtimes to find modified files would end up sometimes picking up changes (because someone else had recently modified another file in their respective home directory while logged onto computer 3K) and sometimes not (because at other times no one had been changing any files in their home directories between the time that you last did so on computer 3K and when you came back to it after having worked on your files from computer 2A).
As a workaround you might say, okay so let’s stop at file system boundaries. And that might work out, but it’s a bit hairy.
And still it could be that a longer component of the remote path is sometimes mounted and sometimes a shorter component. Then again, if the remote system is also doing this recursive mtime modification then that by itself is not a problem.
Furthermore, as a consequence of this, in order to make things workable, you should also be changing mtimes of all parent directories when mounting any file system to the most recent mtime of the file system you are mounting, if it is more recent than the mtime of the directory you are mounting it on in order to make the change visible. And if that makes sense then it means that mounting and unmounting file systems are considered as a change, so actually when mounting always set the mtime of all parents to now, and when unmounting update all mtimes again.
It seems like a lot of hassle, and potentially bad for performance. Though again, measurements would need to be done. And in the end I maintain that it boils down to what your mental model is for how the file systems that you use work.
Yeah, overall this article had a lot of great points and insights, but I think they oversimplified on this. (Not a big deal since it is a side issue.)
It's easy to look at a piece of software and think that, because it doesn't do what you want in a particular situation, it's a bad design. That's not the way you determine whether something is a good design; instead, you have to look at the big picture.
Usually, each different design option will have its own pluses and minuses and will make certain things easier and certain other things harder. The only look at your own situation, you're sort of cherry picking the data.
That's a much easier way to know something's changed if you're going to the trouble of keeping state.
"Checksumming every output after building it is somewhat slow."
"Mostly this is not too serious: the file is probably already in disk cache (since you just wrote it a moment ago!)"
> Checksumming every input file before building is very slow
If you have a very new CPU you might have hardware SHA support.
Related example: on a cloud instance with SSD drive, I have scripts that generate a ~10GiB file, then immediately after (while it might still be in cache), calculating an md5sum still takes tens of seconds (maybe it wasn't in cache? I've not investigated deeply). It's just an example of a case that falls outside of that "Mostly" category.
$ time dd if=/dev/zero bs=1048576 count=1024 | md5sum
1024+0 records in
1024+0 records out
1073741824 bytes (1.1 GB, 1.0 GiB) copied, 2.104 s, 510 MB/s
cd573cfaace07e7949bc0c46028904ff -
real 0m2.108s
user 0m1.891s
sys 0m0.628sEdit: "sum -s" is faster for me, but the man page doesn't give much info on what the algorithm is.