You can find more information about us at https://srcc.stanford.edu
And you can find more information about Sherlock at https://www.sherlock.stanford.edu
You can find more information about us at https://srcc.stanford.edu
And you can find more information about Sherlock at https://www.sherlock.stanford.edu
On a NFS mounted on my PC your LS_COLORS tweak actually degrades the performance. Without modification, it takes 0.5 seconds to list a directory (on a slow consumer grade HDD) with 14448 files and after setting
export LS_COLORS='ex=00:su=00:sg=00:ca=00:'
it takes 4.4 seconds.However, listing a local (SSD) directory with 15k files takes just about 0.12 seconds and gets faster after settings the variable to 0.06 seconds.
For all tests I drop the caches before testing, e.g.:
echo 3 > /proc/sys/vm/drop_caches; time ls --color=always /mnt/nfs/many | wc -lRun "strace -o logfile ls --color=always" and diff the two logfiles.
# du -sh logfile_nfs*
1.7M logfile_nfs
668K logfile_nfs_exported
However, the strace -c output looks kinda different:Without modifying the variable: https://bin.disroot.org/?149fa91c08b27312#0w9O6BAWNEEC4SUXEb...
With setting the variable: https://bin.disroot.org/?149fa91c08b27312#0w9O6BAWNEEC4SUXEb...
https://bin.disroot.org/?56f3dcb618240df0#x3StRGZZgSzM0UJ0BK...
But for completeness, I wrote a few loops which should answer your question:
=NFS=
Normal: https://bin.disroot.org/?8967a0f34b26e512#aTYmbyESfeuqAXS802...
With exported Variable: https://bin.disroot.org/?c405d7aa74b50ee0#KverJYJEzNhct7CmdW...
-----
=Local=
Normal: https://bin.disroot.org/?b94e1d1f58e3edb7#tKszD/tjvwBwepJLun...
With exported Variable: https://bin.disroot.org/?d3ac83f1dff9e767#WVODjGnvL1QbOJdWPp...
I also tried dropping the cache on the NFS server but that didn't seem to have a major effect on the performance (probably because reading a 15k file index from a local disk doesn't take that long after all).
I helped a researched debug a Lustre performance issue a while ago. Each job was nothing special, read a few files (maybe a few GB total), do some (serial, no MPI or such) calculations taking maybe 10 min or so, produce output files, again a few GB. No problem, except when the person ran a several hundred of them in parallel as an array job the throughput per job dropped to a small fraction of normal. Turned out that all the jobs were using the same working directory. Slightly tweaking the workflow to have per-job directories fixed it.
Disclaimer: I work for Dell.
Disclaimer: I work on Isilon
Edit: By "file storage", I'm talking about storage mounted using protocols like SMB and NFS.
IMO. I guess I'm just an old fart.
It's not because it has colours that it's not a serious article.
Postfix emoji, eurgh
IMO. I guess I'm not an ageist.