Naming things (2015) [pdf]
www2.stat.duke.edu
www2.stat.duke.edu
All this ancillary stuff that we'd like attached to files, like dates, client names and projects, versions and so on, are metadata. Some systems keep metadata in files: EXIF, Word, PDF. Some systems have conventions for this instead: header blocks in source code. But if neither of those applies? Only place you can put it reliably is the filename :(
In other times, the industry would solve this by creating a standard for metadata interchange, but one of the biggest players gets a too big to ignore amount of revenue from FAT, and is able to block any attempt at standardizing.
ZFS, ext4, and plenty of others used by BSDs and Linux distros have no licensing issues.
Also, if any player cared to win the filesystem war, they'd open up their spec royalty-free.
I've written several "disk and file catalog" utilities over the years so I inevitably spent a lot of time thinking about the "metadata" problem.
I think the issue is that it's impossible to solve metadata in a universal way that satisfies all scenarios. This is why metadata often ends up being inscribed into the filename. It's the "least worse" solution.
Let's take one example of the scientific data of csv files. Typical Comma-Separated-Value files do not have metadata fields such as author, measuring device, timestamp of readings, GPS coordinates. (Yes, csv files sometimes have a first line for "column names" which is arguably metadata but that's not the higher-level metadata I'm talking about.)
Exactly where does one put that high-level metadata?
1) If one makes a new pseudo-standard that signifies any lines at the top the csv beginning with "//" as metadata, that means that modifying any metadata of a 100GB csv file (e.g. change author from "John Doe" to "Jacob Doe" is rewriting the whole 100GB file to add 1 byte.) As a related issue, let's say you have hash of "e1bb76e7391b93eb12" for the csv file. You really want a stable hash that represents the actual "raw data" of the csv file. You don't necessarily want the hash to change just because the metadata changed. In this case, embedding metadata into the file itself makes certain operations worse since typical hash utilities don't have "intelligence" about which parts of the file is "important" for hashing. (A similar problem is scanning mp3 files for duplicates. If 2 mp3 files have bit-identical audio output but the metadata tags are different, are they the same or different?!? It depends.)
2) if you put metadata at the end, typical utilities won't know about about it. (UNIX has "tail" command but standard MS Windows does not. The tail command is also unstructured and read-only which makes it a non-solution for managing end-of-file metadata fields. Also, the "quick" view of GUI file managers show the top of the file and not the bottom of it.)
3) If you put metadata in a separate file, it easily gets lost. File managers like MacOS Finder and MS Windows Explorer don't know when 2 files are supposed to be "treated as one unit" vs separately.
4) If you try to put metadata in a separate special area using os file system features suchs MS "NTFS alternate data streams" or Mac OSX "resource forks", they will get lost when transferring across incompatible filesystems or uploading to Amazon S3.
If one is feeling uncharitabe, one could say the MS WinFS[1] was a spectacular failed attempt at unifying metadata. (A relational database that makes metadata more of a 1st class concept.) Nobody has tried it on that level since. Even Apple's new file APFS system didn't have the same metadata ambitions as WinFS.
The combination of tradeoffs leads everybody to re-invent the idea of embedding metadata (including namespacing hierarchies) into filenames. The article's suggestions for scientific data filenames looks very similar to filenames that companies end up using for ETL pipelines.[2]
Not sure if it's a good solution but if I were to put the meta data somewhere I would somehow try to put it in the identifier of the file (the name) as it is data that would help me identify the file AND it's content!
https://blogs.technet.microsoft.com/askcore/2013/03/24/alter...
They sort-of do, each in their own way. Mac OS has packages: directories with files that the Finder treats as a single item. I think this is mostly implemented outside of the file system; if you give a directory a file name with an extension that some application claims to be an extension for a package, the Finder treats it as a package.
On Windows, when you export a web page from Internet Explorer as "web page, complete", you get a file and an associated folder containing the images of the file. MS Explorer shows them as separate icons, but knows that the two form a unit; when you delete one, it informs you about the existence of the other. I don't know how this is implemented, but suspect it is 100% outside of the file system, too.
Also, classic Mac OS had resource forks: a single alternate stream for every file. Text editors used it, for example, to store the cursor position, line wrap settings, etc. across saves.
Jakob Nielsen had a post (a link for which I cannot find) recommending that web-page titles put the most specific information at the beginning. Doing something similar with file names (e.g., calling them "H01_MutantFraction...2013-06-26.csv", etc) would trade some of the advantages of the proposed scheme for speed of finding and switching between files when you're actually using them.
If one needs to distinguish groups of files, why not just put them in directories? That's the reason directories exist, no?
I can somewhat understand if some (bad) software is written to look for files only in a single directory and you have to put everything there. But otherwise, it seems pretty pointless to use a common prefix and make filenames longer.
It's like brand name packaging, all the information including nutrition packed neatly on the outside. You don't go to the store and buy 'bread' you buy '2017-08-18-00-natures-own-dbl-fiber-wheat'.
This whole pdf resonated with me because it made me realize I'd developed these almost identical practices without knowing it. Mostly over time, trial and error, and a kind of natural selection, when it comes to sake of ease.
Cool stuff.
Which is to say: yes, you could, but it seems like a worse idea than just not putting them in directories at all, and relying on the OS's (strong) search features to get you want you want.
Directories do not really cost anything, so carrying them around doesn't really matter that much. You can think them just as simple (hierarchical) tags. If you really cared, doing a (scheduled) GC pass to prune empty branches shouldn't be too difficult to do.
Of course the question if tags are actually good way of organizing things is still open, but that is distinctly different problem than "baking in a single specific organization"; I would even venture to say that it is almost opposite problem, simple tags generally being too freeform and unstructured.
I'm a graphic designer, so for me everything is Client/YYMM-Project/_FINAL/YYMM-COLLATERAL-NAME
Within each project there is a _PROCESS folder with a _ELEMENTS subfolder for pieces the client has given me to work with.
For invoices I do YYMMDD-ClientName-Project-Sum.pdf. When the invoice is paid, I rename the file to add -PAID- before the client name. Its simple, but its allowed me to easily track and maintain projects and billing over the years.
If I end up working for another 83 years, I guess I'll pad the year with a 0...
Proper file management is an undervalued skill and should be taught both in school and in corporate environments. In an old tech job we had a public folder on the server that was total chaos. So many people insisted on naming their files MAY-%day%-%contents%-%personsname% -- and, as you'd expect, people spent countless hours per year trying to hunt down that one file so-and-so worked on before they left for another job.
> Within each project there is a _PROCESS folder with a _ELEMENTS subfolder for pieces the client has given me to work with.
Out of curiosity, why do you have the year/month in the names of both the project folder and the collateral file?
And why do you start the names of your final/process/elements folders with underscores?
[Edit] As dkarl mentioned more elegantly: this system helps map a project in time.
YYMM on project files is shorter and less ambiguous than using a project name for the same function.
I'd imagine the sub-directory underscore is also used to aid searching. You can easily visually identify (or filter) sub-directories.
> The initial underscore helps by (a) keeping those directories sorted [separately] and (b) providing a visual clue that they are 'special'. [0]
0. https://www.sitepoint.com/community/t/why-do-some-folders-ha...
I prefer organizing by year and by category(business, code, etc) instead of by client but overall I'm mostly happy to have any organization of my files at all.
Modern operating systems index the contents of your files, so finding all your files on project "foo" is only a search away. If you are a GUI user, then file naming isn't really that important for locating data on your machine.
Where file names matter is because there is 40 years of cruft out there which absolutely refuses to move metadata along with files. So you can touch your files to set dates, organize them in directories or tag them to your heart's content in Mac OS or Windows, but you will lose all that information when you attach the file to an email or put it in DropBox.
So you only have a choice of two places to put metadata in such a way that the metadata will be carried along with the file- the file name or the file contents.
Putting the data in the file contents lacks discoverability and in many cases the applications you use to manipulate the files don't allow for additional metadata anyway. Also, some file types (Word .docx files, jpegs, MP3s) get their metadata updated and/or scrambled when you open them with specific applications. So really your only valid choice is to put it in the file name.
The author's specific recommendations (use underscores and hyphens for delimiting) assume that you really want to access the files with the command line and use globbing. Other than that implication, the recommendations are sound.
It's certainly good advice and I definitely avoid using non-ASCII characters in filenames in practice. But I can't help thinking that advice like that is why support for Unicode is still buggy in many places.
I see nothing fundamentally wrong with using non-ASCII for filenames (and the slides don't give any reasoning), if only random software wouldn't mangle encodings, sorting order, or plain refuse to accept such filenames.
If you're getting into R or data analysis, check out http://stat545.com/topics.html. She has put a lot of thought into the project management aspects of carrying out a data science project that don't get discussed as often as other, sexier topics.
Edit: Glitchmr has a better solution below, "ls -v".
Its so hard to use CI interface and type that filename every time or navigating to in file explorer when its so long to read and the important bits of info should always be at the start.
Iso date conventions arent necessary since most files have metadata associated with it (create and modified date) so adding ISO date format is redundant for human made files . As you can always use a bulk renamer at any point.
But my point still stands just sort regularly and add some sorting identifier at the front of the file. Depending on who is working on file and what context it is a simple number at front suffices 01, 02, 03, etc or it can be a word and version number at end
Lastly the author didnt mention foldernames. Those need to be one or two words at most to help segregate information if theres lots of files in that one folder
If your creating machine made / autogenerated saved reports following a standard ISO state convention makes sense though, with regexable slugs, etc
Example: client_project_element-name_20170818.txt
When sorted by name, similar items are grouped together.
I like to call this middle endian.
i.e. [a-zA-Z][0-9a-zA-Z_]+\.[a-z0-9]+
2017-04-04 #S4907 #Choir List of names.pdf
2017-04-05 #EXCITE #S=5005 Notes on Data Repositories.pdf
2017-04-06 #ARAG #S=5031 #Amount=14.2 #CUR=EUR Invoice.pdf
with the following semantic:* date. Every file name is prefixed with the ISO date to facilitate sorting
* tags. The syntax #<tagname> to categorize documents with tags.
* key-value pairs. The syntax $<key>=value let's us attach structured information to the document name.
and keep them in a single large folder.
On top of this, I have written some shell tooling to normalize, view, and shuffle around those documents: https://github.com/HeinrichHartmann/pile
E.g. `pile extract EXCITE` will extract all files with tag #EXCITE to a separate folder named #EXCITE. There is also a HTML form that helps with proper naming of new files.
File management is still a pain for me, but this at least gives me some confidence that I can retrieve stored documents reasonably well. I hope, that one day I'll be able to auto-generate expense reports and tax filings, from properly tagged up filenames.
Just make sure to quote your variables in bash.
Not so practical for anything else but file naming IMHO...
Second "little endian" dates are inconsistent because the year is still big endian. If you want to remain consistent you would have to write the current year as 1720 (or even 7120!) because the years (or decades) are smaller than the century. To achieve the consistency of ISO 8601 with little endian time you would also have to write seconds before minutes and minutes before hours.
Of course, when we remember that Arabic is right-to-left, we realise that in the original Arabic the least significant digit was read first ...
BRAFWTNEGASSAY doesnt make any sense. To distinguish between the filenames you actually have to read the full filename, and with such long filenames chances are they won't fully be displayed, so that you have 2013-06-26_BRAFWTNEGASSAY_Plasmid-Cellline-100-1MutantFrac... For the first four files.
Keeping cruft out of your filenames seems like a much, much better way to name files. Also, most systems keep track of the creation data, no need to keep it in the filename. I think it's better to give files an id.
True, but often it makes sense to put date into the filename as well. For example, notes for an event that happened on a certain date. You might write the first draft at that date and then edit/move it later. It's still strongly related to the date the event took place, but the filesystem ctime and mtime will be different.
You could have named it differently: 2013-06-26_KUTKLOON7_Plasmid-Cellline-100-1MutantFrac
Creation date can sometimes be lost if you copy/move the file between different mediums
I agree. Maybe at some point, and in his example I think it makes sense to just change some of those '_' to '/' and boom, you now have folders.
You only need one tool that doesn't quite do things right to lose that:
- file copy programs need to explicitly set the creation date back to that of the original file.
- when you do save a file in an editor, it typically writes a complete new file (ideally through a write temp file/delete/rename dance). Again, the program may forget to reset the creation date of the new file to that of the original.
- I don't think got even _stores_ creation time stamps in repositories.
I've seen too many files with obviously bogus creation time stamps to put much trust in the creation dates of files.