Let's Solve the File Format Problem (2019)
fileformats.archiveteam.org
fileformats.archiveteam.org
You are worried that if an application has a directory "cache" with files in it, a user will go into that directory and randomly delete the file "cache/3F3819A1.dat", and that might confuse the application.
But, suppose instead you had a single file "cache.sqlite", containing a table "cache" with an "id" column. What's to stop a user from opening that file with the sqlite command line tool and running "delete cache where id='3F3819A1'"
If a user really wants to muck with your app's data by hand, you can't stop them. (Unless you go with some DRM-like solution in which your app runs in a secure enclave such as Intel SGX)
Whether it is a single file or a directory of files doesn't really make a difference
On Solaris, xattrs really are up to the task; they are basically Windows alternate data streams by another name. Unlike Linux xattrs, their maximum size is the same as that of the main stream of the file. (Likewise, on macOS the maximum xattr size is the same as that of a file's main stream.)
I think if Linux would lift its limit on maximum xattr size then you would have the same basic function across most major operating systems, just with some differences in API. Then if someone could devise a standard API (even as just a portability layer), your problem would be mostly solved. Just also making sure that tools like cp, scp, sftp, zip, tar, etc, know about the xattrs and transport them. (Which is actually already true for some of those tools on some platforms.)
Why not just leave files as a sequence of bytes and let the applications handle this in their file formats? It's not like we don't have standards and libraries there already.
I guess it boils down to what is "application data" and what isn't.
Instead you either need sidecar files, archives, or custom binary formats.
It's as if the posix filesystem abstraction crystallized a bit too soon (same like tcp/ip)
Hmmm... have you ever heard about the concept of archive? It's a single file that you can copy around, but it contains literally several files inside!
Another option is to tell the operating system to deallocate part of the file, turning it into a sparse file. This can be done with fallocate(FALLOC_FL_PUNCH_HOLE) on Linux, FSCTL_SET_ZERO_DATA on Windows.
Linux also has fallocate(FALLOC_FL_COLLAPSE_SIZE) which lets you actually remove a byte range from the middle of a file, shifting up everything that comes after that byte range. So, in fact, you can "move" bytes in a file, on certain filesystems on Linux (ext4, xfs, among others).
Unfortunately, support for sparse files is quite variable across platforms. The macOS equivalent to fallocate(FALLOC_FL_PUNCH_HOLE) if fcntl(F_PUNCHHOLE), but the issue is HFS+ has never supported sparse files. UFS did, but UFS support was dropped from 10.7 onwards. The support is supposed to be back in APFS, but I don't know if it actually works.
I have done exactly this for a few applications. I likely wouldn't have done it for the database.
The reason I've done this is disk space. Some software is bad at getting rid of things that they've cached (eg Chrome). Theyend up taking more and more and more and more disk space. I don't understand why my browser needs to take up 10-20 GB of disk space, if it can't even retain 7 months of history.
But then you just hide your bad caching design from users by only "showing" what you're doing wrong to those that understand SQL. And those users can still change it.
You should test your app for random cache permutations (user changes .dat file in the wrong way) and partial cache loss (user deletes random file). You have to assume that anything that can be broken will be broken.
Curious if you can clarify / specify / examplify what "semantic" means in this case?
Having a table with columns A, B, and C doesn't say a ton about what the values in them mean. Or how they relate to one another. If it's a bunch of decimal numbers, it could be financial data, and the sums expected to sum depending on how another column that indicates a type is set. Or it could be temperature readings. Or distances, with the unit implied or tracked elsewhere.
Generally this is the kind of data and rules for reasoning about data that isn't stored in a database. Certainly not a sqlite database. Column and table names are often vague or unclear, and in my experience it's an exceptionally rare developer who leaves detailed commentary on the relationships between columns and tables and their possible values in their create statements.
IMO Sqlite would be a terrible format for e.g. a word processor. Database style formats are good when partial updates are the norm, but if your application typically needs the entire data model in memory, a relational mapping is unpleasant overhead.
Ideally, this type of invariant would be in the DB, and not as documentation but enforced in code (i.e., by triggers, since it's to complex to be a simple referential integrity constraint.)
> Generally this is the kind of data and rules for reasoning about data that isn't stored in a database.
The ones that are invariants should be enforced in the DB, and the ones that aren't should generally be in in-db documentation. Sqlite in particular lacks the ability to attach the metadata usually used for that purpose, but does store code comments in the schema, which provides roughly-equivalent functionality.
What level of in-db documentation are you in the habit of writing? I can honestly say I've never worked on or encountered a codebase that explained fully the significance of any of its database columns with in-db comments. I would count myself exceptionally lucky to find natural-language prose documenting such details! Generally I find myself puzzling things out by a combination of reading application code, tinkering with tests, and interrogating developers.
I think I've only ever once seen any real use of comments in a real-life codebase. It was a postgres database, where columns that contained PII were commented according to whether it was plaintext or encrypted.
~10 years ago I worked on a database frontend that stored schema metadata (version, additional type info) in database comments ;)
I've maybe seen database comments used a handful of times. But I also worked a couple years as a consultant, so saw a lot of databases...
The argument for this has already been made very well by the sqlite creator.
Basically there is no way to say to you operating system "hey, can I get a no-nonsense, don't what to have to come up with some other conventions that will be broken, way to have some additional storage and file system paths that are just for this file so I don't have to worry about name collisions?" I already have a perfectly good unique name under which I could stash information, but the posix standard basically prevents me from using that name as a way to safely namespace things that I need per file.
However, to your original point I have partially implemented this [0] using zip files on the read side. I treat the internal paths inside the zip file as if they are rooted at the zip file itself. I haven't implemented rewriting though.
0. https://github.com/tgbugs/augpathlib/blob/master/augpathlib/...
For the record, it does. The Boomla OS has exactly that, the ability to store files in files. As in, files and directories are the same concept. It is not a posix filesystem of course and the OS is built specifically for web development. (Disclosure: I’m working on the project.)
https://boomla.com/docs/how-it-works/anatomy-of-the-boomla-f...
This already exists in libmagic (https://github.com/file/file) and can be used on any BSD/Linux system by typing `file [filename]`
People are reading the title thinking it's about filesystems. It's about cataloging all human information storage methods, from phonographs to zip drives to word perfect files.
https://mark0.net/soft-trid-e.html
which has also a couple online options and a database of signatures/magic numbers:
I've spent a long time (> 10 years) developing code for detection of file types (without relying onto the file extension), and extraction of the data... And everything was need to be done very fast & reliably. The last project was the content detection for the filtering web proxy, where every millisecond counts. And for signatures, there was a lisp-like language that allowed to describe very complex detection rules, although it's sometime was still necessary to go to the C++, for things like, OLE2 parsing, XML parsing, or listing files in Zip files, etc.
From the open source alternatives, it makes sense to look to the Apache Tika that has both content detection & extraction capabilities, although it's written in Java...
To “solve” file format, you may need a lot of data beyond the file provided. Does the file come from a floppy disk labeled VisiCalc? Did you read the magnetic media correctly? Let’s say you did. You look at a file and then try, the best you can, to determine what it is, maybe given a lot of information about how the computer that read the file interpreted it.
Sure, if you’re talking about more modern files, maybe the filename’s extension is a good hint.
Some files contain metadata headers or footers. They maybe have multiple layers of them. That metadata may have evolved. Perhaps it wasn’t versioned.
Files were stored on cassette which could be listened to as audio or interpreted as containing file definitions.
What if the cassette were warped in the heat? You could still recover the file partially or fully if you knew that you could speed up portions of the recovered audio, perhaps with additional pitch-bending and error-correction. Could you do that without any background information? Maybe, but it may take much more effort.
Warning: Unknown: Unable to allocate memory for pool. in Unknown on line 0
Warning: require(): Unable to allocate memory for pool. in /usr/local/www/mediawiki/index.php on line 54
Warning: Cannot modify header information - headers already sent in /usr/local/www/mediawiki/includes/WebStart.php on line 63
[..]
Seems kinda weird this stuff isn't cached?http://fileformats.archiveteam.org/wiki/Electronic_File_Form...
At work we have one that creates a “.snapshot” directory with folders with backups. It works pretty well. (Net app?) I would like that idea at home (Apple has time machine which I guess is similar)
One thing about Linux, I thought it would eventually surpass the proprietary OS in a lot of areas including file systems. I think it has in a lot of ways but file systems seem to be a bit of a mess.
I’ve heard good things about BRfs and zfs but like all things Linux I’ve heard bad things (you will all your data). There are so many options and changing off the default seems hard.
http://fileformats.archiveteam.org/wiki/Category:File_format...
From that, I can go to SQLite, HTML, DAG, whatever. It's easy to work with, easy to back up and synchronize, and as human readable as it gets.
`Warning: Unknown: Unable to allocate memory for pool. in Unknown on line 0`
Guess the 'hug of death' has exhausted available memory for PHP to display the page.
> Hosting is provided gratis thanks to the kind folks at Tranquil Hosting
The architecture is totally fine for their use case. Arcitecturing it to stand up to hn hug of death would be over-engineering and not really make sense given their goals.
https://osuosl.org/ https://www.fosshost.org/ http://www.gplhost.com/ https://www.digitalocean.com/open-source/ https://www.gandi.net/en/gandi-supports