This already exists in libmagic (https://github.com/file/file) and can be used on any BSD/Linux system by typing `file [filename]`
This already exists in libmagic (https://github.com/file/file) and can be used on any BSD/Linux system by typing `file [filename]`
People are reading the title thinking it's about filesystems. It's about cataloging all human information storage methods, from phonographs to zip drives to word perfect files.
I've spent a long time (> 10 years) developing code for detection of file types (without relying onto the file extension), and extraction of the data... And everything was need to be done very fast & reliably. The last project was the content detection for the filtering web proxy, where every millisecond counts. And for signatures, there was a lisp-like language that allowed to describe very complex detection rules, although it's sometime was still necessary to go to the C++, for things like, OLE2 parsing, XML parsing, or listing files in Zip files, etc.
From the open source alternatives, it makes sense to look to the Apache Tika that has both content detection & extraction capabilities, although it's written in Java...
https://mark0.net/soft-trid-e.html
which has also a couple online options and a database of signatures/magic numbers: