File format wiki
fileformats.archiveteam.org
fileformats.archiveteam.org
> At HandmadeCon @mike_acton brought up the game data format problem. I've reverse engineered hundreds of proprietary formats from games, and he's right that all I would need to make things simple is a basic description of the file data. Just struct declarations are enough. No encryption has stopped me, I always find the data I need, and your data all has the same general architecture, so there's no reason a game developer should be afraid of the public knowing how their game's data is laid out. This zeitgeist of closely guarding that information is doing you more harm than good — moddability is a key feature of many successful games, and all modders want is docs.
--
There's a lot of other industries that could stand to learn from this. Reinventing the wheel sucks. Doing your own, proprietary thing in your corner, taking on extra maintenance work, extra work to modify existing tools to make them compatible with your own image/archive/whatever format... it all sucks and nobody wins.
As a counterpoint, the community I founded around Hearthstone (http://hearthsim.info/) is extremely open source friendly and on excellent terms with the developers. It's really about how you decide to manage it.
The work consisted mainly in reverse engineering WoW's file formats, protocols etc in order to understand, index and organize data, and to be able to produce "datamined" patch notes as soon as new game files were available.
I got to work with and learn from people a thousand times smarter than me, it was by far my favourite job and it's why I still do various Blizzard-related work today.
It also creates an interesting relationship with the game's developers. I think on that front, many studios could learn a lot about working with the people interested in the technical aspects of their game, rather than working against them. Most get it horribly wrong, few of them get it right at all - Riot is one of the few that do, I wish I enjoyed LoL. :)
http://www.eurogamer.net/articles/2013-12-06-avalanche-gives...
vs
http://www.eurogamer.net/articles/2015-08-10-rockstar-bans-g...
Studios that get too large for their own sake and get completely disconnected from their userbase. They no longer see the players, the fun or even the game - they just see the numbers. Business guys thinking a video game company can be run like a bank.
The devs get blamed for it, too; as if they had anything to do with it. These decisions are made high above them.
I have a couple small projects on the backburner at work that require me to reverse engineer file formats. I've been able to decode a few basic details out of the files by looking for patterns in a hex editor but the finer details are still escaping me.
Your other option is something like QuickBMS - on mobile, can't link.
It gets tougher when there are anti-debugging protections in place but that generally happens only in the video games industry.
If the format is simple enough though, you will be fine with a good hex editor and some pattern recognition.
To echo this with code, I've open sourced my collection of game archive unpackers. Given how similar they all are, this is a typical example of how to add a new decoder:
https://github.com/shish/pyge/blob/master/archive/PackDat3.p...
(Note that two of those 6 short strings are user-friendliness metadata and can be omitted, and one of the others is redundant -- "Magic Bytes", "Header Format" and "Directory Entry Format" are all you really need in most cases)
just go to: http://www.fileformat.info/format/${file_extention}/internal...
Under my limited experience working with wikipedia and wikidata I've seen it's not the best option to a) store structured data and b) edit the pages (markdown is, imho, much better for that).
[Last time I said that, someone recommended Moin. I hope this won't happen this time - Moin is a UX joke]
http://fileformats.archiveteam.org/wiki/Dendrochronology
http://fileformats.archiveteam.org/wiki/Quantum_computer
http://fileformats.archiveteam.org/wiki/TLD_.mobi
http://fileformats.archiveteam.org/wiki/Endianness
How is any of these things a file format? Because the definition of a file format in the FAQ (<http://fileformats.archiveteam.org/wiki/FAQ:File_Format>) is so broad that it can be made to encompass basically everything. The manifesto that started this site is a rambling sermon that doesn't clarify anything in this respect: it just repeats "let's solve the problem" without even properly defining what the problem is.
There are also other issues. The classification scheme is often Procrustean and confused. Error messages were until recently mixed up with error detection codes, which conflates two different meanings of "error". Similarly FUSE shares a category with HFS+, even though the former is an API, and the latter a disk format; distinct things which just happen to share the name of "file system". The pages are rather short and consist mostly of lists of links. Given the above-mentioned lack of clearly defined scope, I suspect many pages seem to be created about topics just because they're in the news and/or just to have a place to put a link to a "neat" blog post: see for example <http://fileformats.archiveteam.org/wiki/Facebook#Links>.
Last but not least, the whole site is rather ugly, and the logo is awfully non-descriptive of what it's supposed to contain; what am I supposed to do with this thumb, stick it up my arse?
It's a shame, really, because documenting file formats is a hard and valuable endeavour. But I don't think these people are going to do a good job of it.
Any successful wiki will have a problem of demarcation; the initial signs being that people add whatever they want, including ramblings like Two cows. [2]
If the project goes along, taking Wikipedia as precedent there will develop two "inclusionist" and "deletionist" cultures that will fight over the definition of what the project should cover, but which would get rid of these extreme cases.
[1] http://fileformats.archiveteam.org/wiki/Quantum_compressed_a...
If you're not and are interested by it or in contributing to it, you should email me - I'll put you in touch.
I currently work on a proprietary l10n/i18n system, so a lot of what I do is under an NDA. However, I'm slowly gaining traction in moving us, at least in part, to a more open model.
If you let me know your email, I will reach out.