Yes, actually they could. They could be readable without any software at all. Like XML.
All of the improvements suggested in the article could be implemented by extending the XML format to allow including by other XML files by reference.
Yes, actually they could. They could be readable without any software at all. Like XML.
All of the improvements suggested in the article could be implemented by extending the XML format to allow including by other XML files by reference.
The problem with formats "readable without any software" is that you usually end up with everything represented as strings, with no knowledge of actual types and constraints. So, you still need that documentation, even though the format is "readable without any software".
If your document editor represents documents as a stack of edit-events (for undo/redo or whatever) and you document every nook and cranny of this "document format", standardize it as "binary event-sourced document format, .esd" then your binary spec will be very close to a full description of your document editor. Everyone who wants to use your format now has to reimplement your editor.
Anyone can leak implementation details into any format, and it is to some extent unavoidable, but if you can make it readable (or semi-readable), some thought has at least been put into the representation.
What an extraordinarily stupid notion. Your filesystem is not readable without a filesystem driver for the specific binary format used to represent it, so you cannot even read your text file without having something in your stack that understands a more structured data organization scheme.
This argument doesn't hold water, since file content doesn't depend on which filesystem it's on.
An archiver may choose to store data in some very particular way, like some specific tape driver. They're free to do that without corrupting the content of any files, since the files don't depend on the filesystem.
On the other hand, if an archiver chose to store, say, unzipped versions of their opendocument files, then they've corrupted the data: opendocument files are zips, unzipped data is not opendocument. The format does depend on the file content.
No, it's not. An sqlite-based document can have as many representations as an XML-based document -- nothing in its implementation prevents this.
>so "readable without any software" is a good test for wether it's "readable with any sfotware"
The OS, filesystem, shell, editor etc needed to read the XML file are still software.
As is the XML parser needed to do anything useful with it in the data realm.
>If your document editor represents documents as a stack of edit-events (for undo/redo or whatever) and you document every nook and cranny of this "document format", standardize it as "binary event-sourced document format, .esd" then your binary spec will be very close to a full description of your document editor. Everyone who wants to use your format now has to reimplement your editor.
You do understand that the XML file can also just be including a stack of edit-events, right?
This is totally orthogonal to the underlying storage (XML or sqlite or whatever).
What if my text editor has an sqlite file viewer?
I despise xml quite a bit but I have to admit, it's still potentially better.
Binary formats start out mostly unreadable by humans [1]. XML and other textual formats at least have the possibility of being made to where they can be read by laypersons.
1] I'll note that after enough immersion, I've seen people read binary core dumps, etc. but that takes much more time and practice than with XML.
Would the average person edit a .svg by hand? No, he'd use Adobe Illustrator or anything else.
Would he edit a .docx file by treating it as a zip archive and edit content.xml? No, he'd use MS Word/LibreOffice Writer...
Suppose Sqlite used XML/JSON instead of binary files, would you modify them with notepad? No, you'd probably use a SQLite browser software.(Or an application that is more tailored to the domain)
Fair point, and in fact, I'd go further - most of the time most people, do not need to directly edit or view most files in most formats. Even if you took 'most' to mean 99% or higher, I'd be comfortable with that statement.
Where I differ is that I think the essence of the argument is really whether or not a binary file format offers enough value to be worth entirely eliminating the direct edit/view possibility for everybody all the time. Even if it's not a common case, it can be game-changingly useful when you need it.
Just to illustrate, you give three examples, and I have counter examples for each:
> Would the average person edit a .svg by hand? No, he'd use Adobe Illustrator or anything else.
I've modified bounding boxes by directly editing SVG, as well as read SVG directly to analyze some plots a library was generating for me.
> Would he edit a .docx file by treating it as a zip archive and edit content.xml? No, he'd use MS Word/LibreOffice Writer...
I've done this recently to extract out embedded documents on OSX (where the native versions of Office do not directly support this.)
> Suppose Sqlite used XML/JSON instead of binary files, would you modify them with notepad?
HSQLDB can use text based SQL scripts to store data, and I've modified and edited them directly for several reasons.
They aren't, though. Or can you read ZIP archives without an unpacker?
For some values of "readable" and "human".
It fits the requirements of reading the file without (external) software since you can trivially reimplement it, and is also a much lower threshold than "you need to reimplement sqlite", which is why pretty much every system out there comes bundled with a zipfile unpacker.
However, all software is not equal. Any OS comes with a text editor and a ZIP utility by default now. The same is not true for SQLite's format.
Normally tar balls are used and zip has to installed afterwards:
apt install zip
I read the article as advocating for more open binary formats. That it is useful to have binary formats is logical. That someone would like to have open ones is I think a nice thought.
That's not entirely true. Anything that comes with Python installed has SQLite bundled/compiled in. It's just the library, not the cli, but still.
• Win 10 apparently does (I don't use it, so can't personally verify). Win7/8 don't.
• Every version of OSX for many years does
• Fairly sure the main Linux distro's do
• FreeBSD does, not sure about NetBSD/DragonFlyBSD though.
Checking a local OpenBSD 6.x box just now, sqlite/sqlite3 aren't in the default user path, so maybe that's one that doesn't.
The same thing is becoming true of SQLite (and pretty much is true if you exclude older versions of Windows). We're talking about one of the two (estimated) most-widely-deployed software projects in the world today (http://sqlite.org/mostdeployed.html). It's FOSS, ubiquitous, easily embeddable, works remarkably well, and has bindings for pretty much every programming language under the sun. As far as a general-purpose file format goes, a SQLite database is thus pretty close to ideal.
But you do have a point that in a rare case, you could get your hands on the raw data which makes text easier.
But everything is a tradeoff, and the question really is if the tradeoff in this case is worth it, since sqlite is so open and portable.
The only difference is that xml is also human "readable". But most xml isn't really readable without a tool anyway.
You may not want to look at XML without a tool, but you could. Or, most likely, you could build a tool to look at the XML. This is a crucial design consideration of XML - even if all XML tooling, open or not, is lost forever, you can still read the XML. Not so with SQLite.
Especially for a document format, this point seems important. That said, OpenDocument is zipped, which, while a simpler and even more ubiquitous format, is still a binary format that requires a library to read, so that detracts somewhat from the point. It's difficult to imagine a future where zip libraries survive but SQLite libraries don't.
ASCII encoded text is simpler than UTF-8, but there is content that can't be represented in ASCII. Zip comes a bit up the food chain from there, and SQLite comes quite a bit further.
XML is about being meaningful without tooling at the lowest possible level. Any added required tooling or documentation by definition makes the data less 'open'.
At this point, you've side-tracked into an argument that anyone could make against any software.
SQLite for all it's significant qualities are none of that, and thus you shouldn't use it if designing for file format longevity.
You can't read XML without software.
Perhaps you take editors and cat and less as a given.
In which case, nothing prevents you from taking sqlite the binary (or 1000 serializers that can be produced very easily) as a given too.
XML is only trivially "readable without any software at all". The same is true of any binary format - the issue is that it can be read, but not easily processed - which is also the case for XML (you need to integrate an XML parser for that).
And if you are including XML libs, why not SQlite libs?
A printout of a binary dump does not provide any clues as to how the data was structured.
> a binary dump does not provide any clues as to how the data was structured
True, but if that binary format is SQlite, the structure is fully documented. You could work it out mentally just as you could work out any given XML schema.
More to the point, no one needs to install a full blown database engine to read a plain text document.
You may argue that specialized tools make the job easier, but you also need to acknowledge that requiring specialized tools just to open a text file is silly.
you already need a decoder