QuickTime as a Tape Archival Format
eschatologist.net
eschatologist.net
Obviously proprietary formats with not many open source parsers is the biggest issue, but still a neat idea to think about.
There are subtle differences in how you should lay out some atoms (e.g. ftyp is subtly different in QuickTime, MP4 uses free instead of wide atom, some Macintosh-specific fields in vmhd(?) and related atoms are always zeroed out and marked as "reserved" in MP4, etc.). But it's easy to gloss over these differences when you do not care about the Macintosh-specific things (like QuickDraw transforms, sprites, ...)
That's definitely not true for LTO which can consume non-deterministic amounts of tape when a write head is acting up or there are buffer underruns.
Of course the physical reality of tape is much more complex, but when considering something like an archival data format you need to take into account which aspects of physical reality to preserve and which aspects to drop.
The advantage of the QuickTime container format is that it allows for a straightforward representation of all of the aspects of tape that matter for archival purposes—including subtleties like inter-block gaps, tape mark sequences, leader lengths, bad blocks, and so on—in addition to allowing straightforward and efficient storage of data blocks of varying sizes and open-ended metadata, as well as allowing for the possibility of non-destructive editing.
Most attempts to represent tape for archival to this point simply haven’t handled any of this, and have instead have just treated tapes as fixed sequences of files containing blocks, often without even recording the sizes of those blocks.
Clearly for streaming purposes they make sense but it seems like you could just time align video and audio streams manually in other contexts?
What this meant was an at least 2x and often 100x read amplification for truly random reads. There were some heuristics to disable that, too, and then drive controllers started doing this directly so some code to managed that interaction, issuing more early in the "pretty sure we want to read more" path but dialing back the scaling. I would not be surprised if this type of logic is implemented into modern SSD controllers directly. When you issue a random read, it actually reads several sectors ahead, just in case those are the next requests.
If we disabled that - How much faster would the random reads path be? How much slower would the sequential be? I suspect (very heavily, given the availability of enterprise drives with precisely this kind of tuning available) that there's a lot of tradeoff to be made.
"truly random" went the way of the dinosaur with the first CPU cache and memory bus wider than the size of read operations.
But as soon as you are pulling in extra data due to alignment you will have lost. Then you have to deal with prefetching being harder and other things that are arguably over-optimizations for sequential access.
The latter is why there are both tracks and track media in QuickTime while a lot of more playback-specific formats combine them. In the QuickTime architecture, a video track is a video track regardless of whether it’s implemented using Motion JPEG or Cinepak or MPEG-4 etc. Same with audio, subtitles, and any other form of data stream.
The high-level types and low-level implementations of those types are kept distinct so most tooling can be written at a higher level, while enabling a developer to reach down to a lower level as necessary or desired.
I'm curious why ISO decided to standardise QTFF for BMFF over RIFF. I also note that RIFF itself is based on IFF, from which Apple derived AIFF, which shares similarities with QTFF, so maybe that Emmy award should have gone to Electronic Arts for inventing IFF in 1985
...also, genuine question: are there any known predecessors of IFF's "chunked-sequential" file-format? (idea: reckon we'd find anything interesting in the docs for LucasArt's EditDroid's file-formats?)
What the QuickTime format supports that makes it appropriate for representing tape data—and for receiving Emmy and Oscar awards, among others, for its classical uses—is provide a bunch of additional abstractions atop its low-level implementation to represent typed and subtyped streams of data that (accurately!) span time intervals using arbitrary timebases. One of the reasons some other formats are less appropriate is that they don’t have that separation between type and subtype (track and track media, in QuickTime terminology), don’t allow arbitrary tracks, don’t support media with different native timebases, etc. This means you can’t easily write code that’s agnostic of the sample representation and CODEC details. (This was my issue with Ogg a couple decades ago—it’s a playback format for finished artifacts, not a format for authoring and editing.)
I figure instead of reinventing all of the stuff that’s already been done the past 30+ years for getting that sort of layering right, it’s much easier to just come up with some new track types and CODECs and fit them into an architecture that already supports most of what the problem requires.
https://www.martinreddy.net/gfx/2d/IFF.txt “Previous Work” sez —
“Our basic need to move data between independently developed programs is similar to that addressed by the Apple Macintosh desk scrap or "clipboard" [Inside Macintosh chapter "Scrap Manager"].”
“Macintosh uses a simple and elegant scheme of 4-character "identifiers" to identify resource types, clipboard format types, file types, and file creator programs. […] We'll honor Apple's designers by adopting this scheme.”
“We may also borrow some Macintosh IDs and data formats.”
And some other mentions too, like Postscript and
Some apps like MediaInfo will report an MOV with prores and pcm audio as MP4 though, so some people still seem confused.