Rearchiving 2M hours of digital radio, a comprehensive process
digitalpreservation-blog.nb.no
digitalpreservation-blog.nb.no
One of the challenges for the project was finding working 1-inch analog video machines. The team scoured the world for machines, working or not, and managed to get several running. There is one particular part that fails, Sony only has a handful left and the machinery to make them is no longer available. When they are gone the media will be unplayable.
The data complexity was due in part to there being multiple version of various quality, a program can be split across multiple tapes, and tapes can have multiple programs. So they need to ensure that all tapes of the best versions were selected, and also know what parts of other programs were on the tapes – bonus material. Finally, it was common to re-use video tapes to save money so it was possible that rare fragments of other material could be found at the end of the expected programs.
Nonono, save it all! Only a matter of time before they can be merged together to make a supercopy.
(But I get it: the practicality of saving even more data isn’t there)
For some program material there were low resolution copies, copies that were shortened for different purposes, copies that were edited for different markets... and so on. There was a decision tree to follow.
There was also a "never to be shown again" list, usually where people that were involved in the program (either on-screen or part of the production team) were later associated with crimes or very unsavoury behaviour, or the material was in some other way extremely controversial.
If so, what tools are they using?
[1] https://huggingface.co/collections/NbAiLab/nb-whisper-65cb83...
The medium and lower models wasn't quite up to the task, miss-interpreting some crucial words here and there, but so far I've been very pleased with the large Q5 model.
It directly translates into English well enough that subsequent English LLMs understand the meaning with high degree of accuracy, at about 5-10x real-time on my 2080 Ti.
Well, unless you have a Mac. Apple AAC is still the best quality encoder, but it's only available for macOS, and even then, the only UI officially supported by Apple is the "Music" app, so you're going to have to use a third-party command line or GUI wrapper. (XLD is good.)
Outside of that, quality of the various alternatives has changed over the years, but the Fraunhofer encoder, which they say they are using, is a good choice, even through licensing problems mean that it isn't included in ffmpeg by default. Frustratingly, the default build does come with an encoder called "aac", which isn't Fraunhofer, and has very poor quality. So, you have to make your own custom build.
Even then, the low-pass cutoff defaults to a weirdly low value, leaving the user to guess at, or consult ancient wikis, to try to divine a suitable value.[1]
It's unfortunate that AAC remains the best (by which I mean, most-supported) choice for modern lossy audio, because making it is still a huge pain.
[1]: https://wiki.hydrogenaud.io/index.php?title=Fraunhofer_FDK_A...
Besides that there a several docker images including static builds of ffmpeg including the libfdk bindings.
However, Am I the only one questioning a lossy codec suboptimal for archival purposes? Maybe the sheer amount of data is too mich for lossless...
If I read this correctly it means that those files will not be preserved due to DRM?
> Some radio broadcasts were stored as mp3 and wav files, with accompanying checksum files. Other broadcasts were only stored as mp3. Before the re-archiving process began, it was decided to generate new MP4 playback files from the wav files to replace the varying qualities of the old mp3 files.
In the cases where we only had mp3 files, the mp3 file was preserved as our master in the preservation environment, with a copy sent to the streaming servers.
That's why it was notable that in a couple of cases where the WAV was corrupt, they kept the original MP3 as being now the 'best available' copy. In no case did they transcode MP3->MP4.
For example, many filesystems from nix will scale just fine, self-check, and de-duplicate. However, accessing a path in a high-branching factor tree can cause problems for ls or rm etc.
Notably, for external BLOBS we found it simple to convert a filename into its sha512 hash with standard formatted media specific extensions, and include the extracted meta-data file in json (details of the encoding, label, and stats etc.) Thus, for quality of life improvements we would pack the sub-paths based on the file hash characters... so the smaller leaf path content entries were present in a given path (char[0] is local index, char[1..k] is the sub paths.)
It is strange, but when the files get big it is convenient to be able to audit each part in a decoupled way independent of the underlying filesystems/NFS/databases.
the file hash stored inside a database also infers the host OS file set label, and packed location
* the metadata is preserved along side the media in human readable utf8 json text, and thus the host node does not require knowledge of the media specific encoding during most operations (i.e. the file-server has minimal dependencies.)
* the file external-BLOB location is set by the size of k-1 hash string length used (k=128 chars in hash512 for example)
* normal CLI still works in each k-1 leaf path, and will be under ((2+w)16files) entries depending on how its implemented. Note, Windows usually demands k<124 sub-paths deep, and caps string lengths to under 260 chars.
BLOB file sets are trivially accessed/locked by many processes on the host OS, and additional files with extensions/analytics may be put beside the original media as json metadata etc.
* operates on top of whatever filesystem you are deploying at the moment, where we assume k<=128 will sit under your file system limits
* user and group read-only permissions are supported on the host NFS or JBOD
* Must turn-off auto-indexing filesystem search routines on the host OS node
* a database BLOB index can be rebuilt/ported from the archive tree leaf nodes
* source file corruption is self-evident (the metadata should match inside the json file as well)
* contents are obfuscated, but duplicate file checks are trivial (note this differs from block level de-duplication, which can be impractical if the data flow rate is high)
* tree-root-node sub-paths may be externally mounted on explicitly sharded volumes to share/backup the workload as needed... given the pseudo-random hash makes specific io-busy areas unlikely.
Mind you we only had to deal with only around 40TiB of sparse video data on that old project. Unsure if such a method would be performant for 2M hours of content.
Sometimes the metadata format inside media is versioned in nonstandard ways... if your intent is long term access support, than storing stats on the program and version used to read the file may be necessary to parse it properly in the future (even if a VM OS image snapshot with the codecs/parser is also archived as a read-only file.)
Best of luck =3