LTO Tape data storage for Linux nerds
blog.benjojo.co.uk
blog.benjojo.co.uk
If you want a LTO Tape solution with more bells and whistles you could check out Proxmox Backup Server's tape support:
https://pbs.proxmox.com/docs/tape-backup.html
We also rewrote mt and mtx (for robots/changers) in rust, well the relevant parts:
https://pbs.proxmox.com/docs/command-syntax.html#pmt
https://pbs.proxmox.com/docs/command-syntax.html#pmtx
The introduction/main feature section of the docs contain more info, if you're interested: https://pbs.proxmox.com/docs/introduction.html If you have your non-Linux workload contained in VMs and maybe even already use Proxmox VE for that it's really covering safe and painless self-hosted backup needs.
Disclaimer: I work there, but our projects are 100% open source, available under the AGPLv3: https://git.proxmox.com/
The whole tape management is in the common PBS API, so that'd be a bit harder to port but not impossible. For example, I made some effort to get all compile on AARCH64 (arm) and while we do not officially support that currently there are some community members that run it just fine.
So, maybe, but could require a bit more hands-on approach. If you run into trouble you could post in the community forum (<https://forum.proxmox.com>).
Retrieval costs are additional of course, and depend on how quickly you need access to the data, but if you just want to store data long term in case of disaster, $1/TB for multi-AZ replicated data seems like pretty reasonable pricing.
LTO-6 tapes hold 2.5TB of data (uncompressed), assuming you store 2 for redundancy, you'd need to find a place that will store them for $1.25/tape/month to break even, plus you're paying $25 for the tape itself, so over 3 years, that's almost another $1/month/tape. Plus the tape drive itself is around $1500.
You can use newer tape technology for better economies of scale, but your buy-in cost is higher due to the higher price of the tape drive, so you'd need a pretty high volume of data to break even.
From their pricing page:
S3 Glacier Deep Archive - For long-term data archiving that is accessed once or twice in a year and can be restored within 12 hours - us-east-2 (Ohio)
All Storage / Month $0.00099 per GB
When I last managed offsite tape backups, I never planned on really needing to retrieve the data -- I had the data on disk and on the most recent tapes. (I did do periodic restore tests)
If I had to restore the data, I wouldn't care how much it costs (within reason).
So, if your 100% confident you won't ever have to restore it, its a great deal. I was considering it for my personal backups but the ~3-5k a full restore would cost me is unpleasant on top of the $30-50 a month the storage would cost me. Over a 3 year timeframe that is $1000-1800 just in storage fees. So, maybe I will just buy another USB JBOD...
I suspect that over 100T, tape is still much cheaper, especially if the offsite storage is a fire safe at the CTO's house. Once your in the few PB range its a no brainier and actual offsite storage services start to look inexpensive. https://spectralogic.com/wp-content/uploads/white-paper-iron... says that its ~$1 a month per tape to store it at iron mountain, but I'm pretty sure there is a nice volume discount, one of the previous places I worked used them for tape storage and the actual storage bill was a joke in comparison to the pickup fee.
I always have local backups which I'd use to recover a deleted file or crashed hard drive, it would take weeks, maybe months to pull down my cloud backup unless I pay for a hard drive to be shipped to me (or in AWS's case, a Snowball, I could have 100TB restored in under a week using a Snowball)
> and you store it on tapes for me?
I mean, we can do client-side encryption and efficient remote syncs, so such a service would be possible to pull of with PBS, but no, we don't got the bunker or dungeon to shelve all those LTO tapes at the moment :-)
Tapes normally are not kept on site, to ensure that even a destruction of the office/datacenter/... like, for example, a fire does not make one lose all data.
With PBS this also allows to leverage the whole thing in a more flexible, efficient and powerful manner, as one can do tape-backups async, the deduplication layer can be used so that subsequent backups only need to add the delta of new backup chunks from the content addressable storage (note, the backups themself are still full backups, and one can configure a schedule for when a new media set needs to be created), besides that backing up file-level to tape is normally pretty slow compared to writing out big chunks that are ordered by physical disk location.
But before writing my backup to the cartridges, I tried reading their contents, and found that they actually came from a major film studio, with backups of raw animated film content on them!
Still did you try to recover any material and wach it?
Home users will really want to think in terms of the "raw capacity" imo. This is normally half of the advertised capacity for the older standards (I believe the newer ones have stronger compression that squeezes a bit more). LTO-5 tapes are 1.5tb raw, for example.
Maybe you'll get a little bit out of it, but a lot of the things you'd want to back up (and especially the bulkier stuff that really eats space) are already compressed. Family photo library, audio/video storage? JPGs are compressed, H264/H265 or MP3/FLAC/etc are already compressed. System images? A lot of application files are already compressed. Home user scenarios are not outlook mailboxes and database backups like the "official" scenarios.
Everyone would be better off thinking in terms of the raw capacity. "Compressed capacity" is nothing but a marketing gimmick. Even in enterprise use cases the compression ratios will vary, and the drive's transparent compression is unlikely to offer the most savings. If your data is at all compressible you should compress the backup yourself before sending it to the drive.
And most enterprises don't really care if their monthly backup requires 10 or 15 tapes. And zipping it all up beforehand requires even more space on the primary storage which is even more expensive than a couple dozen tapes
The tape drives themselves are much more of an issue than the tapes. It's a shame, because it necessitates moving data on older tapes to newer generation tapes after a few generations (which reminds me I have to do that with some LTO-3 tapes).
This is absolutely the pain point. Having to either keep multiple generations of drive around, and hope that they don't die in storage; or continually migrate your old tape to new tape is a huge pain in the arse.
That is my experience too. There is that time I got kicked out of the computer lab as an undergraduate because I'd created a number of newsgroups and they 'wrote' all my files... to what turned out to be an empty SunTape. That time I tried to recover a configuration file from an IBM tape robot and it took 14 hours. When I was successful with tape I always did a lot of practicing and testing. A sysadmin who taught me a lot (esp. how to get things done in a place where you need 'social engineering' to get things done) told me "you don't have a backup plan until you've tested it" and many people learned that the hard way.
Yep. Though that's what makes small-shop disk-to-disk backups easy, depending on the backup software used.
We use rsnapshot, which uses rsync and "cp -l" to make backups. So restoring is as easy as using cd to go into the appropriate directory and copying out the files. No special utilities needed. Yes, we encrypt the backup drives using cryptfs / LUKS.
The closest is things like inFECtious, which is more of just a library.
I would prefer something in go/rust, since these languages have shown really high backwards compatibility over time. Last thing you want is finding 10 years later building your recovery tool that you can't build it. Will also accept some dusty c util with a toolpath that hasn't changed in decades.
https://github.com/vivint/infectious
Ok I just dug up blkar, this looks promising, but the more the merrier.
It allows you to choose the amount of parity on the disk-level (as in: 1,2, or 3 disk parity in raidz1, raidz2 and raidz3). You can also keep multiple copies of data around with copies=N (but note that when the entire pool fails, those copies are gone - this just protects you by storing multiple copies in different places, potentially on the same disk).
[edit] To add another neat feature that allows for granularity: ZFS can set attributes (compression, record size, encryption, hash algorithm, copies etc.) on the level of logical data sets. So you can have arbitrarily many data stores on a single pool with different settings. Sadly, parity is not one of those attributes - that's set per pool, not per dataset.
Eventually i think i will start populating 2-4 sata drives on my 15-25W TDP celerons, atoms, and so on, and just give them away to people who need a computer for whatever. I'll even toss in a GTX 1050ti.
Check out section 4.9 in https://www.ibm.com/support/pages/system/files/inline-files/....
To be clear, this is a "user" level function that basically says "here is a CRC I want the drive to check and store alongside the data i'm giving it". It needs to be supported by the backup application stack/etc if one isn't writing the drive with scsi passthrough or similar. Its sorta similar to adding a few bytes to a 4k HD sector (something some FC/scsi HDs can do too) turning it into a 4K+X bytes sector on the media, that gets checked by the drive along the way vs, just running in variable block mode and adding a few bytes to the beginning/end of the block being written (something thats possible too since tape drives can support blocks of basically any size).
The problem with these methods, is that one should really be encoding a "block id" which describes which/where the block is as well. Since its entirely possible to get a file with the right ECC/protection information and its the wrong (version) file.
So, while people talk about "bitrot", no modern piece of HW (except intel desktop/laptops without ECC ram) is actually going to return a piece of data that is partially wrong because there are multiple layers of ECC protecting the data. If the media bit rots and the ECC cannot correct it, then you get read errors.
The drive can't necessarily even pick "wrong" data to send you because there are a lot more failure cases than "I got a sector but the ECC/CRC doesn't match". Embedded servo errors can mean it can't even find the right place, then there are likely head positioning and amp tuning parameters which generally get dynamically adjusted on the fly. This AFAIK is a large part of why reading a "bad" sector can take so long. Its repeatedly rereading it trying to adjust/bias those tuning parameters in order to get a clean read. And there are multiple layers of signal conditioning/coding/etc usually in a feedback loop. The data has to get really trashed before its not recoverable, but when that happens it good and done. (think about even CD's which can get massively scratched/damaged before they stop playing).
man par2Because of the RTOs and backup windows, the supporting infrastructure was fast. The caching layer stuff was the fastest disk in the data center by far, and the team was a small, tight group of people who basically honed their craft by meeting auditor and other requirements. The management left them alone and they did their thing.
That was about a decade ago now; those guys have all moved on to really big things.
(Spectra Logic's tape libraries run FreeBSD too.)
Also last I checked freebsd was used for their disk product, not tape.
Let's say you're using LTO tapes as an archive. Did you know LTO tape itself is abrasive, but that abrasive is meant to wear over time with the intended use of the cartridges, which was backups?
If you use new tapes a single time, the abrasive doesn't wear and destroys the tape heads. You will go through a drive head at month, running the drives 24/7. I had a library used as a genomic storage archive with 8 drives (always write, almost never read), and two were constantly out of service, as we averaged two head replacements from IBM a week.
This is much less a factor on use tapes that have been run through a drive a few times.
A few factors which may have influenced what you experienced:
* The quality of the tapes could be variable. In my experience, some branded tapes were significantly inferior to others.
* If the drive ran hot, then that may have contributed. IIRC, IBM's LTO-3 drive ran very hot.
* If you don't write data to the tape fast enough, it won't stream. It'll shoe-shine back and forth, as it runs out of data, repositions backwards on the tape, and resumes writing. I think this might affect the tape head life.
We did have shoeshining issues in testing, but increasing the amount of caching fixed that. Never heard of any throughput issues in production, but .. .edu so you know how well we monitored. That was a software issue anyway.
I think it was LTO5 era, but I don't rightly remember.
The IBM dude who handled all the hardware support would take a look at everything, nod, and replace the drive. I took him out for beer once and that's when he told me about the issues with the tapes. I left for greener pastures before that was solved, but it was going on for a good year.
Maybe he liked the food trucks outside the building, or maybe it was cheaper for them to replace the drives than actually help us fix the problem. Anyway, thanks for the insight! Glad I don't work on hardware anymore.
I've not seen it on LTO. Where I work we either had very large tape libraries, with 25+ drives in. We didn't have drive affinity, so if that happened I would get an alert.
The other team used to import bulk data by receiving tapes from all over london and beyond, there must have been thousands of drives writing and reading that data. Plus we didn't buy fresh tapes, and they were dropped, thrown, left in the cold/sun, all sorts.
I think LTO is pretty solid.
The gist is that HP heads were harder material than IBM heads and if you used alot of out-of-spec abrasive MAXELL tapes, especially always fresh ones, then on IBM drives, the drives would suffer from premature "pole tip recession" and could no longer read or write any tape. HP drives seemed less affected.
To find out what manufacturer made the tape, you had to insert the tape into a drive and read the MAM chip over SCSI/SAS. With tapes basically everyone did alot of rebadging, so if you got a "HP" tape it could be anything from MAXELL to Sony to Fujifilm. "IBM" tapes were usually Fujifilm, i.e. good to use everywhere.
Actually, it isn't very significant. Price factor of 2.5. I had thought tape storage was cheaper than that. And then there are the drives: A drive to write (3,000 GBP for LTO-8), and at least a couple more drives for reading tapes.
At this price ratio, I would say that ease-of-use and safety/robustness of the backed-up material are more important considerations.
At least it's likely I can find a USB port 20 years from now, or a DVD reader (they are still being manufactured today, when even more than 20 years have passed since their introduction, and they are even compatible with much older CDs...).
Also related is that tapes can easily be transported around/offsite, literally thrown in the back of a truck as they are. Try doing that to hard drives and see how many start throwing bad sectors after a round-trip.
https://www.newegg.com/global/p/pl?d=hot+swap+hard+drive+bay
... and the actual disks will usually be stored offline. So, no accidental erasure. But I agree that tapes are probably less sensitive to transportation.
Also, some nit-picking: Energy prices in Germany are currently MUCH higher than that. We moved and had to get a new contract. Close to 40c/kWh. This makes your point a bit stronger.
//edit: Also2, when doing the math I realized I should first transcode suitable content to h265 (per TB saved the necessary power is cheaper than a new disk), and as a second step replace my four or five remaining 1 TB HDDs with a single bigger drive to reduce the idle power draw (the NAS is on a btrfs mixed-size RAID1).
As a bonus, they can generally be used to offline clone drives.
Why?
I've personally never experienced that scale, I'm sure the industry has some recommended ratio of drives to tapes.
Also, tape is slow. The MB/s is pretty nice on the latest tech, but a tape is pretty big, so if you have a lot of stuff it'll take a good while. Google says it takes 9.25 hours to write a full LTO8 12TB tape. Which means that if you have a sizable backup, in case of needing a full restore you might well spend a whole week reading tapes.
And that's not accounting for that something might suddenly break, and the time where that becomes important is right when you need something restored urgently.
Old-school DECtapes were actually random-access, seekable block devices! They help 578 blocks of data, each block being 512 bytes (or to be more period correct, 256 16 bit words), so 144kiB. They could be read/written in both directions. When mounted on a tape drive, the OS (like DEC RT-11) would treat it just like how a PC DOS computer treats a floppy: you could get a directly listing, work with files, etc. The random access nature caused the tape to move quickly back and forth across the tape head, a process known as "shoe shining".
Edit: Also, tape formats tend to come in two scan methods since they are generally wider than the tape heads (which frequently are actually multiple heads). Helical scan (think VHS/DAT) and serpentine. LTO is serpentine which means it writes a track from beginning to end, then writes the next track in "reverse" from end to beginning, then the next track again from beginning to end. Back and forth until it hits its track limit.
So basically just about every modern drive reads and writes in both forward and reverse.
Although shoe shining (backing up to start the next read/write) is still a thing despite variable speed drives which try to speed match to the data rate the host is reading/writing at.
CP/M required booting from a block device and as far as I know the Coleco Adam was the only computer which could boot CP/M from a tape. Once booted to CP/M the tape drives were treated just as floppies.
Fun times.
Then there's the (to me, still open) question of how to best use the actual tape storage capactiy... Since my hardware is newer than LTO-5, LTFS (https://github.com/LinearTapeFileSystem/ltfs) is an option for convenient access, especially listing tape contents, but that could make it hard for other people down the line to restore data from the tapes I create.
It's probably safest to assume that tar will always be there, at least wherever there's tape, too. GNU tar also handles multi-volume/-tape archives, which seems like a necessity if you need to back up amounts of data that exceed a single tape's capacity. Then again, if you want to use encryption with actual tar (important for the kind of data I need to archive), your only option seems to be piping the whole archive through something to compress the stream, which will make accessing individual records in the archive opaque to the drive itself... and you can't just dispose of individual keys to make select parts of the archived data go away for good, either.
Also, I would like to conserve as much tape as (conveniently) possible in my archiving adventure. There's "projects" (i.e., top-level directories of directory trees) that consume more than one tape of their own, and then there's smaller projects that you can bin-pack together onto tapes that can fit more than one such project.
I've started implementing a small python wrapper around GNU tar to solve a number of these problems by bin-packing projects into "tape slots" and also keeping track of tape-to-file mappings in a small sqlite database, but a workable solution for the encryption problem(s) is not something I managed to come up with yet... If someone has an idea (or better yet, a complete and free implementation of what I am trying to hack together :)), please be so kind and let me know!
I had previously used Blu-ray for backups, I think they are fairly durable if you have a dry, cool place to store them, but if you have to find date spread over 20 discs, it’s quite a pain. Now it would feel better if Redhat or Suse or somebody cooked ltfs in to their products as a first class thing. I think the catastrophic recovery process would involve building enough of a system to download and install ltfs to access the tapes. I could also create a “recovery system” and then just tar that on to a tape too.
My advice strategy has been to keep things relatively warm and when ltfs starts to feel like a liability then I’m going to move the whole archive to something else, fortunately it’s not 100s of br discs, it’s tens of tapes so it will take some hours but it’s mostly waiting on data to stream.
Rather than having one possible tape that can be corrupt, you know have more than one. Any single failure will corrupt the whole archive, and you've lost (easy access to) the data.
Better to break up the source volume somehow, and have each tape self-contained.
LTFS seems dead and tar will not nearly extract all possible good stuff from LTO technology.
Plug: I repackaged the open-source Bareos tape software for Arch Linux a couple of days ago. Here are some screenshots of the web interface to give an idea:
https://seitics.de/files/bareos/screenshots/
One up in the tree are AUR build files and quite a few moderately complex new examples, covering scripting to run ZFS snaphot create/delete around jobs, how to turn on/off external tape units using a USB box to cut down on power and noise (involves drilling a hole on the backside of your formerly 2000 USD drive!), how to write disaster recovery files to tape as well, in case that's all you got left, etc. Software is AGPL3, being a Bacula fork, my contributions here are CC0. Been running it on Arch since around 2017 with multiple LTO drives.
https://superuser.com/a/71239/37009
For example the M-DISC (https://en.wikipedia.org/wiki/M-DISC).
> Millenniata claims that properly stored M-DISC DVD recordings will last 1000 years.
I would then store those bits of media geographically and politically distributed. And I'd store it with paper documents describing the encoding, the file formats, the compression, any encryption, etc. I'd also include a few physical computers (eg. a raspberry pi or laptop) that has all necessary software to read, decode, and display the data. Set it up to be usable by a non-expert - in 1000 years time, there may be nobody who knows how to use a shell or open a file!
And I'd have a 2nd copy of the whole lot on hard drives connected to the internet for day to day serving of the data to people who need to see it. All the stuff above is only needed in case of organisational failure, war, civilisation collapse, etc.
If you do that, I would really love to read some more details on how you actually organize that.
The above is what I have set up for some organisations who want to keep data for thousands of years.
There are other bits to the process, like every 10-30 years, repeat the process with the new data and the old data. This time, the 'old' data will be much smaller compared to the storage mediums, so keep that data uncompressed, preferably unencrypted, and un-erasure coded in every geographic location. That removes many barriers to access the data, and increases the chances someone that finds it in 200 years bothers recovering the data.
Sadly in the future world there is a high chance some of the data is copyright, illegal knowledge or gdpr-impacted and all records need to be erased. There isn't really a good solution to that. It's almost impossible to protect against future humans wanting your data gone.
With the right tools you can read back the data and see how much degredation had occurred (even though all the data was readable). Using that info, you can have a good chance at extrapolating the time till the data becomes unreadable.
(I believe the final consensus pointed to arrays of HDDs where most of them are powered off, and the number of "live" drives per rack is bounded to allow high density/low cost, hence the need for access time/service level bounds, but the BD-XL idea is still intriguing!)
With the consumer discs, even considering cost per GB, the amount of effort required to handle a large library of low-capacity discs is just too great even if the cost is a little bit better. 128GB discs would have been very usable 5 years ago but again, those discs were never affordable to consumers, and the 25GB was still some effort at that time. Today even 128GB is not all that much, as data has grown. As far as I know there is nothing realistic on the horizon to replace blu-ray with higher capacity either, if movie content started being released in 8K it probably would be something like BD-XL with AV1 encoding (or maybe H265 again), not a fundamentally new iteration like DVD->BD.
The future for consumer storage seems to be SSDs and hard drives for fast and slow/bulk storage, and cloud for nearline storage. Tape is still relevant for enterprises though especially in automatic libraries.
The theory of hard drives being shut off/ powered on dynamically in a rack sounds intriguing. Sounds simple and yet difficult because of the rare usecase, i.e. no commodity hardware available. Maybe something to test out for colo backups to keep power usage down and prolong disk health.
Are there any commercial products that use this technology?
I wonder if Sony's ODA format could ever become more popular in the consumer market. I've never heard anybody mention it before.
Alternatively, I wonder if there could even be a "prosumer" robotic library system for common optical disks, something like a desktop archival data jukebox...
https://hackaday.com/tag/cd-changer/
http://hackalizer.com/jack-the-ripper-is-an-automated-diy-di...
http://hackedgadgets.com/2006/06/07/cd-changing-lego-robot/
Yes, you definitely want sth like that. And further extend it.
Still, aside from it being prohibitively expensive (LTO-8 seems like something of a floor given the size of Blu-Rays), tape backups seem to be a hard area to get into. I did some crappy little DLTs in the 1990s but nothing since, so "what software?" and the like questions are all new to me. And this would be with just a single drive, not even a library.
Tape like LTO 6 is quite cheap these days and will last a long time when stored vertically in a dry, cool, dark place. Also at 2.5 TB net capacity per tape with barely used tape prices of 5-10 bucks one can easily write a couple more tapes just in case one is damaged. I think for leader pin damage there is a repair kit and if tape rips in two I remember there was a repair service by IBM which can fix even that and send you a fresh tape with the data or a repaired tape.
I use an ancient LTO2 drive for last resort backups that are off cloud and off premises. Its more peace of mind than practical on a daily basis but I did find myself restoring a few files a couple of weeks ago as I had fat fingered an rm command. It was quicker than getting them from S3 glacier.
For my archival use (the reason why I got into this in the first place) I do not encrypt nor compress the data going to tape. For server/desktop backups. they are compressed and encrypted.
Just to note that tape drives have built-in compression that generally is done transparently in the background. So while using something like zstd (per the article) may get more bits on a given tape, there is some compression that one gets "for free" without doing anything at all.
* https://en.wikipedia.org/wiki/Linear_Tape-Open#Optional_tech...
* https://en.wikipedia.org/wiki/Magnetic_tape_data_storage#Dat...
> Drives above LTO-4 have built-in hardware encryption, however I would steer away from using it and instead just encrypt data yourself (possibly with the tool I helped make called age!). Like most things, you should also consider compressing your data before encrypting and writing it to tape. LTO tape capacities are often quoted in their “compressed capacity” which is a little cheeky since it assumes basically over a 50% compression ratio, this is not at all likely to be true if you are writing video or other lossy mediums like images etc to the tape. I generally run my data through zstd to compress and then age to encrypt. Zstd and age are quite fast and I’ve not found them to impede performance noticeably.
If someone is not familiar with tape drives, I think it would be easy not to realize that the compression is built into drives like the explicitly called out "built-in hardware encryption".
I used to work for a government agency. We ran backup tapes that rotated out through a degaussing machine that spun them around for like 10 minutes to wipe them. It’s not common to have, but it’s definitely easy.
So, it would take a very dedicated nerd indeed.
A downside of encrypting yourself is that you can't benefit from the hardware compression either, hence the articles suggestion to do that in software before compressing as well.
Personally, my tape writing workflow is: dar (per file compression, skips uncompressable mime types + encryption) followed by par2cmdline with 30% redundancy. For comparison: CD-ROMs have 33% redundancy information (8 bits per 24 bits, CIRC encoding).
A tape-drive is your read-heads. The tape is like a platter. The tape-library / jukebox is just a robotic mechanism for switching tapes into and/or out of the read-head.
----------
If you need a Petabyte of uncompressed storage, you can reach it with a tape-library consisting of 84 LTO8 tapes (12TB each). If read/write of 400MB/s is sufficient, one tape drive is sufficient. If you need faster access speeds, you buy a 2nd, 3rd, or 4th tape drive.
So lets say you need 2GB/s read/write speed and a petabyte of storage. You simply get 4x LTO 8 drives, 84 LTO8 tapes, and stick them into a tape library of some kind.
You then buy a certain amount of SSDs + HDDs sufficient for caching, so that you can read/write to this tape library at sufficient speeds (especially since it could be many minutes before a specific byte is accessed).
That is: if a file A is modified, LTFS "represents" this by marking fileA as dead, then copying most of the contents of A to the end of the tape... with the few modifications in the correct location.
Because tapes have so much capacity, this ends up working much better in practice than you'd expect. Especially if your backup-like software (or Apache-log files or whatever you're archiving) are largely append-only (ex: Wikipedia dumps or whatever huge data-storage requirements you're doing).
-----------
Allegedly anyway, I don't really use tapes myself, but I amuse myself by looking at their specs / usability every now and then. It seems like the details of LTFS depend on each manufacturer's system, but it seems like LTFS itself is somewhat cross-compatible.
Some LTFS-library docs suggest that different tapes are different folders. Others suggest that different tapes are mount/unmounted (which sends a sequence of commands to the little tape-robot to fetch the correct tape, which it does so by scanning barcodes or maybe going to correct slot)
-----------
Tapes are meant to be used in a "jukebox-style" application IMO. If you need just one tape, you're doing it wrong. If you need 80-tapes, then doing it manually is obviously wrong, you should pay a bit extra for that "jukebox / tape-library" robot to fetch the right tapes automatically.
> Does it already exist as open-source?
Red Hat claims LTFS support. But I imagine this is only "disk-level LTFS", and not "library / robot LTFS".
I know IBM talked about it before. I think here's their PDF of some university that did this: https://irp-cdn.multiscreensite.com/b076c94c/files/uploaded/...
---------------
The distributed data-thing is pretty common in the supercomputer world. I know that the nodes on the "Summit" supercomputer each have a local SSD for caching, and a centralized hard-drive array for longer-term storage. I'm pretty sure the software just manually copies the data around as necessary (that is: any software written for the supercomputer can copy data from the central storage into a local SSD).
Supercomputers can have that kind of simplicity, because you buy 3000+ nodes with the same specs. So you know all nodes will have that 1.6TB SSD available. No matter which node your software ends up running on, the SSD location is set and obvious. It doesn't matter how the data is archived: you read it in and wait (it could be tape, hard drives, or any other mechanism)
One method for business continuity would be LTO in a fire-proof safe to bootstrap the business back into life. Malware couldn't infect or encrypt the data on tape, because of the air gap and inherent "offline" nature.
Lots of ways for the "all data was destroyed" situation to come out. Tapes are kinda-sorta useful for this situation because capacity is so so so cheap, you can keep full copies of the important data in your tape library pretty easily.
Anyone who cares about their stuff needs to practice a full emergency RESTORE.
I have met very few people who actually do that. For most systems I've seen, the first full test of the restore process is a very scary first production usage of the restore process.
Which is very exciting, sure. I don't want excitement in my data management life.
(I actually see weekly test of onsite backup power at the local banks, and at some large commercial kitchens. Those diesel generators are very loud. I've never seen systematic test of UPS or generators in a front-office environment.)
Seems like one could go quite far in terms of performance with just some basic HW and an FPGA. Is there significant difference between multiple generation of the tapes themselves, or is it just data encoding patterns that change?
More specifically, I was a bit appalled by the "magnetic erasing" bit. Seems like DRM to me, on a medium that is conceptually extremely simple.
One could probably take a VHS drive and convert it to a data drive, unless I'm being naively optimistic about it?
https://en.wikipedia.org/wiki/ArVid
> More specifically, I was a bit appalled by the "magnetic erasing"
Nobody laments what there is no 'low level format' for HDDs anymore.
A disk image plus compressed, encrypted then forward-corrected `btrfs-send` snapshots sounds quite efficient to me. Take your hourly, etc snapshots to a regular disk, write monthly ones to the tape until fills up, then take another tape and repeat. The downside is that you need to replay multiple diffs.
Or would it be a good idea to make more frequent writes? I'm not sure what best practices are when it comes to tape and backup.
But I don't expect to restore more than a few gigabytes at a time from that.
It would take me a week or more to download a terabyte of data. I have very little power over internet connection speed, and there are very few alternatives here. I believe there are two different vendors providing connectivity to our town, and you can pick between four retail resellers.
With those limitations, I have tested a full restore process exactly once. That's not good enough.
Data at rest on LTO or offline hard disk is something I can control. Distributed offsite storage, too. Restore within 12 hours, I can do that.
The downside to tape or cold disk is more in the management of hourly/daily/weekly backups: you have to provision a media rotation schedule, whereas that's sort of built into an online cloud storage service.
Haha! moment
I've been thinking about moving some backups to tape because of the stated durability and good cost (~$7-$12+/tb for used LTO4 tapes+drive, depending on how much storage I get).
But if the drives don't also last 15+ years, then I think the cost could be significantly higher, especially if I'm shelling out who knows what for drives a decade after their low point on the resale market.
Is useless for a long-term, both from $/GB and from sanity standpoint. LTO-5 is more than 10 years old already, there is no point in trying to use LTO-4.
> How durable are LTO drives?
Just look at the comments here. 99% of time it just works. In 1% you throw the shit out of the window (well, you want to).
Also:
> Both drives and media should be kept free from airborne dust or other contaminants from packing and storage materials, paper dust, cardboard particles, printer toner dust etc.[49]
And "5 Environmental" in [0]
https://docs.oracle.com/cd/E21419_04/en/LTO5_Vol4_E4/LTO5_Vo...
unfortunately they dont seem to have an open vcs for the source... (other than really old versions on github)
other than that there is mhvtl:
Let me know if you'd like an English version :)
Who uses this ?
It’s much easier to store tapes in a fire proof and water resistant safe than to find a fire and water resistant storage.
So you can keep you backups in disk, but last resort disaster recoveries should be on tape somewhere.
Gmail has tapes[1]. And they saved me their asses at least once. This can give you a hint of how important and how much use tapes get.
1 - https://www.datacenterknowledge.com/archives/2011/03/01/goog...
Lets start with tape has two types of head positioning commands, locate and space. Locate is absolute (and mt calls it seek), and space is relative. Mt is generally using space (although one can read the current position with tell then do relative space) for all the commands that aren't "seek". Hence the mt commands are things like "fsf" which is forward space file (mark), or "bsf" for back space file (mark). At some point in the past someone thought that each "file" would fit in a tape block, but then reality hit because there are limits on how large the blocks can actually be (in linux its generally the number of scatter gather entries that can fit in a page reliably). So there are filemarks, which are like "special" tape blocks without any data in them. Instead if you attempt to read over a filemark the drive returns a soft error telling you that you just tried to read a filemark. There are also "fsr" for forward space records with are just the individual blocks forming a "file".
So back to seeking. If you man st, you will notice that each tape drive gets a bunch of /dev/st* aliases, which control the close behavior/etc, as well as some ioctls that match the mt commands. The two important close behaviors to remember are that if the tape is at EOD due to the last command being a write it will write a filemark, then rewind the tape unless a /dev/stXn device is being used, in which case it will leave the head position just past the FM (this is actually a bit more complex too because IIRC there may be two filemarks at EOD, and the tape position gets left between them).
This allows one to do something like "for (x in *.txt); do cat $x >> /dev/st0n; done" and write a bunch of files separated by filemarks (at the default blocking size which will be slow (probably 10k), replace the cat with tar to control blocking/etc). Or if you want to read the previous file `mt -f /dev/st0n bsf 2` to back space 2 filemarks.
Now, the actual data format on tape is going to be dictated by the backup utility used to write it. Some never use filemarks, some do but as a volume separator (eg tar), old ones actually put FM's between files, but that tends to be slow because it kills read perf because it takes the drive out of streaming mode whenever you either read over the filemark (not the part on man st about reading a filemark).
Now you can pick which file to read via "mt -f /dev/st0 rewind; mt -f /dev/st0n fsf X; cat /dev/st0n > restore.file"
There are also tape partition control commands, and tape set marks and various other options which may/may not apply to a given type of tape. Noticeably there are also density flags on the special file (some unix'es) and via mt. LTO for example doesn't have settable densities because its fixed by the physical tape in the drive. Some drives STK T10K/IBM TS11X0/3592 can upgrade the tape density/capacity when used in a newer drive.
That got long...
`man st` does a much better job of explaining the difference between `fsf` and `fsfm` than the `mt` docs:
MTFSF Forward space over mt_count filemarks.
MTFSFM Forward space over mt_count filemarks. Reposition the tape to the BOT side of the last filemark.
I’ve got a project that would work very well with linear data streaming, but it looks like even the cheapest of used tape drives probably only break even with HDDs in the tens of terabytes. Ah well, one can dream...Or at least did at one time.
(I used to work for a four letter computer corporation doing enterprise technical support, mostly on tape-based products.)
A lot of 'enterprise' backup software is also now coming with hooks into cloud storage (e.g., S3 APIs), but then you have to worry about bandwidth and the time it takes to get the bits offsite at "x" bits/second.
Of course you also have to worry about retrieving the data in case of disaster per the Recovery Time Objective:
* https://en.wikipedia.org/wiki/Disaster_recovery#Recovery_Tim...
Also: a backup has not happened until you try and succeed your recovery process.
Reminds me of the saying that the fastest throughput is achieved by a 747 full of hard drives.
> Also: a backup has not happened until you try and succeed your recovery process.
A thousand times this.
(And in particular the high-level overviews are important because tapes are wear items, you only have on the order-of a hundred or two (don't remember the exact figures) full tape reads before the tape wears out, so this is something you want to go into it knowing a strategy and not making it up as you go!)
Since it's complimentary to this discussion I'll link a few:
https://www.cyberciti.biz/hardware/unix-linux-basic-tape-man...
https://databasetutorialpoint.wordpress.com/to-know-more/how...
https://sites.google.com/site/linuxscooter/linux/backups/tap...
https://access.redhat.com/documentation/en-us/red_hat_enterp...
https://access.redhat.com/solutions/68115
That is, unfortunately, essentially the apex of LTO tape documentation in 2022, as far as I can tell.
Do note that in terms of tape standards, LTO-5 is an important threshold because that's where LTFS support got added, and that's the closest thing to a "normal" filesystem abstraction that's available for tape (sort of like packet-formatted CDRWs I guess, in the sense of presenting an abstraction over the raw seekable block device). There is also very little documentation on init, care, and feeding of LTFS iirc - and again, it would be nice to know any pitfalls that might cause shoeshining and tape death. Although I suppose in practice it's mostly going to get used more in a "multi session" scenario where you mostly aren't deleting files, you write till it's full and then maybe wipe the whole tape at once, and it's just a nice abstraction to allow the abstraction of "files" rather than sequential records (tape archives/TARs, in fact!) along an opaque track with no contextualization.
If their local snapshots are dead too, or they look for it and realize they can’t find a copy of something they thought they had, it’s often because they needed that data right away and it wasn’t there when they went to get it. Hence ‘user expectations’.
That’s not in a catastrophic case (which rarely happens) that’s the ‘bob just realized he deleted the folder containing the key customer presentation last Friday’ or ‘mary just tried to open the contract copy she needed and it’s corrupted’.
If it’s a once in 10 or 100 year or whatever event, a 1-2 day turnaround is not unexpected and everything else is probably broken too. The file deleted or something got screwed up happens more often and slow response there grinds things to a halt - and causes a lot of stress knowing it’s not ‘solved’.
A lot of companies are pretty terrible at figuring out catastrophic tail risks like this too.
Tape has it’s niche, this is one of them for sure.
Weren't they even originally marketed as such? At least the external enclosures that people like to shuck are usually called something like "my backup".
Whats grand about tape is that its still faster to dump to your library, eject the magazines and store off site.
Whilst you can do that with HDDs (think snowballs but bigger) its a lot more expensive and error prone.
Tape serves a purpose, but thats pretty niche by todays standards.
is that a DECTape?