Tape Storage Trundles On, Increases Yearly Volume to 128 Exabytes
tomshardware.com
tomshardware.com
The most common incorrect assertion is that drives don't fail THAT often, so why not just back up to drives and put them on a shelf? But to be generally safe, all data would have to be on at least two drives, which changes the cost noticeably.
Older, used LTO drives are actually quite affordable, so I use LTO-4 and LTO-5 tapes to back up all sorts of things for both myself and for my clients. It's surprisingly easy and inexpensive, and pulling files back off is as simple as running pax.
Take the Cartesian product, in fact, of those factors. I want location isolation, isolation from application layer problems, isolation from storage layer problems, isolation from media failure.
Edit:
They do RAID4 on the tape level. 5 tapes, one of them the parity tape.
When that fails: Reconstruction at sub-tape level. "We don't really have data loss."
"We make the backups as complicated and take as long as they need. The restores have to be quick and automatic. Recovery should be stupid, fast and simple."
Retrieval (recovery) from disk is easier and faster - ability to randomly read and write. Tapes are sequential - to get to a data which is only at the end, one has to go through all the tapes that is there before that.
As an off-line archival backup medium, access times really are not a concern in the real world.
https://en.wikipedia.org/wiki/Linear_Tape-Open#Physical_stru...
https://en.wikipedia.org/wiki/Linear_Tape-Open#Positioning_t...
These seek times are seldom a problem. If one would keep on HDDs as much data as it is normally kept on tapes, i.e. at least a few hundred terabytes, one would have the same slow seek times as someone would have to plug and unplug external HDDs, unless the storage would belong to a very large company that could afford the high costs of keeping online many HDDs.
As an individual user, I keep in my computer, on an SSD, an index with the content of all tapes. When I need some file, I get it in at most 5 minutes, which includes starting the tape drive, going to a cabinet and taking the tape from a shelf, inserting it in the drive and waiting for the seek and copy operations.
This is perfectly acceptable, because I do not need files from the tapes every day. Storing the data on external HDDs would not reduce the time wasted with manual operations, it would be much less reliable and it would be much more expensive for great amounts of data.
The sequential transfer speed of tapes is greater than that of HDDs. Therefore, if after a seek you copy large files, e.g. a movie in BluRay format, the whole seek + copy operation takes less than from a HDD.
Tapes are for archival storage, not for 'accidentally deleted a single file and need it back' type backups.
Once, we had to set up a hollow Sharepoint and restore a Sharepoint backup just to get someone's deleted file.
Now, do I think this is a good idea? No. Frankly, people who cause these kinds of things need to see "This cost X hours to recover, stop doing that" as feedback.
but, when you are recovering files, its rarely a random io type deal.
Even then, you generally just dump to the nearline/tape cache and fiddle with the data there.
A decent 25 drive tape library will easily saturate a 100gig network link, and its perfectly possible to add more drives to get more IO.
Tapes as part of tiered storage is something that is really powerful. Yes, flash is cheap, but not cheap enough to be used as long term backup (ie legal hold or general archive.)
Keep your expensive fast storage near your clients, the cheaper less fast, but more voluminous a stage away, and then dump out to tape after n weeks with no access.
My issue with tape is that unless your operation is big enough to justify a robotic tape library as a budget line item with support contracts and all, then you're down to paying someone (or eating the cost yourself) to physically swap tapes which is much more expensive and boring than deploying a ZFS server with gobs of disk that once configured, just sits in the rack and does its job quietly.
The whole point of something like tape is having an off-line copy of your data, ideally in a separate physical location. A second server with a bunch of disk in the same location connected to the same network isn't that, and will not save you from ransomware or natural disaster.
If you use even crude tools such as clusterssh, managing a bunch of machines isn’t linearly harder than managing one.
While tape renders itself easily to be offline storage (it’s offline as soon as you eject it after all), you can have that with remote servers that pull data to back it up instead of receiving pushed data. If no server can push data to any other server, only pull from them, a ransomware attack becomes a lot harder.
Also, while tape in a warehouse (or on the desk) is offline, tapes in the robot are no more difficult to destroy than hard disks. They are just slower.
And hope your online backup doesn't get hit by the same group/malware that took out your primary production system.
Offline backups help with certain types of risks.
It's not about how the data gets there, it's the fact that the system is online and hackable just like any other system.
Good luck wiping offline tapes (Mr. Robot not withstanding).
The trick I’ve always found is to figure out where you are on that inflection point. And it’s hard. Is 1PB enough to justify tape? (Which seems like a crazy question to me - I remember having megabyte sized tapes).
Usability isn’t something you can ignore, but it’s not like hard drives are perfect either—do you buy a big NAS / SAN setup and plug drives into it? Will it get full? And tape has the advantage that it’s completely immune, out of the box, to ransomware.
I think there are four cases that really scream for tape.
1. Data hoarders, who just want to store as much data as possible for the cost. There’s an r/DataHoarder subreddit if you’re curious about these people.
2. Archivists, who want to store lots of data long-term. Tape is a lot easier to store. I recommend that archivists standardize on a specific generation of tape for as long as possible and don’t mix generations (don’t mix LTO4 and LTO5, for example, despite the fact that the drives are advertised as working with both).
3. Companies with recordkeeping requirements, like SOX (Sarbanes-Oxley). Tape is just really good for that. It has a way of surviving problems in your IT department.
4. Companies with enough data that they can put a line-item on the budget for backups, and justify the operational cost of tape—support contracts, keeping staff on hand who know how to use tape, that kind of thing.
Of these, the data hoarders are going to use the 150 TB break-even point just because they want as much data as possible. Everybody else is going to make decisions based on other factors, like staffing or compliance. There are a number of gotchas, like problems with mixing tape generations, the prospect of using robotic tape libraries, and support contracts, that make the tradeoff much more nuanced.
Taking into account that for long term archiving it is necessary to store 2 or 3 copies of the data reduces proportionally the threshold above which tape is preferable.
The storage cost per TB includes not only the purchase price but also the lifetime of the media, e.g. a HDD model with a warranty of 2 years cannot be trusted to store data for a longer time.
Tape is guaranteed for 30 years, but the storage time is normally reduced to about 10 years by the risk that the corresponding tape drives may become obsolete.
At the places where I used tape, we used a more efficient encoding for archival tape backups. Rather than storing 2 or 3 copies, we used forward error correction with something like 30% overhead. This, then, gets even more complicated to evaluate, because it multiplies how “hungry” your backup / archive system is for ingesting data in order to remain efficient & still write data out to tape by whatever deadlines you have set. If you store 3 copies of data on LTO-8, then you write data in blocks of 12TB, with one copy on three tapes. If you use forward error correction, you might do something like write out 96TB at once on 11 tapes. You use less than half as many tapes in the long run, but you need to feed the machine faster in order to meet deadlines.
It is possible to use RAID-5/RAID-6-like encoding schemes that can survive the complete destruction of 1 or even 2 storage centers while using less tapes than with simple copies, but such encoding schemes can be used only by very large organizations, which use more than 3 separate storage centers.
> It is possible to use RAID-5/RAID-6-like encoding schemes that can survive the complete destruction of 1 or even 2 storage centers while using less tapes than with simple copies, but such encoding schemes can be used only by very large organizations, which use more than 3 separate storage centers.
Scenarios where two storage centers are destroyed—that’s extreme. The most paranoid scenarios I’m normally willing to entertain are along the lines of one data center burns to the ground in a generator fire, and somebody drives a truck full of backup tapes into a ditch and they’re all covered with mud and sand.
Tapes have a high enough failure rate that you benefit from forward error correction and you benefit from planning to handle individual tape failures. This includes stuff like the tape leader breaking, somebody losing a tape, damage during transport, water damage in storage, etc.
There’s a layered approach here, where you plan for different disasters at different levels of the stack. Each layer exposes some certain failure rate to the layer above it, and deals with some certain failure rate at the layer below it. When I think of backups, I often imagine a top-level data storage system that has a geographically distributed hot backup, and then an offline cold backup. This lets you survive complete destruction of one data center, or lets you survive a catastrophic software bug that destroys data (and a bunch of tapes are damaged on top of that). Pretty good baseline, IMO.
Basically anywhere where you have a lot of data that has to be retained indefinitely for regulatory compliance or practical reasons is a great case for tape. But yeah, the robotic library and the service costs are pretty high until you hit a huge amount of data.
Source: supported a couple of large medical installations for a couple years many years back. Can confirm that you're dealing with a lot of mechanical complexity per GB until you get to an absolutely enormous amount of data. I genuinely can't imagine breaking even with a robotic library you can't climb into.
To me, 1PB is also where I'd draw that line. Which I would interpret as it never really being worth it to go to local drives for these storage modalities: you start on cloud storage, then move to local tapes once you're big enough.
(Heck, AFAIK the origin storage for Netflix is still S3. Possibly not because it's the lowest-OpEx option, though, but rather because their video rendering pipeline is itself on AWS, so that's just where the data naturally ends up at the end of that pipeline — and it'd cost more to ship it all elsewhere than to just serve it from where it is. They do have their self-hosted CDN cache nodes to reduce those serving costs, though.)
Particularly relevant for personal data storage.
(Offlined SSDs would probably be fine, if those ever became competitively affordable per GB. And https://en.wikipedia.org/wiki/Disk_pack s would work, too, given that they're just the [stable] platters, not the [unstable] mechanism; they would work, if anyone still made these, and if you could still get [or repair] a mechanism to feed them into, come restore time. For archival purposes, these were basically outmoded by LTO tape, as for those, "the mechanism" is at least standardized and you can likely find a working one to read your old tape decades later.)
Even LTO tape is kind of scary to "leave on a shelf" for decades, though, if that shelf isn't itself in some kind of lead-lined bunker, given that stray EM can gradually demagnetize it. If you're keeping your tapes in an attic — or in a basement in an area with radon — then you'd better have encoded the files on there as a parity set!
I think, right now, the long-term archival media of choice is optical, e.g. https://www.verbatim.com/subcat/optical-media/m-disc/. All you need to really guarantee that that'll survive 50 years, is a cool, dry warehouse that won't ever get flooded or burnt down or bombed [or go out of business!] — something like https://www.deepstore.com/.
But if you're dealing with personal data rather than giant gobs of commercial data, and you really want your photo album to survive the next 50 years, then honestly the only cost-efficient archival strategy right now is to keep it onlined, e.g. on a NAS running in RAID5. That way, as disks in the system inevitably begin to die or suffer readback checksum failures, monitoring in the system can alert you of that, and you can reactively replace the "rotting" parts of the physical substrate, while the digital data itself remains intact. (Companies with LTO tape libraries do the same by having a couple redundant copies of each tape; having their system periodically online tapes to checksum them; and if any fail, they erase and overwrite the bad-checksum tape from a good-checksum copy of the same data — as the tape itself hasn't gone bad, just the data on it has.)
Paying an object-storage or backup service provider, is just paying someone else to do that same active bitrot-preventative maintenance for you, that you'd otherwise be doing yourself. (And they have the scale to take advantage of shifting canonical-fallback storage to being optical-disk-in-a-cave-somewhere format — which reduces their long-term "coldline" storage costs.)
Instead, you're just left with the need to do the much rarer "active maintenance" of moving between object-storage providers as they "bit-rot" — i.e. go out of business. As there are programs that auto-sync between cloud storage providers, this is IMHO a lot less work. Especially if you're redundantly archiving to multiple services to begin with; then there's no rush to get things copied over when any one service announces its shutdown.
Its also the software to run the blasted thing as well. As soon as you get into tape shit get's enterprise-y real quick. There are opensource tools to manage tape collections, but its not fun.
LTO tape libraries are fairly cheap to pick up second hand, its the cost of getting the newer drives that hurts.
Oh does it? You'll never be forced to update or maintain that configuration due to shifting sands of company policy, infra, or CVEs?
Many drives I did that with lived happy afterlives after being revived, some having been brought back from the dead more than once.
I have archived data (almost a couple hundred TB in double copies) on more than 60 HDDs, about half from WD and half from Seagate, from several HDD generations, i.e. with capacities of 2 TB, 3 TB, 4 TB, 6 TB and 8 TB, most being from the 4 up to 8 TB generations. All HDDs were more expensive WD and Seagate models (i.e. with longer warranty times), not their cheapest consumer HDDs.
All files were checksummed for error detection. When the HDD content was transferred to tapes after some 3 to 6 years since the initial archival, almost all HDDs had a few errors, which sometimes were not reported by the drives even when they were detected by the file checksums.
Only the fact that all files were stored on at least 2 HDDs has prevented data loss, because even when both HDDs had errors they were in different locations.
A few years later I got a QIC tape for my 486 PC and had similar experiences. There was even the time I tried to restore a several kb configuration file from a Tivoli tape robot and was quoted that it would take 18 hours and figured I could recreate it in much less time.
If I did it all the time I'm sure I would do better, but 2 out of 3 times or so I had tape backups fail on me, contrast that to almost always being able to read HDDs, even if there is some bit rot.
Would using something like ZFS done effectively the same thing? Not being critical. I genuinely want to know more.
* https://www.jedec.org/sites/default/files/Alvin_Cox%20[Compa...
I'm sure most last longer, but as a CYA I wouldn't necessarily want to rely on that assumption.
They're not used as a backup, rather I would start with the data disk and as the price goes down and my needs grow the drive would be replaced with the higher capacity one. And there is an extra set of drives (copy of those) I always back them up.
The biggest expense in TAPE storage has always been the tape-changing-robot. (and if you ever worked with SUN HW -- the software was super expensive and the UI/UX of the backup software (logato? I cant recall the name -- was a PoS (not point of sale) and I hated it...
However, with that said, the price/GB/TB for tape is very cheap...
What would be a good idea would be to use tape duplication robots to take one tape to N tapes on a fairly regulated basis (like every 2,3,5 years) where you read and copy to new such that the medium (the cheap tape) is cycled through...
For a home/small business it's still too much effort and expensive.
I tried revisiting this topic so often and HDD was always so much cheaper for just 5-30tb which is quite a lot for quite a lot of companies.
They are, but If I'm not mistaken you need a SAS controller, and those go for ~1k euros in my area, even used.
Fiber channel drives also exist.
My archive is getting to the size that optical discs are becoming cumbersome. Do you know of a tape solution that works with macOS?
Any tips? Particularly on finding something I can interface with SATA?
Maybe 40 Mbytes of data. It's archived with the write-ring removed.
Somehow, I don't think I'll try to rebuild the scattering matrix for Jovian cloud particles...
Mainframes in 1968 were handling much bigger data sets than you could handle with a micro until 1990 or so.
For a laptop SSD, all of these matter. You can't do much compression because compression consumes a lot of power and latency upon access, some compression schemes make random access harder (i.e. compression schemes without an index where you have to scan through intermediate checkpoints or, in the worst case, sequentially through the entire media), and it may lead to write amplification as well if you need to add a piece of data in the middle of a file (basically the issue with shingle HDDs). As a result, no compression possible, and so SSDs are advertised with the raw capacity (or, in fact, they are underadvertised because SSDs need spare block capacity to account for wear).
There's no such thing as a standard distribution of compressible data. Customer data varies wildly.
Even things like Windows Bitlocker and LUKS/dmcrypt on Linux totally ignore the drives ability to do any encryption, and do all encryption using the CPU before the drive sees the data.
Bitlocker/LUKS could easily just calculate the encryption keys and send them to the drive and trust it to encrypt/decrypt data for them... but they don't.
Source? Last I heard many companies rely on the encryption supplied by tape drive systems.
The reason is that all encryption is based on the separation of place between the encrypted data and the encryption key. Whoever can access the encrypted data must not have any way to access the encryption key.
Whenever you give your encryption key to a hardware device, or worse, when the hardware device also generates itself the encryption key, it becomes impossible to ensure that the attacker will not be able to access the encryption key.
It is impossible to know how the encryption keys are stored inside a hardware device and how and when they are erased and how easy or how difficult it will be in the future for an attacker to retrieve them.
It is impossible to believe any marketing claim of the vendor of a hardware encryption device about how tamper-resistant the device is, because such claims have very frequently been proven to be lies (even when the claims come from the largest companies, e.g. Microsoft and many others like it) and it is too difficult to distinguish truth from lies in such cases.
The only reliable means of encryption are in software, under complete end user control (or equivalently, in a custom FPGA).
That's probably one of the items on the list of reasons there are so few vendors in the space, LTO drives are essentially fungible and don't do anything interesting across vendors.
In the end, it’s encrypted, compressed, decrypted, and decompressed multiple times on the way from the platter (or flash) to the tape recording head. It feels kind of dumb, really, but I’ll agree that if the computer doesn’t have the private key of the drive, stealing the data on tape will be a lot harder.
If they didn't, they'd be vulnerable: https://en.wikipedia.org/wiki/Distinguishing_attack
However, something close to your point is true. If they compress the data and then encrypt, the ciphertext will indeed be smaller than the original data.
If you compress the symmetrically encrypted data, you will indeed gain some capacity. Not as much as you would compressing the raw data, but a visible amount, because the symmetric algorithm doesn't have the property of indistinguishability.
That property is not needed for tape storage because anyone who has the tape can safely assume it contains a ciphertext. The tape has header information saying whether it's encrypted or not.
If you read the links you posted, you will find out that it's not "most" cryptosystems, it's some. Some applications just don't need indistinguishability and LTO Tapes are one of them.
All that said, IBM does compress first then encrypt to make the most of the data compression, but the resulting ciphertext isn't completely random.
You can't really say indistinguishability makes one cryptosystem better than another because it prevents distinguishing attacks because for some applications it's both not needed and computationally expensive.
The tape sort algorithm read tapes backwards to avoid the delay of rewinding. Pretty cool.
If you put a ping pong ball atop the blast from the vacuum pumps, it would hover magically in the air.
If a tape got worn out from use, you cut off the first 100' and pasted a new silver Load Point Marker.
As a hobby, I've looked into second-hand LTO tape backup for my 71 TB NAS but the drive alone is pricey, and if it dies, how can I restore my data without forking over too much money?
In the unlikely event your NAS dies at the same time as your LTO tape drive, then you buy a replacement drive, restore the tapes, and sell the drive again on eBay, recouping most of your costs.
https://aws.amazon.com/snowmobile/
Or you can get a Snowball that holds 80TB. These you just ship via UPS.
AWS charges 5 cents/GB at the highest tier. That’s $50/TiB (roughly). 18 TiB out from AWS costs $900. You can buy a nice WD Red Pro 18 TiB drive from Amazon for $298, and shipping it next day is not particularly expensive.
Of course, you can’t load data onto that drive from AWS, and you won’t pay anywhere near $900 to send 18 TiB over a network from any reasonable provider.
https://www.cnet.com/tech/computing/carrier-pigeon-faster-th...
(You can occasionally get away with slightly lower feed speed -- the drive will write "invisible empty data" to keep the motors running and tape movement going -- but that cuts into capacity of the media in unpredictable ways. And eventually the drive decides it is facing data starvation and stops writing. The result of that is a few seconds of tape repositioning to restart the writes.)
Even my now decommissioned 8-drive NAS could do 800MB/s.
And my first 20-drive NAS (long since gone) from 15+ years ago could keep up. (Peak at 1GB/s)
We used to run "Dell" storage and tape (rebranded ADIC changer and Quantum/IBM drives). A tape changer would autoload tapes under command of the backup software. It would run overnight, automatically changing the tapes two or three times, as required, without issues.
Restore was another problem, mostly the fault of the software.
Used 10G nics are pretty inexpensive these days, if you really want to write network to tape.
Then you'll need the tapes (can be expensive, buy in bulk)
and then spend the time doing the software. You'll also want to make sure you can stream data in reliably at >300megs a second
after that you'll be good. then its a case of putting the spare tapes somewhere safe.
Or does it rotate the tapes?
There are single hyperscalars that buy more than this much capacity in HDD in a single quarter.
The conventional wisdom is that the tape industry continuously backports hard disk media innovations that are a few years old, so the end game is that there will be a period after HD innovation ends that allows end-game tape to create an insurmountable advantage over the end-game HDDs.
Put another way, if HDDs stop getting bigger this year, then they'll end up competing with tape drives that continue improving density at current trajectories until ~ 2028, and choosing between 2023 HDDs vs. 2028 tape for archival storage will be a no-brainer.
Having said that, the rapid drop off in tape unit shipments combined with the stagnation of total shipped capacity suggests conventional wisdom might be wrong.
I wonder if HDD density improvements have already stopped, so now we're seeing the endgame.
That's why, while YoY growth of tape is (checks blog...) 5%, that's lower that global HDD sales growth even this year. This is a marketing piece from the LTO consortium.
Lastly, HDDs do way way better in the DC, operationally. From a much wider environment/humidity envelope of operation, to IO scaling linerally across data corpus, to siteops management, to having more than literally 1 vendor, etc.. That's why there are only two hyperscalars left deploying new tape, and even there it's dwindling or being deprecated.
Now HDDs have their own problems, and their density scaling with hamr/mamr/smr hasn't gone great, but no one is really asking for less IO from devices given power cycling drives tends to reduce IO just as effectively as fewer tape drives, and for far less money.
What's in a LTO drive that costs so much to make, and why can't it cost the same as a hard disk? When I think of it a hard disk seems rather more complicated.
And the guts of one don't look that much more complicated than a VCR. Does the head cost $3000 or something?
I mean I found that nearly instantly and there are probably others.
But backups aren’t a “one and done” you want to keep doing them.
The TCO of tape only makes sense at scale, and is much harder to assess than HDDs, as tape cost includes media, drives and libraries, each of which have different life times. Tape libraries can be used across many tape generations.
I wonder if IBM is going to introduce a new generation of 3592, or is LTO the “last man standing” too?
The latest IBM 3592 drive generation (20TB uncompressed), came out in 2018. At the time, their roadmap had future generations promising 30, 40, 80 TB capacity-but I’m thinking, five years on, if they were going to deliver any of that, surely they would have by now?
NRE: non-recurring engineering expenses. Every product needs to be researched and designed before manufacturing. Take all of those costs and divide by the number sold (which in the case of tape drives is very low).
Whatever happened to that area of research? Did SSDs and flash drives kill off that idea?
That's how modern SSDs achieve so much density. Alas, the machinery to etch silicon chips in this way remains very expensive, so I bet that Tape and Hard Drives remain the price-per-TB king for the foreseeable future. I'm pretty sure SSDs have won the density crown however.
------------
IIRC, BluRays achieved 4x layers through focusing the laser, reaching 100GB per disk back when BluRay research was still popular (15 years ago or whatever). I don't think it ever was a popular format, but it existed as a crude 3d layer.
I guess Hard Drives are like 8x platters or more, achieving a degree of 3D as well.
But I presume you're talking more about technologies that achieved hundreds of layers (Tape, thanks to "rolling up", or SSDs/modern Flash with 200+ layers, etc. etc.)
-Odo
Rest in peace René Auberjonois.
Unfortunately, LTO is not truly backwards-compatible, not even for reading. https://en.wikipedia.org/wiki/Linear_Tape-Open#Compatibility
That being said though, tape generations advance infrequently enough that this would never really be a problem. And if it is a problem, the secondary market is absolutely loaded with older gen drives going all the way back to LTO 1 if you really need it.
Re-branded Tivoli Storage Manager in 1999, I think still maintained, alive and kicking. https://en.wikipedia.org/wiki/IBM_Tivoli_Storage_Manager
Sort of like how Oracle is named, and I'm sure others.
Now that number is real.
Do people expect that to always be the case, or is anybody expecting tape to have been supplanted by something else (optical or other) 20 or 30 years from now?
Or is there just something about the physics of tape technology that we can't even imagine anything else in the medium-future that could even begin to compete with it for archival needs?
(Knowing nothing about the area, just as a consumer I would have expected that with the advent of 2-layer DVD's and then 4-layer Blu-Rays, that before too long we'd be getting 1,000-layer discs and essentially moving into 3D storage cylinders by now or something that would have left tape in the dust...)
The vast majority of even backup data is legally required to either be readable in a very short amount of time (say one week) or is required to be deleted within 90 days. Neither of these are ideal for optical (or tape btw, thus why it's YoY growth is only 5% in this LTO marketing piece)
Wouldn't an autoloader solve the the former, and a fire solve the latter?
In reality, the data is all mixed up because data gets backed up usually around the time it's written, and gets deleted in a completely different pattern. The art of compaction is very important, and things like optical and tape make inaccurate compaction extremely expensive due to the reduced IO.
And even if we pretend you do have perfect compaction, you still don't have nearly the IO that HDD would provide. Considering hypothetically you have perfect compaction, that also means you have the perfectly smallest live data set, and thus the HDD premium is even smaller for wildly better throughput in disaster.
A common usecase for tapes seems to be to store a big archive of data. Imagine having some API to my archive, which says for each record/file stored whether I want to keep it stored, read it, or discard it.
Over time, many tapes start to be only half full of data I want stored (the rest has been discarded). It makes sense to read all the important stored data, consolidate it, and rewrite it to a new tape, perhaps with some new data.
For this usecase, you only need half as many tape drives if a drive has the ability to read data while rewriting a new data stream.
I suspect your use case is strictly for backups, whereas my use case is more archival (where all copies of the data are on tape, and there are no copies on hard drives).
These tape libraries can also defragment tapes. For this all the data that needs to preserved is read from the tape and written to another tape.
Even for backups you have a consolidation/defragmentation task running every so often that frees up space on the tapes and only keeps the backups you specified in your retention policy.
There have been people who put a filesystem on a tape with sectors - they could then do this, but since tapes are always sequential access it isn't practical, just an interesting hack to show off.
[1]: https://en.wikipedia.org/wiki/Log-structured_file_system
And there is no avoiding reading all the data on the tape when rewriting to another tape. The media is cheap after all. And therefore "discard" is largely pointless, unless it's 90% or more of the data.
I doubt anyone makes anything like that though?
Wikipedia says that a C90 cassette tape was 135 meters.
Now ignoring all the other differences (width, serpentine winding, chemistry, etc. etc.), that gives us roughly ~1.6 TB of data if a C90 (90-minute) cassette tape had the rough bits-per-inch of a modern LTO-8 tape.
Oh the joys of tape barcoding and offsite tape vaulting.
The lock on the office data safe, a large FireKing, was an interesting disk lock.
It'd take a lot to make me give up my tapes at this point.
However it feels irrelevant for a company who exclusively operates in the cloud.
Tapes can burn. Give me multi-location online/nearline disk storage any day.
We had nearlines that were basically archive targets, that stored about 6months to a years worth of versioned data. We also dumped to tape at the same time (one for data interchange with other companies, and one for our actual backup)
WE arse saved multiple times from both tape fuckery (mostly the robot being sad) and nearlines going whoopsey (they had reduced redundancy, and a whole bunch more nodes to allow for the size)
Given these are intended for enterprise usage, there is an expectation of quality, service, and support associated that you ARE paying for with that cost.
I thought it was only IBM still making drives today
I'm reaching the point I'm starting to consider tape backups in addition to my regular systems.