Seagate Announces 12 TB HDD: 2nd-Gen Helium-Filled HDD
anandtech.com
anandtech.com
(fortunately, it'd only take 3-4 months to bittorrent down all those, ummm, linux ISOs...)
Oh, I would, but I guess you meant per byte.
At current rates it's $3K monthly to keep a backup of this HDD on tarsnap and that's not including bandwidth. That's six 10TB drives.
I don't mean to discourage from tarsnap though. Cloud storage is unfortunately still pretty expensive, and most people are probably only truly paranoid about much less than 1GB of their data.
I'd consider a backup solution that takes 3+years to make a full backup (or 4 months for a restore) to be of negative value to me...
(I wonder how many months I'd get away with uploading 300GB/month before my "unlimited" ADSL plan hit me with it's "acceptable use" disclaimers? Given that I'd be using 100% of my upload bandwidth to take this 3 year backup, I suspect the cost of 2 identical drives to make local backups would probably be cheaper than the network costs of sending it to tarsnap!)
I've got terabytes backed up, no issues at all.
I can upload several hundred gigabytes per day without any throttling or other issues, so I suspect your issues are with your ISP and not Amazon.
sounded good, and got depressed from the Australia Tax. US$59, but AU$100.
AU$100 - $10 = $90 (ish) as taxes are built into our prices, not added later
90 Australian play dollars = 69 American real dollars
US$69 - $59 US price = Australia Tax of ~US$10, approx 16%.
In your case, even with the exchange rate as it is, price parity would be ~$80. If the taxes are ~$10 as you suggested, that means you're paying another what? $10 for essentially the same service? In a way, I'm surprised it's "only" that much more, but it's still terrible for Australian consumers.
edit: I wonder if Amazon goal with this was to get their hands on as many photos with their full metadata as possible (for training).
It's concerning to hear, since I know people using it for backups.
In my case it was 90GB, a whole single machine. The fact that they were pretty causal about it, is worrying. But I still use them for some use cases because they have unlimited storage with software for Linux, although I don't depend on them.
Quote from the support ticket:
"I have looked again for your data, but we are unable to locate the archive anywhere in our system.
I am sincerely sorry that we have let you down as there is not a reason that I can find that this data should not be here.
If there is anything that I can to for you, please do not hesitate to ask and I will do so."
Will report back.
I downloaded a full restore of the state of my Macbook Air as of roughly one month ago for testing. I don't have the storage space for a total restore (which would be tens of TBs I guess, since it's an incremental backup solution [unless you could restore all the files instead of all the backups]), but I've tested everything I care about.
I actually just deleted most of the backups from Amazon (just to be nice), since I don't really need to have backups of everything that has ever touched my Downloads folders or web browser cache for example.
I hate to pick nits... But my email alone is about 8 GB for the past 25 years.
That doesn't include documents, photos, etc...
I think we underestimate what Joe Average cares about.
On the other hand, 1GB is pretty tight. I suspect there's a lot of people who could come up with 1GB of photos they don't care to part with, and coming up with 1GB of valuable home videos is child's play, even assuming efficient encodings of the video rather than the rather flabby stuff that tends to come straight out of consumer cameras or phones.
But I bet if most people had to fit the Family Jewels into 10GB, they could, just one more order of magnitude away. Of course HN will have a way-above-average number of people who have way-above-average video recordings and such. Or, perhaps more realistically, call it a single 32GB SD card. At the moment my "family jewels" clocks in at about 90GB, however, that is with no judgment whatsoever about what photos we keep and with a lot of the aforementioned "flabby" videos, not to mention, no judgment whatsoever about what videos we're keeping either.
It mostly comes down to how many photos you take, really.
When I worked for Nationwide insurance (50 to 60 thousand people) we could keep as much email as we wanted locally, but no more than 100 mb on the server. This is a stupid policy, but it is a common policy.
I would like to use borg backup, but it isn't as flexible as arq.
Personally I have close to 7TB of backups and personal video footage stored, everything encrypted, never had any issues.
However there's always the danger of them pulling a onedrive at some point in the future and reverting from unlimited back to something like 1TB if people actually use it.
I always wished a 5.25 mass storage HDD would come back on the market. How much would those hold relative to this?
Hm, I'm guessing about 2.5x more data per platter, about 50% more platters... 45GB?
Best case is 12TB / 254 MB/s ~12 hours, * 2.5 = over a day.
However, random reads are a lot slower.
Very few things do, really. It's pretty much just rebuilding a RAID or reorganizing your data storage hardware that care about full-disk transfer speed, and in those cases two days isn't a big deal.
The only 'obvious' improvement you can make is having more arms or figuring out some way of micro-aligning multiple heads at the same time on a single arm. There's not much benefit from larger single units.
You can already fit way too much SSD into a 3.5 inch drive, and nobody will ever buy it because they want more ports and performance per dollar. There's no benefit for either technology to go up to 5.25 inches.
If that's to pricy, you can still fit 16TB in a 5.25" drive bay with the same dock and 8 2TB rotating drives[3] and 20TB with this dock[4] and the 5TB version of the same drive[5]. Note that all 4 and 5TB rotating drives I can find are 15mm thick, so the 8x drive-bay is a no-go.
1: http://www.icydock.com/goods.php?id=192
2: https://www.newegg.com/Product/Product.aspx?Item=N82E1682014...
3: https://www.newegg.com/Product/Product.aspx?Item=N82E1682217...
4: http://www.icydock.com/goods.php?id=184
5: https://www.newegg.com/Product/Product.aspx?item=N82E1682217...
Google presented a paper at FAST16 about the possibility of fundamentally redesigning hard drives to specifically target hard drives that are exclusively operated as part of a very large collection of disks (where individual errors are not such a big deal as in other applications) – in order to even further reduce $/GB and increase IOPS: https://static.googleusercontent.com/media/research.google.c... .
Possible changes mentioned in the paper do actually include new (non-backwards-compatible) physical form factor[s] in order to freely change the dimensions of the heads and platters. The only market for spinning rust in a decade or so will be data centres (or anywhere else that needs to store a shitton of data) -- everything else will be flash.
Other changes mentioned in the paper include:
* adding another actuator arm / voice coil with its own set of heads
* accepting higher error rates and “flexible” (this is a euphemism for “degrades over time”) capacity in exchange for higher areal density, lower cost, and better latencies
* exposing more lower-level details of the spinning rust to the host, such as host-managed retries and exposing APIs that let the host control when the drive schedules its internal management tasks
* better profiling data (time spent seeking, time spent waiting for the disk to spin, time spent reading and processing data) for reads/writes
* Caching improvements, such as ability to mark data as not to be cached (for streaming reads) or using PCIe to use the host’s memory for more cache
* Read-ahead or read-behind once the head is settled costs nothing (there’s no seek involved!). If the host could annotate its read commands with its optional desires for nearby blocks, the hard drive could do some free read-ahead (if it was possible without delaying other queued commands).
* better management of queuing – there’s a lot more detail on page 15 of that PDF about queuing/prioritisation/reordering, including the need for the drive’s command scheduler to be hard real-time and be aware of the current positioning of the heads and of the media. Fun stuff! I sorta wish I could be involved in making this sort of thing happen.
tl;dr there is a lot of room for improvement if you’re willing to throw tradition to the wind and focus on the single application (very large scale bulk storage) where spinning rust won’t get killed off by flash in a decade.
You might have a few drives with bad seals that fail early. But say the majority are sealed properly, but the helium still leaks out at x small rate over time, you might have almost all the drives suddenly fail at around the same time after say 6 years or whatever.
I don't think you can really know if they're truly reliable in that sense until you can take one that's been sitting one in a drawer for 10 years and plug it in.
A high-altitude environment might put more pressure on the seals than other places, that could shorten the life-span, but other than that it shouldn't be a huge deal. Most drives have a commercial life-span of no more than 4-6 years anyway. After that you're on borrowed time.
Helium also goes through whatever the sealing gasket is made from, at a much higher rate. http://lpc1.clpccd.cc.ca.us/lpc/tswain/permeation.pdf is a neat chart, probably the gasket is buna-n because it's cheap. I think that chart and https://en.wikipedia.org/wiki/Permeation are enough to get an order of magnitude answer to OP's question, but I'm not quite up to doing the math.
Serious vacuum systems have to worry about this sort of thing more than one might expect. I once had a project where two test engineers and myself spent around a week hunting for leaks in a vacuum system that turned out to be caused by my specifying Silicone o-rings instead of Viton.
No. Helium hard drives are laser-welded shut. See http://www.seagate.com/files/www-content/product-content/ent... for the details about what parts are made out of what and the basic details of fabrication.
I'm not sure why I thought there was a rubber seal, I have vague memories of seeing an exploded view with one somewhere.
http://www.engineeringtoolbox.com/thermal-conductivity-d_429...
Big difference between different gases.
But, check this out. Helium also has a high heat capacity. Table on same website:
http://www.engineeringtoolbox.com/specific-heat-capacity-gas...
Six times the conductivity of nitrogen (0.142 to 0.024) and five times the heat capacity: 5.19 to 1.04.
Big conductivity and capacity should translate to "convective cooling monster gas". :)
The servo and associated electronics surely generate many orders of magintude more heat.
http://www.hgst.com/company/media-room/press-releases/HGST-H...
As I understand it, the crucial factor motivating helium-filled drives is the reduction of vibration due to internal turbulence. Reducing turbulence inside the drives allows platters and heads to be made thinner and packed more closely together without affecting reliability.
Fun-fact aside: The head fly height in helium drives is ~1 nm.
http://www.engineersedge.com/heat_transfer/thermal-conductiv...
Back in the stone age disks were "fast" in random access but they couldn't hold as much as tape could. So to get the storage you had a mix of disk and tape. That started as you had a command you would send to the operator (another historical concept) which was the person who was responsible for 'operating' the computer. That would print a message on a hard copy terminal that would say "Load Tape XYZ on Drive 2" or something similar, the Operator would go to the shelf of tapes, pull out the one that was labeled XYZ and put it on drive 2, "mount it" (which would, on vacuum readers, suck in and tension the tape and read the first block (the label)) and make it available to the operator who could verify it was XYZ and then send pack (again on the console) "tape loaded." Then your batch program would start and you'd get the classic video of a computer with the tape being read in bits and bursts and often another tape being written in bits and bursts.
Computers got bigger (able to process more data) and tape libraries got bigger, one of Sun's big customers was Fingerhut (mail order catalog) in Minnesota and they had a room where there were lots of 'tape operators' and when a customer was on the phone and the operator said "let me bring up your account" it lit a sign in the tape room with the needed tape, someone would jump up and grab it and put it on the nearest tape drive and 'tag' (there was a clock showing time from request to mount) and the tape would identify itself and send the customer record to a disk so that they were "online" and the operator's screen would light up with all the customer details.
IBM and StorageTek made robotic libraries that did the same thing, but without the human in the loop.
Then in the early 2000's NetApp and other storage vendors started offering ATA disk drives (dense, cheap) as storage offerings and slowly, eventually the tape libraries were crushed because this dense storage was more cost effective.
Primary disk storage has now gone over to solid state disks. They have even better random access and generally and do reads at the limits of the interface they are attached to.
But there is still a market for dense, cheap, read mostly data stores. That which used to be tapes, then tape libraries, and now spinning rust in a bath of helium. A typical SATA drive can really only do about 110 "IOPS" per second, in a large part because the mechanics of moving things around is inviolate.
So the use for dense data stores behind a small i/o pipe is read mostly archival and reference data. Oil & Gas sonar dumps, credit card sales transactions, backups of data which is 'live' elsewhere, etc. The role that Tape used to play but no longer does very well.
You really can't have too much local storage. Price is the only prohibitive factor, though rebuild times are starting to be too.
Another factor to consider is that uplink, if you wanted to store stuff in the cloud, is almost always terribly slow and often capped, plus the cost of storing multiple TBs on any given provider.
the use cases for raid5 have pretty much been replaced with raid6 plus hotspare.
http://www.zdnet.com/article/why-raid-6-stops-working-in-201...
The wise sysadmin mantra of "RAID is not a backup" must be kept in mind at all times.
Unless 3 drive controllers die simultaneously, you can survive finding faults during a full rebuild.
(Though probably something with checksums instead of naked redundancy.)
Why is that happening? There's no reason to perform a bunch of writes to those drives.
It would be interesting if someone offered a similar plan, but priced on energy consumed instead, giving incentive for managed storage arrays that power down drives that idle (for some algorithmically-derived definition of "idle", trading off against projected lifespan of the device due to increased power cycling), and other energy-saving practices. Slices up the granularity of cloud hosting even finer.
One example from a previous job, these in a Ceph cluster as an origin for a CDN delivering video. We transcode the video to various bit rates and chunk em, then clients request them through CDN, there is not a lot of writing, with some reads when the asset is new, but once it is cached it just basically sits and is idle.
Larger capacity meant we could store more data in a single rack, our limitation for the origin wasn't CPU or bandwidth throughput, it was literally our storage that forced us to keep expanding.
The deployment I did was using 8TB drives. We used HGST with their Helium because of the lesser power draw, allowing us to install more servers in a rack without going over our maximum power draw and heat density for cooling.
So right now I have 3 copies of everything.
1. Masters are on 4 bay FreeNAS mini (ZFS baby!) 2. Rsync to Thunderbolt drive on desktop (well laptop but...TB monitor with drive attached). 3. Everything goes up to Backblaze auto-magicly.
But if you aren't (or for others who might be reading this), that's a crazy price.
In case noise and idle power consumption aren't as big of an issue:
You can get a 12-bay Supermicro with two six-core CPUs + 48GB and 12 HDD bays for <$400 shipped: http://www.ebay.com/itm/2U-Supermicro-12-Bay-826TQ-R800-Serv...
Or the very popular SA120 DAS for <$300 to hang 12 bays off another machine (Dell R210II is quiet and cheap): http://www.ebay.com/itm/Lenovo-70F10000UX-THINKSERVER-SA120-...
(Random eBay links I found just now, have no affiliation with either of them)
Yes. It's what Amazon Glacier runs on: https://en.wikipedia.org/wiki/Amazon_Glacier#Storage
The Backblaze blog is also a great place to learn about a large storage deployment: https://www.backblaze.com/blog/hard-drive-benchmark-stats-20...
If all you need are two hard disks, go for it.
Since then the price went up by 20% for reasons that are beyond me.
Edit: apparently it's not easy even to take the lid off of these drives :) See https://youtu.be/ANMtvYnI1gQ
https://www.forbes.com/sites/timworstall/2015/06/18/were-rea...
https://www.wired.com/2016/06/dire-helium-shortage-vastly-in...
> In 2014, the US Department of Interior estimated that there are 1,169 billion cubic feet of helium reserves left on Earth. That’s enough for about 117 more years.
That's how markets work.
We will need something because fossil fuels are not sustainable. In 100 years we should be pumping much less oil than we do today. This also means finding substitutes for all of the byproducts of oil production.
With DM-SMR strange noises and paranormal drive activity has to be expected since the drive can spend considerable time flushing it's PMR buffer (~16 GiB) to the shingled zones.
[1] I divided the rated workload listed on the page by the size of the disk.