The future is now! Instead of you sending money to Amazon and them sending you disks, they keep the disks and you send them the money anyway.
Progress! :-D
The future is now! Instead of you sending money to Amazon and them sending you disks, they keep the disks and you send them the money anyway.
Progress! :-D
Amazon's even more expensive. Storing one petabyte on Amazon, at the slowest, cheapest, glacier level, will cost you $10000/month, so at about 10 months, you've paid a hundred grand. At thirty six months, you've paid three hundred sixty grand. Plus, you have no hardware to show for it, while amazon sells that broken outdated junk to a reseller to regain some of their investment. Had you hosted it yourself, you could have sold all that equipment off to refurbishers and resellers and regained at least a portion of your investment back. So yeah, it is pretty terrible to choose Amazon.
n+1 means you need two of them in case one gets the flu.
Tell me again about how Amazon's a ripoff, please.
So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they're only spending 10% of their time actually working on that particular issue. In reality, these are unrealistically small numbers. Each employee will have a cost to the corporation of $125,000-150,000 a year. Employees of that caliber spending 10% of their time on supervising a cluster for a year is $75,000-90,000. I'm amortizing the amount of work over your 3 datacenters by imagining you already have the staff and only counting out the number of hours needed for just this work.
So the reality is that you left off something like $225,000-270,000 in actual cost of running your own cluster from your analysis, because while it's not "much different", it is a few hours a week from at least 6 employees if you're really talking about running managed, highly reliable storage.
When you have your own servers and staff, even if they are just on-call with a pager, you know they are going to work for you on your problem until it is fixed.
In comparison, the sentiment about AWS is this: better build redundancy into the application because at the server layer, you get whatever you get.
My argument would be that part of the system is always down, the question is which part and how long, and how that impacts the system performance on the whole.
AWS services work well if you build a stateless system which is in some senses "embarrassingly parallelizable", because you can talk about the capacity of such a system, and the impact of non-functioning components is easy to predict. This is how most standard engineering is done, across disciplines.
Traditionally, it has not been the case in computer systems, but most modern techniques advocate using such systems, because they're MUCH more reliable.
You just sound like you need a safety blanket for emotional reasons, not that you're making sound engineering points about how to most cheaply engineer a high availability system.
I mean, do you really believe AWS engineers aren't working hard to keep their system fully functional?
It's like forgetting to count the price of hardware, and only talking about the relevant cost of electricity.
Ed: Accidentally a negative.
This is something that has been puzzling me. Many years ago I purchased 4x 2TB 5900 RPM drives for a 4 bay ReadyNAS (cost about ~$300 for drives plus ReadyNAS). They have been spinning nonstop for ~4 years [1] and haven't had to replace a single one. Not even an increase in errors to signal that the drive is going.
Yet - I've worked on a SAN that cost hundreds of thousands of dollars and would have to replace a disk about every month.
Granted the disks in SANs probably spin faster (thus faster data access/lower MTBF) - but that high failure rate seems rather suspicious to me.
In the Blekko cluster we have just under 10,000 drives. We have a two 20 drive 'boxes' (40 drives) from Western Digital, as drives fail we pull replacements from the 'new/refurbished' box, and we put the dead one in the outgoing box. When we get up to 20 we RMA them in bulk, 20 go out, 20 more come in. That becomes the new 'new/refurbished' box.
It really isn't SAN vs non-SAN it is all statistics.
That said, if you're running your ReadyNAS with raid 10 (mirrored drives in a RAID 0 config) you may find some unpleasantness when a drive does fail. Statistically you have a 1/10 chance of not being able to re-silver the mirror for a 5900 RPM desktop SATA drive. That gets a bit painful.
"In addition to presenting failure statistics, we analyze the correlation between failures and several parameters generally believed to impact longevity."
There is also a more recent open source dataset from Backblaze[2] that includes:
"Every day, the software that runs the Backblaze data center takes a snapshot of the state of every drive in the data center, including the drive’s serial number, model number, and all of its SMART data"
which forms the basis of an article correlating SMART data with drive failures at Backblaze[3].
[1]: http://static.googleusercontent.com/media/research.google.co...
[2]: https://www.backblaze.com/blog/hard-drive-data-feb2015/
Which is why for important data I always use some sort of RAID (or cloud syncing). If I lose a drive I won't lose all my data (presumably though if I bought both drives at the same time there is a chance that both could fail at the same or close to the same time).
> That said, if you're running your ReadyNAS with raid 10 (mirrored drives in a RAID 0 config)
ReadyNAS and other products use a special type of RAID that is actually kind of clever (that will use all space regardless of disk size). I know many would criticize the special RAID but I've had more issues with Linux software RAID than I have had with ReadyNAS (nothing related to data loss). But suffice it to say if 1 drive fails then in theory the data should still be ok (at least that is what their claim to fame is).
My personal opinion - I feel like the ReadyNAS will die before the drives. I'm not saying they couldn't or won't die - but I feel like that would be highly unlikely and even more unlikely for more than 1 to fail at once.
Maybe I'm just paranoid - but when I worked on my last SAN it felt like the company used the cheapest possible drives with a short MTBF for the simple reason that my employer would have to keep using them and keep paying for their warranty service (this wasn't a big name like Dell).
Just a friendly reminder that RAID != backup. There are numerous data loss cases that RAID does not deal with.
Personally I use striped ZFS with important volumes periodically snapshotted, replicated to external (and encrypted) disks and then stored offsite (cycle through a couple of sets of disks). Most important data is also periodically synced to cloud storage (as well as offsite disk).
This accounts for:
- Single disk failure (striping)
- Bitrot (ZFS scrubbing can reveal bitrot on disk and correct it from parity)
- Human error (snapshots)
- Catastrophic damage to home NAS (offsite backups)
RAID alone (depending on the particular implementation) will generally not deal with the 3 latter failure cases.