“Let's talk about a hypothetical public-facing service”
reddit.com
reddit.com
The future is now! Instead of you sending money to Amazon and them sending you disks, they keep the disks and you send them the money anyway.
Progress! :-D
Amazon's even more expensive. Storing one petabyte on Amazon, at the slowest, cheapest, glacier level, will cost you $10000/month, so at about 10 months, you've paid a hundred grand. At thirty six months, you've paid three hundred sixty grand. Plus, you have no hardware to show for it, while amazon sells that broken outdated junk to a reseller to regain some of their investment. Had you hosted it yourself, you could have sold all that equipment off to refurbishers and resellers and regained at least a portion of your investment back. So yeah, it is pretty terrible to choose Amazon.
n+1 means you need two of them in case one gets the flu.
Tell me again about how Amazon's a ripoff, please.
So we're talking about $75,000-90,000 a year in salary to maintaining the cluster if you want 24x365 coverage (which Amazon provides), as you'll need 2 people per shift, 3 shifts per day at a minimum to have people in house, even if they're only spending 10% of their time actually working on that particular issue. In reality, these are unrealistically small numbers. Each employee will have a cost to the corporation of $125,000-150,000 a year. Employees of that caliber spending 10% of their time on supervising a cluster for a year is $75,000-90,000. I'm amortizing the amount of work over your 3 datacenters by imagining you already have the staff and only counting out the number of hours needed for just this work.
So the reality is that you left off something like $225,000-270,000 in actual cost of running your own cluster from your analysis, because while it's not "much different", it is a few hours a week from at least 6 employees if you're really talking about running managed, highly reliable storage.
When you have your own servers and staff, even if they are just on-call with a pager, you know they are going to work for you on your problem until it is fixed.
In comparison, the sentiment about AWS is this: better build redundancy into the application because at the server layer, you get whatever you get.
My argument would be that part of the system is always down, the question is which part and how long, and how that impacts the system performance on the whole.
AWS services work well if you build a stateless system which is in some senses "embarrassingly parallelizable", because you can talk about the capacity of such a system, and the impact of non-functioning components is easy to predict. This is how most standard engineering is done, across disciplines.
Traditionally, it has not been the case in computer systems, but most modern techniques advocate using such systems, because they're MUCH more reliable.
You just sound like you need a safety blanket for emotional reasons, not that you're making sound engineering points about how to most cheaply engineer a high availability system.
I mean, do you really believe AWS engineers aren't working hard to keep their system fully functional?
It's like forgetting to count the price of hardware, and only talking about the relevant cost of electricity.
Ed: Accidentally a negative.
This is something that has been puzzling me. Many years ago I purchased 4x 2TB 5900 RPM drives for a 4 bay ReadyNAS (cost about ~$300 for drives plus ReadyNAS). They have been spinning nonstop for ~4 years [1] and haven't had to replace a single one. Not even an increase in errors to signal that the drive is going.
Yet - I've worked on a SAN that cost hundreds of thousands of dollars and would have to replace a disk about every month.
Granted the disks in SANs probably spin faster (thus faster data access/lower MTBF) - but that high failure rate seems rather suspicious to me.
In the Blekko cluster we have just under 10,000 drives. We have a two 20 drive 'boxes' (40 drives) from Western Digital, as drives fail we pull replacements from the 'new/refurbished' box, and we put the dead one in the outgoing box. When we get up to 20 we RMA them in bulk, 20 go out, 20 more come in. That becomes the new 'new/refurbished' box.
It really isn't SAN vs non-SAN it is all statistics.
That said, if you're running your ReadyNAS with raid 10 (mirrored drives in a RAID 0 config) you may find some unpleasantness when a drive does fail. Statistically you have a 1/10 chance of not being able to re-silver the mirror for a 5900 RPM desktop SATA drive. That gets a bit painful.
"In addition to presenting failure statistics, we analyze the correlation between failures and several parameters generally believed to impact longevity."
There is also a more recent open source dataset from Backblaze[2] that includes:
"Every day, the software that runs the Backblaze data center takes a snapshot of the state of every drive in the data center, including the drive’s serial number, model number, and all of its SMART data"
which forms the basis of an article correlating SMART data with drive failures at Backblaze[3].
[1]: http://static.googleusercontent.com/media/research.google.co...
[2]: https://www.backblaze.com/blog/hard-drive-data-feb2015/
Which is why for important data I always use some sort of RAID (or cloud syncing). If I lose a drive I won't lose all my data (presumably though if I bought both drives at the same time there is a chance that both could fail at the same or close to the same time).
> That said, if you're running your ReadyNAS with raid 10 (mirrored drives in a RAID 0 config)
ReadyNAS and other products use a special type of RAID that is actually kind of clever (that will use all space regardless of disk size). I know many would criticize the special RAID but I've had more issues with Linux software RAID than I have had with ReadyNAS (nothing related to data loss). But suffice it to say if 1 drive fails then in theory the data should still be ok (at least that is what their claim to fame is).
My personal opinion - I feel like the ReadyNAS will die before the drives. I'm not saying they couldn't or won't die - but I feel like that would be highly unlikely and even more unlikely for more than 1 to fail at once.
Maybe I'm just paranoid - but when I worked on my last SAN it felt like the company used the cheapest possible drives with a short MTBF for the simple reason that my employer would have to keep using them and keep paying for their warranty service (this wasn't a big name like Dell).
Just a friendly reminder that RAID != backup. There are numerous data loss cases that RAID does not deal with.
Personally I use striped ZFS with important volumes periodically snapshotted, replicated to external (and encrypted) disks and then stored offsite (cycle through a couple of sets of disks). Most important data is also periodically synced to cloud storage (as well as offsite disk).
This accounts for:
- Single disk failure (striping)
- Bitrot (ZFS scrubbing can reveal bitrot on disk and correct it from parity)
- Human error (snapshots)
- Catastrophic damage to home NAS (offsite backups)
RAID alone (depending on the particular implementation) will generally not deal with the 3 latter failure cases.
I knew EMC storage was utter shit when, upon attempting to create a new RAID group, I realized that the configuration tool's default was to stripe across drives within a shelf, not to create stripes that span shelves.
Worse, to create the more fault-tolerant, shelf-spanning RAID volumes, one must manually add drives, one by one to the array, in a process that involves about 44 (slight hyperbole) clicks per disk.
And then there was the fact that the configuration tool was Windows-only.
Yeah, screw those guys.
Having said that, EMC is stupidly overpriced bloat.
It was a 20-something disk RAID 10 [1], arranged so that every mirrored pair of disks spanned different enclosures, in order to mitigate the failure of any one shelf — that is, interleaving mirrors across controllers and shelves, exactly as you suggest I should have done — and further, such that any one shelf failing only affects the mirrors that had disks on that shelf.
EMC's software wanted to allocate the drives from two shelves, with an unequal number of drives per shelf. It was just grabbing the next however many disks, linearly.
So, no, they weren't trying to balance the mirror across enclosures or controllers. They just weren't thinking.
[1] By "RAID 10", I mean "striped mirrors" — that is, create a bunch of mirrors that span shelves and then stripe across them — not "mirrored stripes" which is what you appear to be suggesting, with "it is better to have two disks on the same controller/chassis in raid 0, then raid 1".
A striped mirror is recommended in everything I've ever read on the subject, because it puts the redundancy at the lowest level of the array's geometry.
Using a mirrored stripe, on the other hand, means that when one disk fails, any other disks striped with it, still presumably perfectly functional, can't be used; the controller must instead read from and write to the mirror. If a disk in that mirror subsequently fails, you've lost data — and remember that when striping, the chance of failure is multiplied by the number of disks in the stripe.
EDIT: Footnote.
In a big system I am usually more concerned with avoiding any SPOF (that will cause downtime), then adding redundancy for the most likely failures.
Individual mirrored pairs on split chassis will mitigate against long rebuild times for a single disk failure but it requires an all software architecture as the individual HBA can't do hot spare rebuild. It also reduces the benefit of the HBA cache. So intra chassis, striped mirrors; inter chassis, mirrored stripes. Of course, mirroring across striped mirrors would be better, but the cost to peak write and capacity might rule that out.
Also, yes I believe you, EMC sucks. But they do have some good engineers.
It's his right to make up ridiculous licenses. Like the sisterware license (you can use the software if you send me a pic of your sister if you have one), people shouldn't take it seriously and avoid code licensed like that. That Crockford doesn't get this is either him trolling or being clueless.
In some ways the bundleware served like a SMART warning, causing people to back up and migrate their projects.
I found myself in the odd position of hoping SF would come back up so I could finish what I was doing, when I'd normally welcome the news that they had shut down.
Lots of google searching for what's wrong with Sourceforge revealed nothing (other than complaints about their packaged installers/crapware). Now I have the answer to the question I was really asking.
I just need to wait and see if the developers of the packages I need will somehow migrate to github so I can get the source...
What's happened to the mailing list archives for all these projects, are they just down the toilet now? And of course the source, this is bloody terrible !
[Disclosure: I work for SourceForge]
My condolences.
As usual, one's respect for BigCompany is inversely correlated to one's use of their products.
I'm having the same issue with the number of hardlinks, which, for linux ext4 systems, is limited to 65000.
Offworld, Roy Batty reference.
(continues reading)