Why we built a 40TB photo server in-house instead of using S3
tech.oyster.com
tech.oyster.com
I can only imagine how much they will be scared each time they need install updates or reboot THE BOX. They will eventually decide to build identical BOX and mirror their data on daily basis. Then they will notice that mirroring such big volumes of data is wasting tooo much system resources and start evaluating in-house distributed storage solutions, such as OpenStack Swift. Then they will notice it is way too overcomplicated and finally decide migrate their data to Amazon S3.
I'm writing it as a person who walked the same path over the last few years. :-)
1) Single Box
2) Single Location
3) Single 40TB RAID 6 Array on single RAID card with 22 Drives (assuming 24 2TB drives, 2 parity, 2 hot spare = 40TB)
4) Single bonded network link means single switch, no redundancy against switch / network device failure
Honestly this may have been cheap but your getting what your paying for: An unreliable backup solution.
Also: Norco? Really? I wouldn't be trusting anything they produce with your company's critical data. Supermicro isn't much more expensive all things considered.
> An unreliable backup solution.
Nope. RAID is not a backup solution. It provides redundancy so the array can survive an event like a device failure (and so the data survives as a consequence) without significant downtime for repair (with zero downtime if you have hot-swap hardware) but it does not, and is not intended to, protect the data from the huge list of other things that can affect it.
RAID is redundancy for reliability purporses, not backup purposes.
OP isn't saying RAID is a backup solution. RAID is being used as a component within the backup solution (the server).
My knee has been playing up and jerked a little there.
RAID is certainly a valuable part of many a more substantial backup solution.
Wether you are swapping out a replacement for an active drive or swapping out a dead drive with what will be a new hot spare (a drive that was previously the spare now being an active drive in teh array) whether you have hot-swap or not makes a difference. Without, you have to power down the array for a short time which in an environment that needs availability to be as high as possible could be an issue.
PS: Honestly, if their custom solution only add's one 9 of redundancy that may be fine. Their CDN is probably at 98+% up time or so which bumps them into the 99.8+ range which is better than most services out there.
Maybe, but not necessarily. Could be a pair of stacked switches (making one logical switch), with the bonded link having one port on each physical switch.
Excuse me?
As a customer I'd feel slightly uncomfortable about my data by now. You make it sound like stuffing 24 disks into a box is rocket science to you, all the while SuperMicro and others sell plug&play chassis for up to 45 disks[1].
Also, you didn't mention it in the post, but you do have at least two of these in two physically distant racks, right?
[edit: deleted snarky comment about people running windows on a fileserver]
They are backing up their curated photos (which seems to pretty much be the whole business as they say in the article).
It appears they are using Akamai services to server their images so hopefully their is some extra redundancy there already.
Other than that it seems like it would be a steal to use s3 as a backup system at this point because from this article it looks like they need to hire another employee to tell them why this backup solution seems a bit silly (I mean having trouble setting up the hardware).
Being a single big box, it had a bunch of single points of failure, and boy did they fail; we probably had 5-10 hours per month of downtime due to the photo server falling over (flaky RAID controller firmware, mostly).
Also, since the big box was expensive, we only had one in production. There was code for taking a newly-uploaded photo and copying it over to the photo server that only executed in production, which meant the only way to functionally test it was to ship it and hope.
We switched to S3 about a year ago with different buckets for prod, staging, and dev; the production-only code paths went away, and there hasn't been any photo-related downtime since. Definitely worth it.
However, that's the risk you run with single points of failure. Put all your data on one big box, and any failure in your RAID hardware, RAID firmware, RAID drivers, network drivers, kernel, RAM, OS, et cetera will take down the big box and thus take down anything relying on it.
The lesson I learned wasn't to make a super-robust single system, it was to have enough redundancy to stay up when something inevitably fails.
Fortunately, this can usually be rectified with a simple application of money. Good hardware is its own reward.
However, it sounds like the problem that S3 cured was caused by a bad architecture.
Supermicro chassis re a big improvement overdoing your own wiring. The Areca controllers, especially in raid 6, re great. For raid 5 I'd also look at 3ware.
For single gig e, you can get away with esata expanders, building something like the back blaze pod. I've done that kind of thing for personal use, and to have an onsite mirror of something, but I'd want several, in several different colo facilities, to compare with s3. The exception is if you need some kind of scratch storage, but even refilling a 40TB archive with downloadable content takes a really long time over a 1Gbps link.
I'd build a few boxes like this now, but the Thai floods pushed hard drive prices up to the point I have to wait. Hopefully fast 4-5TB drives will be $200 by summer 2012.
Before anyone replies about needing a "24/7 sysadmin", if you are running stuff at this scale:
A) You already have a sysdmin, or somebody competent enough to do the setup/maintenance work. Running off s3 doesn't magically mean you never need deal with sysadmin issues.
B) Rack providers will swap out hardware if you put in a ticket, so you can provide multi-regional redundancy for much much much cheaper than amazon would charge you ($60000*2 for 2 regions gets expensive fast)
Anything you haven't tried, doesn't work. That's a truism in computing. I don't think Amazon tries failing-over entire data centers very often (have they ever?), so when it needed to happen, it didn't work.
Anyway, I'm thinking this guy has only to back up his photo store about once a day to (something big) and put it in his bank box, and he's good to go, at least for a photo site.
(Aside) I wonder if this is more accurately stated: Anything you haven't tried recently, doesn't work. Even if you have tried it recently, that's no guarantee.
Or, to paraphrase Hofstadter's Law, the probability of something working some time after you last tried it drops surprisingly quickly, even when you take into account this rule.
Buying a ton of parts, carefully assembling them, and having it be your problem when something breaks is simpler than paying Amazon to solve the problem nearly perfectly?
Building your own solution to beat S3 is certainly viable but at this scale I doubt it.
Its not. The hard part of storage is to understand the capabilities and limitations of your storage and more importantly how those fit in with your computational needs. You have to do that at amazon on in house. A3 just makes it easy to ignore evaluating their service because you're never actually forced to. Maintaining servers is butter.
I wonder if they subconsciously undervalue their photos, or if this is just a naive 'we can do better' moment.
I'm all for building your own, but there's a good reason enterprise server hardware costs more, and there's a good reason to go for enterprise hardware when you say 'The one most valuable asset at Oyster.com is our photo collection'.
For instance a Dell or HP storage server with 24 disks would be around 18K list - and be really engineered for that as opposed to hacked together.
I wonder why no one ever factors the cost of having a knowledgeable person handling the system into their calculations. TBH 40TB doesn't sound as much, but once you start growing you'll want someone familiar enough with the subject to take care of it (especially if it's their most valuable asset).
Also for some applications it really does make sense to move away from S3, and have a solution in house. For example The Broad Institute has about 6 petabytes of storage [1]. They in particular benefit from local storage, since all of their data is generated on-site. However, even at this scale, they don't build the boxes themselves [2].
[1] http://www.genome.gov/27538886
[2] http://www.isilon.com/press-release/isilon-iq-powers-data-st...
enough with people defending s3. Go buy a IBM mainframe already :)
If I needed this much space in a start-up I'd totally do it myself - it would be a good use of my time. No, my concern with this setup is that it looks (to me) to be a "w00t I got $6000 so lets build a 733t boxen!" It looks cool and fragile, instead of boring and robust.
UPDATE: To be fair, I'm not hammering those drives continuously. How much does that change the equation? Even if it goes from 45 minutes every two years to 45 minutes per day, you'd still be net positive. Basically, you can't prevent drive failures. But if you can keep the failures restricted to the ones that just require popping in a new drive, then your "operator time" is minimal. Its when you lose an array that you're in trouble. Again: Do not use one big box!
- A box from abmx.com. Really nice people, good hardware.
- An Areca RAID controller (might have to talk to them about this).
- FreeBSD, which has good support for Areca controllers, thanks to Areca's friendliness towards the open source community.
- ZFS.
This particular box has had almost no downtime in the past year, and the downtime it's had has been caused either by administration tasks (software updates or configuration modification) or a USB flash drive failure (we've been experimenting with BSD and Linux server configurations on flash media separate from the server's data storage, so that we can do things like show up with replacement or upgraded configurations, plug them in, and be done. USB flash media is not reliable enough for this though, even if you're not doing swap on it).
I just added a couple of 3TB Hitachi drives to the array last week. Resizing the volume was a little fiddly because the Areca firmware needed an update first before it could recognize the 3TB drives (not a problem you'd have on newer models), but otherwise, everything happened live, while the box was up and running and doing its job. When the Areca controller software finished resizing the volume, zfs happily said, "oh, I'm bigger now! I can handle that!"
OpenBSD's softraid is also pretty good stuff, but I don't think it makes sense to use it unless you're trying to cheap out with a small 2-bay box and no controller in a RAID 1 or something.
SBS 2008 is not my favorite thing in the world.
1. If you are using a hardware raid solution, then RAID1, and only if one of the drives can be pulled out and used by a standard SATA controller to get the data off. Controllers that munge the MBR just go in the trash.
2. Otherwise software raid1 or raid10.
Raid5 and Raid6 IMHO are more trouble than they are worth. Now, true, I'm not spending $6k on drives, so maybe at that quantity, Raid6 makes more sense. But there I would argue that the money you save on drives, you throw away as soon as it goes tits up. I've tried Raid5 with promise and areca cards and the performance was not as good as software raid10. Plus, on several occasions the transfer to the hot-spare failed and brought the array down. The data was there, but it required operator assistance (me). Until it went down, it was staggeringly slow.
If you really want to make sure that the array is available, use Raid1 with three or more drives. They are just so cheap. I was responsible for a windows server at a start-up a long while ago, and Raid5 sucked ass. Bonus: one day two drives failed. Tape restore is slow. Oh yeah: Don't just buy four drives and slap them in. Quite often a whole batch of drives will be bad. Buy drives from different manufacturers and beware that the sizes will be every so slightly different, so make your partitions smaller than the smallest.
However, all of this is irrelevant if you really want a lot of storage. For that it seems that a distributed, redundant file system would be best. If you want, you should be able to pull a drive or two out of every server, or all the drives out of a few servers, and still have all your data.
I used to run RAID hardware cards but changed to software-only solutions after a hardware card crashed. Unless you are going to buy two of the exact same RAID card (one as a hotspare), or unless performance is a big deal (which it shouldn't be on a backup system I would think), software-only solutions are the way to go.
I complement this with off-site continous backups of irreplaceable data (about 200GB) to CrashPlan.
I really wanted to use ZFS, but the Linux support isn't that great, and also I don't think it's possible to remove drives from ZFS arrays. Both of those were deal killers for me. In any event, LVM+XFS has worked out great. XFS is a very stable file system and has given me no trouble.
http://tech.oyster.com/how-to-build-a-40tb-file-server/?#com...
"Didn’t mean to give the impression that this is the only backup, it is not. It is the “warmest” one — first line. It does not share the same physical location with the primary storage box either, but is less than 2 miles away. So we can easily have a 2 hour recovery time without throwing away those astronomical monthly service fees. (Although many technologists will always prefer paying big bucks for the comfort of being cushioned from every angle by SLAs and such — nothing wrong with that, just a different approach.)
As far as TCO goes, it cannot get any lower since we already have one or two system guys handling all servers, as well as office workstations, etc.. This backup box takes up such a small fraction of their time that its almost negligible — several thousand annually at most. Same goes for power, etc — it is just one of many servers.
The disks are all Enterprise Class 2TB SATA-II, several different models. We were purchasing them right after the monsoon floods in Thailand constricted supply so our choices were somewhat limited as time was a factor.
Raid6 has come a long way since it’s early inception days, but is still a trade-off between raw storage capacity and processor utilization. HW RAID industry is now old enough to not have to wait for new products to mature as we used to when the technology itself was in its infancy. Old habits certainly die hard, but getting the “latest and greatest” was a conscious choice made for this specific problem, not submission to some immature fascination with “elite” new products, or however that may be.. This card has the best specs for Raid6 currently on the market — bottom line, period.
Big Kudos to all who made suggestions and participate in the discussion, keep it coming!"
Considering the total cost of 6000$ for a single server, you would be dropping 12k from the get-go for two servers (though you should really lease over a 12 month period if you can't afford it) and then you need to add colo costs for the machines, which can be around 300-400$ per machine in a single unit colocation environment, depending on your location. This brings us to 21600$ in the first year for two machines colocated in separate datacenters, without any sysadmin costs;
The article states S3 would cost 60,000$ per year, and this doesn't include any bandwidth costs, whereas the colocation setup would include some decent bandwidth per month. Over time, it's easy to see how a lot of money can be saved.
Also, S3 is designed for 99.99% availability, but they only guarantee 99.9% through SLA, which isn't extremely spectacular. Hitting 99.9% isn't particularly hard if you have a sys admin worth his salt to set things up right, and with your own solution you can have more direct access to the file system as well as the ability to adhere to any regulations.
My old desktop hardware (Intel Q6600, 8GB RAM, Asus MB, pair of gigabit links bonded)
Supermicro 4u tower case with 8x hot swap bays + the 5.25 bay filled with a 5x hot swap cage, 13 total hot swap bays.
3-Ware 9550SX RAID Controller, 4 x WB RE 320GB drives, RAID 5.
8 x 1.5TB "Green" Drives, mixture of WD and Samsung drives, RAID 6 using mdadm (linux software raid)
1 x WD Raptor (system drive)
Ubuntu with KVM for virtutalization
Originally i was in the "must have raid controller" camp which is when I bought the 3-ware controller and the RE series drives. When that array filled up I did some research and decided to just go the mdadm route and have minimal complaints so far. I use the 3-ware array still for "critical data" and back it up to S3.
Monitoring: You have to work a bit more to get proper alerting of issues from mdadm but its not hard to setup. Doesn't matter for me much, i sit in the same room as this server most of the day working so if something goes wrong i generally notice before i get the e-mail.
Performance: As mentioned I'm using Green drives, this system wasn't built with speed in mind but rather large amounts of nearline storage. Never-the-less, with some basic tweaking and making sure the array's partition alignment is correct I get around 450MB/sec read speed and ~85MB/sec write long term, I have the system setup to cache writes aggressively however as most of the data on this array isn't critical and its on a UPS, what this means is that writes under a few gigabytes usually complete at wire speed (~200MB/sec) then get flushed to the disk later. Most of the time I'm limited by network bandwidth to this system, unless I'm writing a very large amount of data all at once.
One negative is rebuild speed, here i'm very limited by the 'Green' drives I believe. It runs at about 50MB/sec so rebuilds do take awhile.
As far as CPU usage goes, I've never seen it be the limiter but haven't watched it that closely, it doesn't peg during rebuild. This machine acts as an SMB/NFS file server and a few development VMs 24/7 (database, and a couple other things) and I've never really had an issue with cpu usage.
One really nice bonus is that if something in the system fails, you can just plug the drives of your array into to almost any other linux system:
apt-get install mdadm
mdadm --assemble --scan
Poof, working array.
tl;dr
If your going after 1GB/sec transfer speeds, get a high end raid card.
If you just need some large redundant storage that can saturate a Gbit link or 2, then mdadm software RAID is just fine IMO.
http://blog.backblaze.com/2009/09/01/petabytes-on-a-budget-h...
and version 2
http://blog.backblaze.com/2011/07/20/petabytes-on-a-budget-v...
http://lists.basho.com/pipermail/riak-users_lists.basho.com/...
One more question: you saturated the network link with a sequential read/write, but is that how you actually store the data? If not, how long would it take you to be up and running on another CDN in case Akamai goes down in flames?
This matches my experience, RAID performance in linear IO is an order of magnitude below what the disk should allow. This guy is relieved to finally get 1 gig of useful bandwidth out of 24 disks (about 1 gig a piece). So it's no faster than a single disk.
(i know it's linear io in this case because the screenshot shows a large file copy)
Nothing to worry about here, I think we are in good hands.