How We Built A Data Center With Commodity Hardware And FOSS
searchenabler.com
searchenabler.com
I worked for a design agency back in about 2004 which had the same sort of infrastructure internally. They had a cheap desktop Exchange 2003 box, cheap desktop Fileserver, cheap desktop SQL + App server, cheap desktop mp3 server, big old DLT drive, SDSL router, switch, a KVM, cranky old monitor, keyboard, mouse, mounds of cable, a well underpowered UPS and a garfield all wedged under a couple of tables.
Some muppet knocked over a coffee cup and it spilled through the gap between the tables into the monitor causing a cascade failure and a small fire which melted pretty much everything under the table. It took the company out for 2 weeks until Dell could ship new kit in.
There is a reason to do it properly IMHO (and experience!)
We tried to control cost by assembling stuff, using commodity hardware etc. But did not use low quality stuff. Our UPS is from Eaton, Server components are standard ones like Seagate, Intel, Router/Switches from Cisco, D-link.
Best of luck - sounds good.
I built a couple of racks for clustered decryption using commodity hardware; after a few months "burn in" the failure rate proved to be way higher than expected. Which, when you think about it, is logical - they are running 24/7 (compared to if you were using the hardware "normally").
After having to replace three servers inside a month we started to cycle into server hardware. This was a lot more expensive up front, but is reliable and the failure rate is almost non-existent in comparison.
So; buyer beware.
Even so; it is good to see someone else hacking together their infrastructure :)
We had a couple of large RAIDs using commodity TB hard drives and these dropped like flies.
In our case, we, among other things, do a lot of hardware manufacturing (as in buy the chips and build our own boards). Because of this we invested in custom conductive flooring and everything else above the floor is also static-aware. For example, all rolling carts and shelves are wired with redundant contacts that touch the floor, thereby making a connection.
The above is an extreme for the occasional builder. At a minimum you should setup a working environment that has a conductive surface like an static-dissipative mat and wear a conductive wrist strap connected to the mat as well. There are also ionizers and other tools that would be good investments.
Handling components around server racks should never happen without regards for static electricity safety. The rack should be grounded and you should wear a conductive wrist strap connected to the rack.
I'll bet that a lot of failures in consumer/commodity builds are due to this very issue and not necessarily connected to bad or questionable components. If you buy components built by reputable manufacturers they will have been manufactured in tightly controlled environments already. Don't break the reliability chain by handling these components in uncontrolled environments.
We stuck the servers on top of a few pallets so they wouldn't be submerged in water when it rained. The power was fed from the living room above and the network cables pushed through gaps in the floor boards to all the bedrooms. The adsl router was nailed to the wall on the top floor above the beer fridge and ran smoothwall I think it was.
We even tried running a serial link for a monitoring terminal in the living room over standard audio cable .. it kinda worked .. and was called teethanet for some reason.
Overall it worked like a charm and I don't remember there ever being downtime .. even when there was a foot of water surrounding the servers.
Also, I just realised that I may have come across as a bit patronising! That wasn't the intention, I was just relating to what can be done with a lack of resources and a bit of ingenuity :)
The main reason why you should go with real server type components is the support lifetime. You can be fairly confident that you can get replacement parts be it CPU, memory, raid controller, motherboard, or disk drive, for at least 5 years after purchase. Otherwise you will find that when your components die, you will be completely unable to repair them due to lack of available parts.
Good luck to you, but I've been down this road before and if you cut corners in the beginning, it usually costs you more money to do it right later on.
Hardware availability after the initial purchase is certainly a big longterm consideration. I can easily search for "dell 2950 raid" or "dell perc 5i", match the individual part number to the I have, and be good.
What RAID controllers require replacement drives with the exact same model and disk geometry? I've not seen one that wouldn't take a bigger drive as a replacement.
I also recall back in college setting up a 110 node linux cluster. Before they upgraded the HVAC system in the server room the cluster generated so much heat you the power cables started to shows signs of heat damage (we had to shut the cluster down until the HVAC was beefed up).
I've also had a bunch UPS where the batteries fail, and need to be replaced periodically. I also wonder how that figures into the cost of their data center (it's an issue at a big data center as well).
I love what you've done for so little cost.
I have UPS and generator at home; the cost of running a power generator exceeds $1,000 a month but it is still not reliable.
Wouldn't servers at Hetzner cost cheaper and more reliable than hosting in your garage in India? You can get an i7 server with 16 GB RAM is for $60 and 32 GB for $72, and no hardware setup & maintenance cost.
http://www.hetzner.de/en/hosting/produktmatrix/rootserver-pr...
From Hertzer pricing page, add 1Gbit port cost of 50$ a month, data transfer charges etc would make it much more expensive. Considering our usage, we felt private infra would be a better choice.
I commented from what i got from pricing page. Apart from basic server cost, there's always associated extension hardware/data/software cost which generally makes it much more expensive.
If you have more details on cloud setup using Hetzner and associated cost, we would be happy to explore.
[1] http://devblog.supportbee.com/2012/06/04/building-a-dependable-hosting-stack-using-hetzners-servers/Whoa, this is pretty shocking to me. I knew there were outages, but I'd assumed they were more like 10 minutes at a time. Do all tech startups there have to deal with this problem? Are there office parks catering to businesses that need continuous power?
In my experience the computer failures are relatively easy to manage. Power will be your primary issue, and to a lesser degree, internet. When we ran a lot of servers on UPSs, the UPSs start to become a major pain. UPSs fail much more often than hard-drives. And when the battery dies it cuts off power to the computer. You're using low-end desktop machines, so you probably don't have dual PSUs in each machine (if you did you could run each PSU to a different UPS). UPSs are also much more likely to fail when the grid power fails. So the power goes out, your UPSs will kick in, then 30 seconds later one or TWO fail. Can your power distribution can handle that? I'm no UPS expert, but this seemed pretty common across a variety of UPS products.
As for internet, multiple links is a good start. Load-balancing outbound traffic is fairly straight-forward. Inbound traffic is much harder. Really you need BGP. A BGP-based set of links isn't cheap, and doing the configuration yourself is hard if you don't have experience with it. The alternatives that aren't IP routing based are usually based on DNS, which is better than nothing. But the effectiveness of it is dependent on how your users servers/clients/browsers implement DNS caching. It will work OK for some users but some will cache the stale link IP for 24 hours -- not cool.
You certainly can work this way, and probably get three nines of uptime. But I can say now that we switched to a decent colo I sleep a lot better. We can focus more on building product, adding value for our customers, than screwing around with the utilities.
FWIW I'm basing this off of my experience in the United States.
Edit: Never mind, I just realised colo won't work for you because they'd probably expect servers that fit correctly in their racks and your computer cabinets probably won't.
Interesting project and I wish you luck with it as your project scales up.
http://en.wikipedia.org/wiki/Jugaad
I only know the word from Sikh prayers in the sacred language of Gurmukhi. For instance, it's part of the Mul Mantra, in which it means "throughout the ages". As I'm not well versed in Punjabi or Hindi, I didn't know it had other meanings.