So long AWS, and thanks for all the fish: FathomDB (W08) focusing on its new DB
blog.fathomdb.com
blog.fathomdb.com
FathomDB customers saw outages because the implementation Fathom did using MySQL on AWS went out when AWS went out.
So in conclusion Fathom is going to build something which Oracle, IBM, SyBase, and every other database company in the history of the planet has tried and failed to do, build a relational database on a distributed infrastucture that is more reliable than the underlying infrastructure? Uh, good luck with that. Seriously, its Turing Prize material if you succeed.
I was thinking the punch line would be "Gee Netflix uses AWS and they didn't go down so we're going to more what they did." I guess they are going to compete with MongoDB and Riak.
Incidentally, Oracle and IBM both have clustered database products, though they can certainly be improved upon.
You might find Greenplum and VoltDB to be more interesting models to study.
The key question is: are you aiming at OLTP or OLAP? Because, as those boring mainframe-era guys have discovered, it's difficult to serve two very different purposes.
In particular, how do you see the tradeoff between durability and performance?
However, the AWS outage has (we hope) put reliability back on the map. So maybe now we don't have to be the fastest database, if we're the fastest reliable database. That's a liberating change of perspective, which has been behind many of the recent advances.
For example, if durability is an absolute must (as it is for current databases), then your main bottlenecks would be a) spinning rust and b) network traffic to distribute copies of the data throughout the network. SSDs will help this a lot, but already you're back where almost everyone has already started.
If you're prepared to accept highly replicated in-memory copies as "durable", you can already make your overall system hundreds or even thousands of times faster.
So how do you define durability?
1 copy on 1 disk?
x copies on y disks in the same computer?
x copies in y disks in different computers?
x copies in RAM of y different computers?There were multiply redundant switches and massive investments in transactional data recovery. The big finance guys were particularly intimidating when they talked about several hundred billion dollars in transactions in a 24 hour period, that is $4,000 - $5,000 per mS, or 4 to 5 million dollars per second. Combined with the whole finance game of not confirming receipt until the last possible moment to keep your liabilities managable, ouch.
I would love to see a relational database emerge that could do ACID on an unreliable cloud, and I think the folks who achieve that should become gazillionaires. I was just noting that a whole lot of money and research has already been poured into that hole. We're still waiting to here the 'thump' of it hitting the bottom. :-)
It totally rocks that Fathom is taking on that challenge.
why does it seem that everyone expects 100% uptime from a VM just because it's in "the cloud"? shouldn't "the cloud" be used to make fault-tolerance even easier, because you have access to multiple geographic regions and multiple providers with little fuss?
they're still computers, and they're still bound to go down occasionally. i don't see how running leased VMs somehow absolves you of doing basic operations work and guarantees a bulletproof experience. regardless of where you host it, writing a nicely distributed and fault-tolerant system is and always has been difficult. the only part that's gotten easier is finding rack space.
However, if the provider doesn't keep those promises, then all your hard work and calculations go totally out the window.
At that point, you have to figure out a way to run a database on a system that effectively offers you no guarantees. That's what we're working on.
If the datacenter you are in has a slight conditioning problem this summer and 10% of your drives breaks down due to excessive heat, how quickly will you be able to re-provision the data center?
About a week after AWS outage the italian ISP Aruba had a UPS failure due to a fire in the UPS room and the entire datacenter switched off automatically during the night. For the following 8 hours that datacenter was off for every customer. How would your new solution handle such a situation?
Designing a hot standby replication solution for MySQL/PostgreSQL that works across regions seems easier to me rather than implementing a database from scratch that should solve a very complex problem.
I want you to succeed, but you're dealing with seriously hairy deep magic issues that the best minds in the industry and academia have been chipping away at for decades.
it's like when you have a big building project and there's a provision that the contractor will refund $5000/day for every day past march 1st. doesn't mean march 1st never gets passed.
failure modes in the cloud are known too: your VMs are either working or they aren't, or maybe they're somehow degraded but you should fail over anyway. it's very similar to physical hardware. what will you think when a backhoe takes out your datacenter's fiber for 10 hours? that physical hardware and datacenters are now unreliable too?
as for a database that runs on a system that offers no guarantees -- isn't that all of them?
Separately from the SLA, technical promises/guarantees that AWS did make e.g. isolated AZs were broken in the April outage.
I think that your proposed model ("machine is online or not") may be sufficiently simple that the AWS cloud can satisfy it; however I think it is very difficult to build anything interesting if that is the only axiom you have. In particular, I would want something related to persistent storage in the model, or else storing state becomes very difficult.
What cloud provider offers a more reasonable SLA?
In terms of who offers a better cloud, clouds based on OpenStack are 'the great hope'. IaaS is about commoditizing infrastructure, so open source seems necessary; rather than AWS's very closed, secretive approach.
An SLA (even with severe penalties) is certainly setting up the right incentives on paper, but in practice it has little effect on the results. Why? Because enforcing a promise on a piece of paper is bothersome, expensive and sours the relationship.
Also, saying you'll pay 50x penalty is only half the story. What are the triggering events that entitle customers to payment? Is the SLA measured end-to-end (i.e., from the user's endpoint to yours)? Is the SLA penalty cumulative? Can your customers actually profit from your downtime (i.e., you would end up owing them more cash than they paid you)?
Without more details, your post suggests that each fraction of a second of downtime experienced by any single user of your service entitles them to a credit of 50x what they paid for that fraction of a second. Is that your policy?
Most cloud providers (including FathomDB) charge by the hour anyway, so you have termination rights for any reason, and there are no prepayments to refund.
We think that 'doing the right thing' is the only workable policy, and it shouldn't require an outcry from a vocal userbase. We should have handled the outage better, and we're making a big payment to demonstrate that. We're not going to let any new customers use a system where we can't be confident in its uptime (MySQL on AWS); we're building a new system in which we can be confident.
Cheers!