Percona XtraDB Cluster: Setting up a simple cluster
mysqlperformanceblog.com
mysqlperformanceblog.com
It's a great setup (although you have to work with it, just like most sql clustering setups out there, you can't just throw anything on it and hope it sticks).
I'm a mysql newbie but am very fascinated with scaling relational databases. It's just seems so challenging and exciting.
For backups we use another Percona tool: http://www.percona.com/software/percona-xtrabackup It's free and works great.
Disaster recovery is standard stuff. Backups/hot backups to external and recover as necessary (generally we try to keep that to never).
Monitoring is outside of the scope of XtraDB, but take a look at MONYog.
Now clustering is more XtraDB related. The clustering is nearly dead simple. If you use the (technically deprecated) wsrep_urls configuration option you can spin up a cluster from scratch without needing to worry about master location. Not sure WHY they deprecated this, I guess some people had issues with it, but it has worked fine since the beginning for us.
You don't need a master/slave relationship unless your writes are really sky-high, or same with reads. Basically make sure you have a good number of wsrep_slave_threads defined (double CPU cores seems to work well for starters).
Sharding, partitioning and archiving are kind of outside the direct scope of XtraDB. You'd be better served looking up documentation online for more info than I'd be able to provide in a short post here. Insofar as it ties in to XtraDB though, I think there's nothing wrong with jumping straight from a single DB node to this solution. You just really have to have the additional CPU/memory resources to add more nodes (you need at least three to keep the cluster alive if one dies... two will get you split brain if one dies). I believe e-mail is in profile if you have more questions, but this should cover the big stuff.
It is not. `email` field is not publicly visible, you need to add it to `about` text field.
It's an entirely different scenario for a data-warehouse type situation obviously, and having 30+ second writes with enough threads won't be an issue then either. Basically though, this is one area where I'd like to see Galera change in one way and it's kind of the 'one' valid criticism to have for the platform. If your workload falls within the good use-cases for it though, I'd recommend it well before NDB/MySQL Cluster due to if nothing else, the sheer complexity of scaling Cluster.
We have a high-write application and while individual writes aren't really that slow we have A LOT of them (we receive our writes off the network from a 20+ hadoop-machine cluster so they can send quite a bit of junk our way).
What's wrong with the complexity of scaling Cluster? From what I've experienced it's been pretty easy to add machines, take individual data nodes offline while cluster is running, etc..
We also needed something that would speed up writes much more than just 2x (and the way ndbcluster does writes was the key).
How many nodes did you deploy? Are they within a single DC or cross WAN? Do your clients write to all nodes, or just one? Basically, how much can you share about your topology?
I can say that getting it working with selinux is a pain. I have a TE file available if someone needs to get it working which I can upload somewhere.