MySQL Cluster 7.2 GA Released, Delivers 1 Billion Queries per Minute
mysql.com
mysql.com
1) They used Infiniband interconnects. Running on ethernet is likely to yield less impressive results.
2) Their benchmark does simple primary key lookups. If you start doing joins or transactions that need to hit multiple data nodes, things will slow down. Depending on your workload, this may or may not be an issue.
3) NDB is an in-memory storage engine, so you're limited to the aggregate RAM in your cluster for max storage size.
4) AFAIK, MySQL Cluster doesn't re-balance, you need to pre-determine how data is partitioned and changing it at runtime is hard. I don't know if this has changed in the later releases.
I'm sure for the benchmark it was all in memory though.
Can someone who knows this stuff give a bit of an overview on the significance of this release?
I read the NoSQL stuff as Oracle/MySQL trying to compete (at least in terms of marketing-speak) with the wave of competition that's arrived in the DB market. Is there any meat to it?
This release adds an additional memcache api layer which sits on top of NDB API, so that you can more easily program against it.
Cluster has "sql nodes" (that you query against) and "data nodes" that house the actual data, which "sql nodes" talk to to actually process the query.
Cluster can now tell the data nodes to only return necessary subsets of data back to the "sql nodes" when doing a JOIN, instead of pulling down more information (and filtering on the sql node side) when answering complex JOIN queries.
You can read more about this at "push down" feature at http://www.clusterdb.com/mysql-cluster/trying-out-mysql-push...
Note that the linked article also outlines more about this feature under the heading "70x Higher JOIN Performance with Adaptive Query Localization".
Automatic sharding and memcached integration are pretty awesome features, and could definitely ease code at the application level (sharding code is a particularly special pain in the ass, not so much getting it working, but allowing for re-sharding migrations if you decide you need more shards, especially trying to do so without downtime which involves all sorts of nasty tradeoffs).
But bang-for-the-buck and reliability wise, I'm still unsure, I've heard very little about this in the wild. Is this more suited for high write v. read ratio situations, or is it aiming more at the Vertica/Greenplum big-data uses, or something else?
NDB was originally created by and for telecoms. So very high write/read rates with very fast response times and very high availability required. Generally not extremely large datasets.
If you know what you are doing, NDB can work extremely well. It supports all of the highend cluster goodies including online software upgrade, online node addition, automatic handling of node failures, geographic async replication, etc...
However, it certainly has a pretty steep learning curve from the admin point of view and it is a bit easy to mess things up. It is a bit brittle due to this, but once it is setup properly and running, it can deliver on the promises.
Check out Postgres or Cassandra.
If this has improved lately, I'd be very interested to hear about the changes!
Since our complexity requirements were low we settled for a home-grown solution (basically memcached with write-through) and so far didn't regret.
I, too, remain curious if anyone is running this at scale and has bumped into the various corner cases (exceeding capacity, hardware dying, etc.).
I know I must have been doing something horribly wrong, but I never could really figure out what it was.
But no info on what disk setup they used. But for every instance node there was a storage node.
You know who has scheduled maintenance? Apple, Verizon, Sprint, etc.
0) Always a good idea to backup! 1) Create your cluster... http://www.clusterdb.com/mysql-cluster/deploying-mysql-clust... - stop before starting your mysqlds 2)Configure your mysqlds to use the same data directory as your existing ones 3) Stop your existing mysqlds 4) Start up your new (cluster) mysqlds. At this point you still have your existing tables stored in the original engine (e.g. MyISAM or InnoDB). 5) Migrate appropriate tables to cluster... "ALTER TABLE <tab-name> ENGINE=ndbcluster;"