Aurora - New MySQL-Compatible Database Engine
aws.amazon.com
aws.amazon.com
To me, there is a very interesting contrast to be had between this announcement and Microsoft's announcement: it feels like Microsoft is discovering the business value of being open at the same time that Amazon is living in the time warp of proprietary everything. Has Microsoft internalized that open source is (or can be) a differentiator in the cloud? Amazon is clearly still oblivious to it -- and it will be very interesting to see if this service generates fear of vendor lock-in...
And in Amazon Web Services, the biggest thing Amazon has to fear is the OpenStack initiative producing a platform that lots of little hosts can host.
At the same time, Amazon is stock market time bomb that will either start extracting monopoly rents and draw regulatory ire, or whose stock will drop, because of the amount of revenue multiples.
That said ... this is the way it nearly always works. Microsoft invented .NET over 10 years ago, and has kept it proprietary, leveraging the dominance of Windows, extracting rents. Now that Mobile is the new hotness and Microsoft needs to catch up, it goes open source. Google released Android for the same reason.
In short ... the companies that devote a lot of resources to R&D usually like to recoup the costs, but eventually PLATFORMS get opened. Frankly, the world benefits more when this happens sooner (Netscape open sources its codebase, the Web is designed, etc.) but the engine of progress in capitalism is still speculative investment, and that needs to pay off for years before the public at large can benefit, if investors are to continue investing.
Once we switch to crowdfunding, however, this model of capitalism and risktaking may be disrupted as well.
I highly doubt OpenStack alone can make it possible for smaller players to replicate AWS's economics of scale, or even come particularly close. From what I understand Amazon doesn't just plop their software on commodity hardware, they use highly specialized gear optimized for running specific services. And that's not to mention buying power, the ability to run on razor-thin margins and other business factors that impact pricing.
It's commodity in the sense that Amazon buys it from commodity vendors, but they use custom designs built to address the specific parameters of their data centers. Very few other companies can afford to employ an army of engineers to tweak every last component for optimal price-performance, nor for that matter do they have the benefit of accounting for a noticeable share of a supplier's revenue when entering the negotiations room.
> Smaller companies can and do offer much better pricing, they just don't have the marketing budget to convince you of it.
I am aware of at least one player (DigitalOcean) offering lower pricing than AWS, but that's only in a fairly limited niche. So while they may have the "economies" part checked they don't have the scale and scope to compete with Amazon on a particularly substantial level. I'd also be curious how they match up to Amazon in terms of operational efficiency since they lack many of the the advantages I listed here and in my previous post, which give AWS an edge in the race to zero as time passes.
So when the moment comes and you want to move away from AWS and you have scaled well beyond the capacity of a single box, what then? Suddenly you have to figure out sharding, replication and HA setups from scratch. While that may be the right choice for you if you're an early stage startup and it lets you defer engineering costs until a point where you're hopefully better funded, it's also dangerous:
It's far easier to get things like sharding and caching and use of read-replicas and multi-master replication right if you think about it from the beginning (if you've ever tried to retrofit any of this onto an application not designed with it in mind you know what I'm am talking about). It'd be awesome not to have to worry about this, but if you don't, you tie yourself into a platform you'll find it harder and harder to move off.
As a product, this solution would be a killer - it sounds fantastic. As an AWS-only service, I for one will treat it as if it doesn't exist, because getting tied to AWS for the long term is not an option for most of the projects I do.
"Baseline storage performance is rapid, reliable and predictable—it scales linearly as you store more data, and allows you to burst to higher rates on occasion."
You mean to say that I can scale my single MySQL instance linearly with data storage by just throwing bigger SSDs at it? That has not been my experience, please share how you have been able to accomplish this?
They've even mentioned it's pretty much a click-through to migrate from mysql... to me that seems like the opposite of lock-in.
As to only being able to use their infrastructure to scale this way, that's a value-add.. that's part of why you pay for SaaS instead of developing your own infrastructure. This is why you're renting instead of buying.
This just leaves the other smaller players (DigitalOcean, Joyent, Rackspace, etc) who are mostly offering something akin to EC2 and then partnering with other vendors to offer the missing pieces on top (frankly what other choice do they have? They can't get into a race with Amazon/Google on who can build the most no. of services - they will never win that race).
OpenStack's model seems to be "clone what Amazon does, 18 months later then hand over to vendors who don't have the scale to match Amazon's prices"
It's kind of nice in theory for internal clouds.
But I'm increasingly seeing tooling targeting Docker instead of OpenStack at the service level for doing the same kinds of things OpenStack is supposed to offer (ie, the service utilities Docker to offer automatic deployment/availability/loadbalancing instead of doing it using the OpenStack APIs).
At the cloud vendor layer, the differences between vendors can be abstracted using libraries, which removes an alleged attraction of OpenStack.
Given that, I see a few vendors challenging Amazon by building unique selling points (Google has some innovative things as does Microsoft, and Digital Ocean pushes the price/performance thing).
I see Docker taking away a lot of the "cross cloud deployment" attraction that OpenStack had, and doing it better.
So what does OpenStack offer end users? (I understand it's attractions to vendors, and maybe the internal cloud use case).
If not, then what utility is OpenStack providing?
Bare metal works well. I see many people using either Docker on bare metal or maybe a VM layer as an additional security layer, but using Docker as the deployment target.
But I agree with you that OpenStack is not good enough to be worth it for that with perhaps the exception of very large multi-tenant deployments (in which case, my first goal would be to start rewriting large chunks of it). And parts of it are just so horribly over-engineered it is scary, because it makes me wonder what it is they're missing to think it was necessary to make things that convoluted, and what they're missing because they've made it so convoluted.
I agree entirely with you on the rest of your comment.
So the source code for Amazon Web Services is clearly architected, free of cruft and under-/over-engineering? I assume you've worked with AWS's source code in order to be able to make such a comparison. In my experience, the design and architecture of closed-source internal applications delivered to customers only as a front-end is usually a nightmare.
<rant>
There might come along something that is actually good, but Openstack isn't it. The architecture is just bad, it will never be as reliable as something like AWS. (not because of scale, purely because of lack of error handling capabilities)
Reality is these sorts of orchestration systems need to be written by specialists. The vast majority of Openstack was written by people that admit they have no clue about systems level programming. This is what hype does, it forces a whole bunch of bad programming and architecture down everyones throats.
The marketecture and hype machine have done horrible things to Openstack. Not only have vendors riddled the thing with lockin and crap code that needs to be supported, but they have pushed entire projects that should have been shot in the head. cough Ceilometer cough
There is security, performance and plain availability problems everywhere, mostly embedded deep in the architecture.
Dumb decisions like "All systems must be Python" when Python is clearly not suited to a large number of the things they want to do is painful. As is the general "not invented here" syndrome and boys club that leads to certain libraries or patterns being pushed over others, usually to the peril of the project due to the 'blessed' thing being incomplete and unproven.
</rant>
I don't intend on returning to Openstack if I can avoid it, unfortunately my skill-set does tend towards that sort of thing so we will see how I go.
Amazon puts all its SDK stuff open source on Github, and I don't see Microsoft open sourcing its tech behind Azure.
I will agree on a lot of the tech.. I think, by comparison DocumentDB is horribly locked-in without an API-compatible version you can self-host (even if it doesn't use the same tech, or perform as well) is a big warning flag imo.
In this case, being MySQL compatible is anything but a lock-in.
Compatibility at the connection level is just one part of a whole lot of issues to take into account when considering whether or not you can migrate elsewhere easily.
Given how expensive AWS is, that's something to seriously consider.
It is the same old play book as they used for Exchange and IIS.
By comparison, look at Azure's DocumentDB or Search services both of which obscure the interfaces to the underlying ElasticSearch ensuring lock-in.
Honest,y both have value, but this really isn't the same by any stretch.
On the infrastructure level, the "lock-in" part is not so much the technology, all of that is relatively easy to replace. But it would take me two additional FTE's to configure and manage everything AWS offers as a service ourselves.
But that applies to a lot of things these days, AWS just happens to be a one stop shop for a growing range of commodity services.
And not to talk about making them highly available, scalable, having to continuously monitor and so on takes time, effort and involves significant opportunity costs.
Imagine business folks telling engineers about an upcoming promotion campaign that bring in 3X more load. Outside of AWS (or cloud provided service) one has to sit and plan for scaling up all the individual components (rabbitmq, haproxy etc.,) which really becomes pain after a few iterations.
and Frequently Asked Questions: http://aws.amazon.com/rds/aurora/faqs/
At $200/month for the entry level, their lowest price is many times what the cheapest geo-replicated "SQL engine as a service" from Google or Microsoft is. I'm not sure how the performance differs, but I am guessing theirs are no slouches.
Microsoft offers "SQL Database" geo-replicated for as low as $15/mo., and it scales up from there. Not sure about performance, but it would be apples to oranges (MySQL versus SQL Server) and difficult to compare. I wonder what the TPC numbers are, but apparently the TPC organization doesn't allow publishing that yet.
Google offers "Google Cloud SQL", also geo-replicated, and their cheapest pricing is between $10 and $18 dollars a month.
A 50GB database with 10GB of RAM usage and 2 cores costs 700USD/Mo with SQL Database. This is replicated twice within DC. Double that price if you want georedundancy. Their entry level costs less but is also unusable for most real applications
1) If you are having trouble with tweaking my.cnf, give the perl script `mysqltuner` a go. It is pretty good.
2) The biggest improvement would be moving the storage to SSDs. I personally prefer Linode to DO, but both are cheap, and good.
My linode servers can do full text non-blocking backups at about 100MB per second, so depending on your outage recovery requirement, you may be able to get away with a simple cron job.
Run "explain" on all your queries and review your indexes to begin with.
That dataset is so small that there should also be no need to be continuously tweaking the config file. 10 minutes looking at a performance guide, and ensuring you have enough ram to load everything into cache should be enough to get your configuration to a "good enough" shape - there's just nothing reasonable you'll be doing with a 300MB dataset that should need anything but getting the very basics of the config right to perform decently.
$200/mo minimum pricetag is extremely high for 300mb of data, or even 10GB of data.
We're looking at migrating a product of ours to Aurora and it has an operational dataset on the order of terabytes.
(inb4 nosql: Why not NoSQL? It needs transactions that don't suck (I'm looking at you Cassandra). We use C* for other large datasets, I do wish we could just use that.)
Using an existing name for a product in a similar space is just confusing and hurts everyone.
A lot of the good names are already taken; there's going to be overlap, and they're very different products.
I'd agree with you from a technical perspective it would have been better to offer Postgres compatibility, rather than MySQL. But I believe the intent is to target both a bigger market (as of today) and challenge Oracle.
It hasn't been claimed, but the article is filled with MySQL comparisons and references. I would not be surprised if it was a MySQL fork.
Would love to see if aurora fixes this for us.
I don't vouch for any of these claims, but MySQL is certainly not the ultimate in DB performance.
"Amazon Aurora delivers significant increases over MySQL performance by tightly integrating the database engine with an SSD-based virtualized storage layer purpose-built for database workloads, reducing writes to the storage system, minimizing lock contention and eliminating delays created by database process threads. Our tests with SysBench show that Amazon Aurora delivers over 500,000 SELECTs/sec and 100,000 updates/sec, five times higher than MySQL running the same benchmark on the same hardware."
People will benchmark this themselves ASAP. If you do, try to include some basic system metrics. The output of "vmstat 10", "iostat -x 10", and some "pidstat -t 1" would be a great start. This may only be possible for the MySQL benchmark, if Aurora is only visible via an API, and the database system can't be accessed directly (?).
The cheapest is a db.r3.large for $0.29 / hr
Is that included in the server cost? Or will you actually be paying ~3x the server + storage costs?
"Storage consumed by your database, up to 64 TB, is billed at $0.10 a GB per month and IOs consumed by your database are billed at $0.20 per million requests"
Sounds like the replication is included in this, but Amazon really should clarify that.
If you go to the normal RDS pricing, there's "Single AZ" and "Multi AZ" pricing. The Multi is 2x Single because it transparently runs the 2nd instance. http://aws.amazon.com/rds/pricing/
I certainly don't want to spend $7 a day playing around with RDS for fun!
I guess since this is MySQL-compatible it wouldn't be that hard to migrate from MySQL to Aurora, other than moving the data...
Moving data from one instance of MySQL to another is usually straightforward as a one-liner, piping a SQL dump from the old database into the new one over the network...
mysqldump -hOldHost | mysql -hNewHost
...and waiting.So what can Aurora do for that workload? Do the support multi-table transactions and referential integrity across all 3 Availability Zones? Similarly, they mentioned Durability targets; what's their targets for Consistency (ie ACID).
I trust your best intention, but even Dropbox was dismissed on HN as user-friendly rsync.
We will see it shortly.)
EDIT: actually it looks like cloudsql might actually be a mysql.. also on gce they max out 16gb of ram and 100gb of storage
Irrespective of that you shouldn't be posting as joe public when in fact you are a Google employee playing cheerleader for Google products on threads about your competitors products. Keep it classy.
Adding "disclaimer: googler" to every post I make on HN seems pretty obnoxious to both me and anyone reading. I just didn't feel that "neat" + a fact qualified. Clearly opinions on that differ, and I'll probably just post less in the future.
"But this stuff wasn’t obvious at all. The MongoDB docs tell you what it’s good at, without emphasizing what it’s not good at. That’s natural. All projects do that. But as a result, it took us about six months, a lot of user complaints, and a lot of investigation to figure out that we were using MongoDB the wrong way."
The film/actor/career example they wanted to do could have been solved pretty easily with a document that had an _id for an actors collection and a name... that's just the document paradigm... I'm a huge fan of both postgres and mongodb and I know that's not a popular opinion around HN. I just get tired of seeing clickbait headlines being cited as the criticism within is correct.