Netflix Billing Migration to AWS
techblog.netflix.com
techblog.netflix.com
I feel like Apple's approach of utilizing multiple providers makes more sense (though they do this for uptime and redundancy.) Maybe I'm being a pessimist.
This Ars Technica article from February summarizes the multi-year effort: http://arstechnica.com/information-technology/2016/02/netfli...
Their billing system has now joined its siblings in living on AWS.
EDIT: yes. [1] [2]
[1] https://openconnect.netflix.com/ [2] http://www.datacenterdynamics.com/content-tracks/colo-cloud/...
https://www.youtube.com/watch?v=tbqcsHg-Q_o
very cool video.
[1] https://media.netflix.com/en/company-blog/completing-the-net...
I can't speak for others, but for me, the decision wasn't, "should I choose Netflix or Prime Video?" Instead, it was, "Can I find enough programming to satisfy me for less than I'm paying for cable?" The only way my answer was yes was with multiple streaming subscriptions.
So you're saying Amazon is going to risk millions of dollars so they can make a few more bucks on video streaming, which is like 3 levels down from their primary business?
Prove it. I had an argument with my last company about this very issue. If Amazon's primary business was video delivery, then that would make a lot of sense. But where does Amazon's primary revenue stream actually come from? That's right, it's AWS. Amazon may be an excellent retailer, but they spend just as much money as they make on the shipping and fulfillment side to get shit on your doorstep faster and cheaper than anyone else out there. Each person that spends money on Amazon can only really spend a few hundred dollars per year. But even a small company that's entirely hosted on AWS, like 70% of the companies I've worked for, pays Amazon thousands of dollars per month for hosting. There's definitely more people than companies, but shipping stuff to people costs a lot more money than Amazon needs to pay in order to have your stuff hosted by them. Plus they do a lot of R&D into making their own systems faster and more efficient, eventually passing savings down when they re-work their pricing tiers.
Basically, my argument is/was that the whole idea of Amazon stealing your IP to make a few extra bucks on whatever they happen to be doing is totally bunk. Amazon is really in the infrastructure business, and if you're a video startup...you're NOT. And take it from someone who watched a company try and fail to build a competing private cloud with no budget and a skeleton crew...it's stressful and not fun at all.
The main reason to use multiple providers is, as you said, uptime and redundancy...and "not putting all your eggs in one basket", so to speak. It's an engineering, not political, decision. My last company could have probably saved themselves by moving everything to AWS and shutting down their Level3 internet-backbone connectivity and direct fiber from the office to the datacenter (which is the same technology AWS is using anyway, except they have an actual cloud API and not just a pile of servers), but they were too busy conflating this engineering/performance decision with one that must be made for political reasons.
But there are a lot more people than businesses, and an order of magnitude more people than businesses who both need custom hosting infrastructure and choose AWS.
The first part is irrelevant, because almost none of those people will ever buy a service from AWS. The second part is false.
An order of a magnitude more businesses are utilizing AWS than 'people.' AWS doesn't exist for consumers or average people, it's for businesses. That was the whole point of its existence: the primary customers of AWS are businesses, and always will be. It's where the radical majority of their sales come from, businesses with greater than 10 employees; not from solo developers spending $37 per month, and absolutely not from your average person that doesn't know anything about cloud computing.
I don't know if the extra effort of using eg. Google and AWS is worth in order to "stay up when AWS goes down" - but it might be worth it to stay up if/when AWS cuts you off for some reasons (quite possibly due to human error, billing, a take-down notice or other legal dispute etc).
None of that helps if you might accidentally find yourself on the wrong side of the "war on terror" by publishing news - in such a case all your US funds and assets might be frozen, and you would need a non-US presence in order to stay up while sorting out the potential error. But I suppose it's no worse than being subject to other kinds of arbitrary censorship...
Netflix is a marquee AWS customer. The PR damage of Netflix even making noise about leaving AWS would be terrible for the service as it fights for market share against Azure and Google. Netflix will get their way.
Let me try. Earlier Amazon's primary business was selling books, then it "became" selling almost every object that can be legally sold, and now you are saying that it "is" AWS. What about tomorrow? Tomorrow, it easily may "become" selling videos too. With Amazon it's very much possible.
So, it's not just technical issue, it's a political/business issue too. Of course, as you have said, they must take into consideration the trade-off. If the trade-off is more like "killing yourself under the technical burden of setting up a good network" vs "potentially allowing/helping Amazon to take advantage of your hosted service on their AWS and thus to become a future competitor" then they may go to AWS and/or other cloud provider(s).
edit: allowing/helping
As far as this not being Amazon's primary business, well, they are a lot bigger and diverse. But let's use a Wal-Mart analogy. Wal-Mart is so powerful (maybe not anymore, because of Amazon) that if it doesn't like your wholesale pricing to them, they can move your shelf space and practically destroy your business. They have the leverage in those relationships. Now, with AWS being the #1 cloud provider you have a similar lock-in, at the very least your switching cost would be quite high.
So, let's just for example sake say Amazon make a decision to give their own video streaming priority on their own cloud. That's not breaking net neutrality laws, because...it's their servers, this has nothing to do with telecom. So now Amazon Prime Video, streams 4K at a much better rate than Netflix (and Netflix would likely never know.) Then, perhaps another assumption, AWS decides to increase their pricing tier for their media servers for streaming. So suddenly they have squeezed you in two ways. Quality of Service and pricing.
Amazon could easily do a calculation to compare themselves to Azure and GCS to tell what is the proper amount they could get away with and still make it more expensive just to migrate to a competitor. You're locked-in, and just to make up for your switching, it would cost, let's say 1 years worth of AWS service. Hard to explain your sudden blip in earnings to short minded investors.
Anyhow, all I am trying to say is...it would make more sense, to me, from a business stand point to diversify across multiple providers instead of being all in on AWS. I never said you had to have an on-premise set-up. It would be much different if they were a small start-up, but Netflix is not. It can at times take up more internet traffic than torrenting.
I will point out though that services like Docker are making cloud lock-ins harder to do, but they don't solve reliance on APIs or just proprietary offerings. There's a reason why Google, Microsoft and Amazon all underline, push and constantly improve those offerings as a form of lock-in.
It is possible Netflix has some sort of pricing agreement with Amazon that locks in a rate for x amount of years. But either way, when you're as large as Netflix, I think diversification is the better long term strategy.
Looking at Amazon's Q1 2016 Financial Results [0], page 8, we see that net sales of non-AWS amounts to ~$26B, and AWS is ~$2.5B. From the Segment Highlights section (same page), AWS sales is 9% of total sales.
AWS beats non-AWS income only because there were losses in international; ignoring international, it's a $16m difference.
Looking at page 14, Media sales in North America and International sum to $5.6B. Media sales alone is double AWS's sales of $2.5B (page 13). Profit margins are way higher for AWS, so there's still a lot of room for a larger income difference between the two segments.
Based on this, someone doing media streaming as their primary business needs to be aware of who they are in bed with infrastructure-wise, but I agree that it doesn't mean that AWS isn't worth using just because Amazon is in the same market/area. Media sales are a big portion of Amazon's non-AWS sales, and being digital are most likely fewer headaches than physical goods; but not so big of a portion that one needs to worry about Amazon being the 800lb gorilla that would need to be contended with. It's more likely that Amazon is a threat to Netflix on licensing deals and content offerings (IMO, Netflix is slightly stronger, but neither are great, in content offering).
Is Amazon going to pull the rug out from under Netflix? Not if they want anti-trust attention. Is Amazon going to disallow the Netflix app to run on their devices? I suspect no, since they seem to want to create a platform rather than a walled garden when it comes to devices and media. Is Amazon going to undermine the trust in AWS by using AWS customer's data, or metadata about their customers, for their own gain? Probably not. If I'm going to put money on who's going to do "the right thing", for businesses and tech as a whole, it's going to be on the likes of Amazon rather than, say, Oracle.
Looking closer at (just) this (document), the growth numbers are such that it could go either way as to which revenue stream will eventually be the major contributor to Amazon revenue streams. That's a significant reduction in losses (-60%) over Q1 2015, so International turning around over the next year or so could keep AWS and non-AWS income neck and neck.
[0] http://phx.corporate-ir.net/External.File?item=UGFyZW50SUQ9N...
This was directly addressed in a 2011 talk[1] by a director at Netflix:
> Competition drives cost down
> [...]
> We would prefer to be an insignificant customer in a giant cloud.
[1] https://justinmk.github.io/2016/03/17/selfish-cooperation.ht...
The blog post covers a lot of the high level stuff really well, but I'm interested to learn whether you experienced any issues along the way, what they were, how you dealt with them.
In CloudFlare's case our migration was made more complex by also adding PayPal, and changing our processing gateway. Both of which created risk that we've had to work hard to understand and mitigate, i.e. how different gateways may, with the same card, return different results, etc.
Happy to chat offline too if you want to go into specifics.
[0] http://killbill.io/ [1] http://docs.killbill.io/0.16/migration_guide.html
I always just set up master/master replication with mysql. You can get free distributed reads that way if you architect your application right.
Ah, that's probably it. ~2004/2005 InnoDB wasn't nearly as popular as it is today, IIRC (but it was there, and was used).
In any case, multiple master with failover setups always worked well in my experience, as long as you take care to track the replication state.
MySQL's master/master has a number of complication and problems: 1. data loss due to async nature of replication, 2. update conflict on same data on multiple masters, 3. two masters mean two IP so all the clients need to know how to fail over to different IP, 4. complication in adding or removing master.
With DRBD, the disks are mirrored so that's no chance of data loss once a transaction is committed. There's only one master so no complicate conflict resolution. Linux HA's virtual IP means the standby will take over the primary IP so the clients don't need to know there's a server failover. Adding or removing standby is easy. DRBD will sync the disks automatically, no downtime on the primary.
That depends on what you mean by async. The replication itself is synchronous (statements cannot happen out of order), it's just not lockstep with disk writes and commits. I think it's more illustrative to say it's delayed.
> 2. update conflict on same data on multiple masters, 3. two masters mean two IP so all the clients need to know how to fail over to different IP
I'm not referring to multiple live masters, I'm referring to a set of servers where each is both master and slave to the other, and one live master server which gets the HA IP address. In that respect, there is no difference to a DRBD replicated setup. Clients just use the HA IP.
> 4. complication in adding or removing master.
I never found it that complicated, and I deployed it at least 5-6 times for multiple companies, and even with slightly different topologies (master/master where each master had an additional dedicated slave reserved for intensive read-only queries). What did you find complicated about it?
[1] http://dev.mysql.com/doc/refman/5.7/en/replication.html
4. For complication, you need to find out and configure the binary log position. Also because replication is on the binary log, if you ever truncate the log, you can't simply add a brand new master. You have to do a backup on the primary and restore to the new master, and then set up the log position to just prior to backup. Just lots of extra complication.
That's what I was talking about. It's just a matter of what aspect of it you are talking about, but I'll give you that it's in their official documentation, so there's no point in me pressing the issue.
> Committed data can be lost if master's disk is destroyed before a slave replicates it.
That is true. It's a trade-off you can make for slightly different CAP assurances, or nuances in the failure states at least (mostly in what you might expect to do in a split brain scenario).
> For complication, you need to find out and configure the binary log position.
Your backups should be logging the binary log position as well (--dump-slave or --maser-data). If they aren't, you aren't doing yourself any favors.
> Also because replication is on the binary log, if you ever truncate the log, you can't simply add a brand new master. You have to do a backup on the primary and restore to the new master, and then set up the log position to just prior to backup. Just lots of extra complication.
If you have to do another backup because it was truncated recently, you aren't much worse off than doing backups with DRBD replication (which even if you have a slave configured and do backups off that, you can truncate logs and need to backup from the master then as well). The downside is that you may not want to immediately do a backup of the master due to reasons of load, which will leave you without a failover for a short while. Whether an extra queryable resource available is worth that is up the the architect.
I remember a Percona training I was at a few years back there were a few more clustering options available I hadn't played with (and still haven't). Percona XtraDB Cluster was one, and it's supposed to support synchronous master/master replication. That might be the best of both worlds, if it lives up to its billing.
High availability is when you can afford to lose one or more nodes, and there is no service interruption. Moreover, the processing and I/O are symmetric and performance scales linearly by adding more nodes.
Edit: My guess would be that they still keep galera in mind, but since they didn't shared they why, one could only guess. And Transaction Wraparound maybe.
If you're doing things with replication logs, mysql replication logs are a heck of a lot simpler, and there's more tooling for them in the OSS world
That said, I do use postgres a lot :)
MySQL is extremely buggy and silently corrupts data. It also does not enforce explicitly requested referential integrity, a core mandate of a relational database management system.
https://www.youtube.com/watch?v=emgJtr9tIME
MySQL lacks adequate authentication model (like OS authentication or SmartCard authentication in Oracle).
Auth is a pretty aside argument. Those sorts of auth are not a particularly common use case. All sorts of auth can go in front rather than be built in, and if those don't apply, sure, might affect your DB choice. Not inherently a reason to not use MySQL though.
https://aws.amazon.com/blogs/aws/cross-region-read-replicas-...
https://aws.amazon.com/about-aws/whats-new/2016/06/amazon-rd...
I am interested in non-read replica replication...
MySQL is so completely riddled with bugs and utterly baffling behavior — it is the PHP of relational DBs. What I remember off the top of my head includes attempting to subtract datetimes causing the punctuation to be removed from the datetimes and the resulting "integers" to be subtracted, FK integrity being violated in a number of easy-to-hit corner cases, the "utf8" encoding not being able to encode UTF-8 (and the default being latin1…), GROUP BY allowing obviously (i.e., catchable to the parser) broken queries, `SELECT * FROM table` on <10k row tables taking minutes in some cases, the SQL dialect swapping the words "key" and "index" inappropriately; Read [1] if you want more.
Thus far the only other tool I've seen come close to this level of insanity is MongoDB. The thing about a tool so willfully discards any sort of reasoned approach to its topic area is that the people who use it — who inevitably are not well versed in how a relational database works "in theory" — cannot derive from its behavior its rules, because its behavior is irrational, bordering on psychotic, and the people using it tweak query after query while having no understanding of why one query might work better until some abomination that someone usually works is crafted; `-- Don't touch`. A good tool — I believe, somewhere deep inside me — will teach the novice user. MySQL will not; it will drive you insane.
I'm not disagreeing with your arguments (quite the opposite in fact), but that might be an important point to use MySQL over Postgres.
A large company describes their very real efforts and shares their experience and knowledge, and the first comment is "Why not use <preferred thing> instead?".
Postgres is usually the first choice for applications like this so knowing why they chose something else over that could influence others that need to make a similar decision.
Netflix has super smart people. They arent just picking random tools off the shelf and implementing things without constraints for the sake of it.
Do you have any data to substantiate this claim?
PostgreSQL was indeed a very attractive option, but we wanted to keep a path to Aurora open. When we were working on the migration, Aurora was still in beta, so instead of going to Aurora directly, we decided run our own MySQL instances on EC2.
Which security models does "Aurora" offer, for instance, are you able to use OS authentication like in Oracle? What about PKI on SmartCards?
Oracle has a nasty licensing model where they charge you per core regardless of if that core is a physical one or not (hyper-threading). While I was there, it suites told all the engineering managers that Oracle was out and the going forward solution was Microsoft SQL which, as I understand, has more relaxed licensing model.
Another thing I'm wondering about is I would figure Netflix to be big enough to have SAN storage. Just about every large company I worked at always used SAN replication technologies instead of open source stuff. And it's not a debate about open source solution vs. commercial. It's more about support. Large companies want a throat to grab when things break bad.
Sorry, should have been more clear.
As a follow up. Many projects I was involved with did use Oracle replication but it was a rolling log file type that was purposely delayed to account for mistakes. Rolling replication happen across geographic locations while SAN replication dealt with the hot fail-over situations locally.
That was a sad day.
That's why you then switch to their T5 or M-series with up to 3,072 processors, and then Oracle comparatively charges you peanuts, since those systems have relatively few physical sockets. And if you know SPARC, and you know Solaris, you can squeeze some serious savings. The problem, it seems to me, is that most system administrators and managers today are neither familiar with SPARC nor with Solaris, so they end up paying more in the long term in licensing and maintenance for running something else like MySQL on GNU/Linux.
But that's their problem, not the guy's who knows SPARC and Solaris and how to save money, isn't it?
Large companies have an enterprise wide license of the Oracle database, so they have thousands of instances (worked at a place like that, that's how I know), and at that point, with that kind of scale, Oracle becomes dirt cheap, considering it delivers the only database I have seen that can withstand the kind of abuse a large institution can throw at a relational database management system.
source: I work for Amazon Web Services.
In fact, I can pretty much guarantee that at the first opportunity where the lawyers agree it is a usable hole, they'll try to kill netflix through denying it service. Taking out netflix for a week or two while the engineers rebuild the backends with a different provider would be excellent for amazon's video division.
Considering that Cassandra is not ACID compliant with her "eventual consistency", and that MySQL is notorious for corrupting data and not functioning correctly, I am compelled to wonder just what kind of people work at Netflix. And who gets the idea to go to AWS and pay the full virtualization on Linux performance penalty?
Now, I've done Oracle engineering at some very large databases (hundreds of millions of rows, OLTP and DWH), and I know that Oracle is a smoking fast database when the right people develop on it. Also makes me wonder what kind of code they had running, and what kind of people selected it, when they managed to gum up what is essentially the Bugatti Veyron of databases.
Given this information from "Netflix" I won't be considering them as a potential employer any time soon. It has to be a mess over there.
I'll admit I'm confused about picking Cassandra as well, but not for the same reasons you are. They're only storing subscriber data (billing address, subscription type, etc). That data is going to remain static for months at a time. When changed the only potential problem that could occur is the billing process using old data, but I'm guessing their system is smart enough to try again in thirty minutes.
Oracle may be fast, but it's also expensive. This is billing, which means batched jobs running in the background- they don't care how long each individual transaction takes, and I'm positive it's going to be cheaper to roll up more servers to compensate than it is to pay Oracle's licensing fees.
Do you have any sources for the corruption comparison between Aurora vs InnoDB?
So? That's capitalism: one gets what one pays for, and Oracle is not just fast, it can do a lot, and it can be configured to be paranoid about protecting data, and it has clustering technology meant for scaling, RAC.
Truth be told they could have picked PostreSQL and it would have still been a better solution than Cassandra and MySQL.
Sorry about the c/p but I don't see a permalink on lapitopi's comment.
lapitopi 1 day ago I work on the Netflix Billing Team. PostgreSQL was indeed a very attractive option, but we wanted to keep a path to Aurora open. When we were working on the migration, Aurora was still in beta, so instead of going to Aurora directly, we decided run our own MySQL instances on EC2.
edit: begun/began