A State of Xen – Chaos Monkey and Cassandra
techblog.netflix.com
techblog.netflix.com
How does one go about writing a user-facing application with eventual consistency?
For more complex applications (think Facebook), there are useful consistency models other than strong consistency, with causal consistency being one of the most promising: http://queue.acm.org/detail.cfm?id=2610533
Likely if they have things that require stronger consistency they use another DB.
...seriously, though. All of our views are out of date by the time we see them, and all of the user inputs have to be dealt with for races and double-submits. Eventual consistency is just the patterns you already know, but bigger.
The same is not true for eventually consistent systems - checking against a version number can't save you unless all your updates are trivial, single-object ones: some of your updates may succeed while others fail - and there's generally no way to perform an all or nothing operation. In general, there's no one good mechanism for maintaining eventual consistency on nontrivial operations - you have to think (hard) about it on a case by case basis.
For example, using the rails TODO list, if you add a task to "buy milk" and then on another client view the list and don't see "buy milk" there (because consistency hasn't been reached), you might be inclined to think you forgot to save it or something. Then you might enter it again. If this is not idempotent (as is the case on the classic tutorial) you will end up with 2 "buy milk" tasks instead of just one.
Making the primary key "buy-milk" would do the job. Getting a consistent key requires a bit of UX but isn't actually that difficult.
buy-first-milk
buy-second-milkThis is exactly the type of case where I'd much rather have the app be explicit about it's actions: allow both to be added, and at most show a notice after synchronisation saying "I've noticed you've added two buy-milk; merge them or leave them separate?".
I've had nothing but bad experiences with apps that thinks they know best and try to merge records behind my back.
If you say, "Buy_first_Milk" and "Buy_Second_Milk" - and the order is reversed, and the "Buy_Second_Milk" is applied first, and the "Buy_first_milk" is applied twice then everything happens as desired.
Being explicit is what allows everything to work without confusion.
From a user interface perspective, if the user ever enters "Buy Milk" - and there is no existing milk, that is always, "Buy First Milk". If they go to a different client, and the milk order is there, and they enter "Buy Milk" - that then becomes "Buy Second Milk". But, if "Buy Milk" hasn't synced yet, they can explicitly flag it as "Second Milk Order".
I don't think it would be very useful in a TODO list situation though.
(At some point for efficiency reasons you might need to prune the transactions, but this often involves trade-offs like not allowing any new "lost" transactions older than date X.)
For example, for your viewing history you might not entirely care that it is 100% up to date and the last video you watched isn't available. And for the times you do care, most users will simply usually refresh the page (and since, in almost all cases eventual consistency means resolution in seconds rather than instantly, most users will have the correct info after a page refresh).
I can't find them right now, but Christos Kalantzis (@chriskalan) has a number of talks on using Cassandra at Netflix w.r.t eventual consistency.
Video: https://www.youtube.com/watch?v=lwIA8tsDXXE
But, it's important to note that Cassandra does recognize that sometimes you do need linearizable consistency. Cassandra provides lightweight transactions for this scenario: http://www.datastax.com/dev/blog/lightweight-transactions-in...
from: http://xenbits.xen.org/xsa/advisory-108.html
MITIGATION ==========
Running only PV guests will avoid this vulnerability.
Did amazon reboot all of it's VMs? or just the HVM VMs? why was neflix running on HVM VMs?
Why shouldn't Netflix be running on HVM?
HVM, at least in the past, had a bunch more code that the guest DomU interacts with vs. fully pv guests. This has security implications.
Now, my knowledge of HVM is a few years... or more like half a decade out of date, for example, I don't even know how to force a HVM guest to only use PV drivers (which would solve 90% of the problem.) and i know that more and more of this has moved into hardware, so it's possible that what was true five years ago is not true now, but... yeah, I don't let untrusted users on HVM guests for the same reason I don't let untrusted users use pygrub or load untrusted kernels directly.
http://www.brendangregg.com/blog/2014-05-07/what-color-is-yo...
Forcing an HVM guest to use only PV drivers sounds like PVH, which is coming in Xen 4.4:
PVH is actually PV on top of an HVM container and is a bit different. You can think of it as PV sitting on top of enough HVM bits to take advantage of the hardware extensions Intel and AMD have invested so heavily in while still being majority PV. This gives you the best of both worlds, including the remaining PV performance benefits related to interrupts and timers that PV drivers on HVM can't utilize.
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance-...
but on this page, they don't say that specifically, but mention the possibility in the PV on HVM section:
http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/virtualiz...
Perhaps/presumably it also depends on the AMI you're using.
PV drivers on HVM have been a thing for a while now, so your IO and network go through PV even with an HVM instance. SRIOV/"Enhanced Networking" is even better than PV networking drivers, so HVM has another win there if you have NICs that support it (The larger Amazon instances do)
PV is significantly slower with system calls than HVM on 64bit hardware, as the x86_64 spec removed two of the four CPU protection rings, one of which being where PV lived in 32bit architectures, below the guest, allowing it to 'intercept' these system calls. Now since it shares a ring with the guest, it cannot intercept these and you are left with doubling the amount of context switches that you were previously. Virtualization extensions from Intel and AMD allow HVM to bypass this and skip the context switch. Other hardware advantages like EPT also give HVM an edge when it comes to memory related performance.
About the only area where PV is still better from a performance standpoint is interrupts/timers
They wrote a blog post about it here: http://techblog.netflix.com/2011/07/netflix-simian-army.html
I don't think they have one that introduces network partitions, but inside a datacenter, network partitions are rare.
It's ... complicated.