Netflix open sources resilience engineering library
github.com
github.com
(I just drank an entire bottle of wine, so please excuse any apparent lack of reasoning, typing ability, or general coherence in the above comment.)
I also agree that libraries are preferable to frameworks and Hystrix is in fact just a java library that can be used as little or as much as one wishes.
It purposefully tries to have a small number of dependencies so should be easy to pull in without significant impact.
@benjchristensen
The code written by someone who's just completed reading the GoF book feels remarkably similar...
Can you elaborate a bit?
The API has a hell of a lot more surface area but is trivial in complexity compared other systems at netflix, and therefor this library has some huge gaps in design.
The two biggest issues IMHO are putting the throttling/fallback handling at the outermost edge of an external service rather than at the lowest level (i.e. the actual rpc) and a very C/errno like method of handling errors.
I'm also quite unhappy with the API. It require creating boilerplate classes to implement the commands. Yuck. A bit of magic with annotations or code gen would've been much cleaner and much less prone to errors caused by programmer fatigue or boredom.
The concepts behind Hystrix are well known and written and spoken about by folks far better at communicating these topics than I such as Michael Nygard (http://pragprog.com/book/mnee/release-it) and John Allspaw (http://www.infoq.com/presentations/Anomaly-Detection-Fault-T...). Hystrix is a Java implementation of several different concepts and patterns in a manner that has worked well and been battle-tested at the scale Netflix operates.
Netflix has a large service oriented architecture and applications can communicate with dozens of different services (40+ isolation groups for dependent systems and 100+ unique commands are used by the Netflix API). Each incoming API call will on average touch 6-7 backend services.
Having a standardized implementation of fault and latency tolerance functionality that can be configured, monitored and relied upon for all dependencies has proven very valuable instead of each team and system reinventing the wheel and having different approaches to configuration, monitoring, alerting, etc - which is how we were before.
As for incremental backoff - I agree that this would be an interesting thing to pursue (which is why a plan is in place to allow different strategies to be applied for circuit breaker logic => https://github.com/Netflix/Hystrix/issues/9) but thus far the simple strategy taken by tripping a circuit completely for a short period of time has worked well.
This may be an artifact of the size of Netflix clusters and a smaller installation could possibly benefit more from incremental backoff - or perhaps this is functionality that could be a great win for us as well that we just haven't spent enough time on.
The following link shows a screen capture from the dashboard monitoring a particular backend service that was latent:
https://github.com/Netflix/Hystrix/wiki/images/ops-social-64... (from this page https://github.com/Netflix/Hystrix/wiki/Operations)
Note how it shows 76 circuits open (tripped) and 158 closed across the cluster of 234 servers.
When we are running clusters of 200-1200 instances the "incremental backoff" naturally occurs as circuits are tripping and closing independently on different servers across the fleet. In essence the incremental backoff is done at a cluster level rather than within a single instance and naturally reduces the throughput to what the failing dependency can handle.
Feel free to send questions or requests at https://github.com/Netflix/Hystrix/issues or to me @benjchristensen on Twitter.
It's more than just aborting a request. It also handles caching with different fallback modes, collapsing multiple requests, monitoring and most importantly it provides a consistent model for handling service integration.
Probably not the most innovative thing in the world but definitely useful if you do have a large SOA system.
https://github.com/Netflix/Hystrix/wiki/How-it-Works It's more than just aborting a request. It also handles caching with different fallback modes, collapsing multiple requests, monitoring and most importantly it provides a consistent model for handling service integration. Probably not the most innovative thing in the world but definitely useful if you do have a large SOA system.
* maintains a semaphore of currently processing requests, and will immediately fail a request if the semaphore is full.
* tracks service latency and other statistics for you
* maintains a circuit breaker to immediately fail requests if the breaker is open (based on statistics)
* watches for health to be restored of the 3rd party service and allows requests to it again.
* built in request isolation via threadpools
* built in request collapsing and caching
Overall I like this cause it's a great demonstration of how to think about, and engineer for, failure in a distributed system.
Of course this is all assuming you have some sort of useful redundancy.
It could be part of the default library or perhaps a custom strategy for the circuit breaker (once I finish abstracting it so it can be customized via a plugin: https://github.com/Netflix/Hystrix/issues/9).
At the scale Netflix clusters operate they basically get this randomness already because circuits open/close independently on each server (no cluster state or decision making).
In this screenshot https://github.com/Netflix/Hystrix/wiki/images/ops-social-64... you'll see how in a cluster of 234 servers that about 1/3 of them are tripped and the rest are still letting traffic through.
Thus the cluster naturally levels out to how much traffic can be hitting the degraded backend as circuits flip open/closed in a rolling manner across the instances.
Also, doing this makes sense even when a dependency doesn't have a useful redundancy and must fail fast and return an error.
It is far better to fail fast and let the end client (such as a browser, iPad, PS3, XBox etc) retry and hopefully get to the 2/3s that are still able to respond rather than let the system queue (or going into server-side retry loops and DDOS the backend) and fail and not let anything through.
We prefer obviously to have valid fallbacks but many don't and in those cases that is what we do - fail fast (timeout, reject, short-circuit) on instances where it can't serve the request and let the clients retry which in a large cluster of hundreds of instances almost always get a different route through the instances of the API and backend dependencies.
@benjchristensen
also, huge thanks to you and your team (and your employer!) for releasing an amazing volume of production-quality open source projects this year.
People who don't know CS are doomed to poorly reinvent Lisp or Erlang again and again..))
But why not, if someone pays for it..
And the whole idea of using Java for serving media content, while there is a specialized, well-engineered solution, created especially for this purpose in the telecom world, is such a brilliant management decision.. In Java we trust.)