Netflix Blog: Tips for High Availability
medium.com
medium.com
Not trying to be a pedant, just curious.
Interesting that you shift US-East to the EU. Any more color on why that direction, versus say, US-West? Like maybe less common components in AWS? At face value, it seems like an odd choice.
I can only imagine what debugging misbehaving clients looks like, it's got to be north of 1,000 hardware/software profiles.
I thought of some questions:
How is the dev/test env kept up to date for integration/canary tests? When a new app/service is pushed to production does its build/image (AMI?) become the new dev/test env base image for everyone?
Do engineers decide which metrics are tracked by Kayenta? Does ops? How does this come together?
What about post deployment service monitoring / alerting? When is the dev team off the pager hook?
Do you assume that a successful deployment in us-west means it will be successful in us-east or is there an [integration] test done per region?
Just curious.
https://medium.com/netflix-techblog/sps-the-pulse-of-netflix...
It is an SPS graph (as indicated by the title). Spinnaker displays the graph on that page to give engineers a visualization for what SPS looks like during their deployment windows. If your service is critical for streaming you'll have a preference for deploying during lower traffic hours to minimize potential impact. Fortunately as the post mentions it is very regular, and has different praks in different aws regions, which allows regional deployments to be staggered.