Lessons from 10 Years of Amazon Web Services
allthingsdistributed.com
allthingsdistributed.com
https://www.usenix.org/legacy/event/lisa07/tech/full_papers/... (somewhat condensed)
Slides from a talk: http://mvdirona.com/jrh/talksAndPapers/JamesRH_AmazonDev.pdf
He writes some of the more practical and useful advice I've seen about how to run a successful business based on high-scale distributed systems.
One of his mnemonics that's stuck with me is the "Four Rs" of recovery-oriented computing: Restart, Reboot, Reimage, Replace (slide 6 in PDF). A human shouldn't be engaged to troubleshoot a problem until the platform has first tried in successing to restart the software, reboot the OS, re-image the machine, and finally replace the hardware entirely. After these auto-recovery steps have failed, only then is it time to engage a human. He describes the connection between these techniques and significantly lower operations costs.
This is why Amazon beat Google in the "as a Service" space. Google provided a framework with AppEngine, but customers wanted the flexibility of the primitives offered by AWS.
The narrative shifted away from PaaS for a number of reasons, but the value prop is still present.
(Disclaimer: work on BigQuery, a DBaaS, which is, in my opinion, is clearly the future)
Plenty of people in the web hosting industry offered competitive x86 linux instances or dedicated servers, way prior to 2006.
I think AMZN deserves credit for delivering a price compelling option that delivered (1) elasticity (per hour billing) and (2) automation (API, available in minutes), among other things.
Offer IaaS first, the PaaS. Google had it backward.
While AWS does have a small vendor lock in too, in the end it's still recognizable , I mean it's doable for example to move an AWS architecture to open stack without a major code rewrite. If you built your web app on GAE and GWT however, your investment to move it out of Google will be much bigger.
Amen. We couldn't agree more: https://boxfuse.com/blog/no-ssh
In a perfect world production and development would be identical, all production issues would be easy to reproduce in development, and logging and error handling would be sufficient to diagnose any problem.
But what do you do when you have a problem that doesn't happen in a perfect world -- that you can't diagnose normally, doesn't occur in dev, shows nothing abnormal in logs, etc.? At some point you need the ability to inspect the actual thing. It's like diagnosing a problem with a bridge by doing experiments on a scale model replica. Doesn't always work.
http://techblog.netflix.com/2012/06/scalable-logging-and-tra...
https://github.com/Netflix/vector
Not terribly easy, but possible. I agree that "never have to SSH" is a bad target because there will always be edge cases, but for the majority of things it's possible to troubleshoot at the cluster level given immutable infrastructure; all the systems will be doing the exact same thing given an input. Reality shows that this is a bit idealistic, but if you start finding that e.g. A certain set of requests cause weird behavior, you can look at what makes those requests different, then set up some routing logic to send them to a debug farm (using something like https://github.com/Netflix/zuul ).
Instead of ssh vs. no-ssh debate, how about:
* Use SSH AllowGroups to limit logins to the development manager/lead or on-call engineer
* with a policy that for each time you have to call them beyond N times per week, you have to buy their team a round of drinks of their choosing?
Deal?