Here is a recent example: https://status.heroku.com/incidents/930
Source: I used to work at Heroku on the team which managed this process.
641 karma · joined January 30, 2010
Here is a recent example: https://status.heroku.com/incidents/930
Source: I used to work at Heroku on the team which managed this process.
At the top of the list is "Managing Humans: Biting and Humorous Tales of a Software Engineering Manager" by Michael Lopp[1], which was recommended to me by a manager who helped me get my start in engineering management. This book touches on a lot of the nuances in dealing with people and, as an introvert, I found this really helpful. The same author blogs under "Rands in Repose[2]" which has much of the content from the aforementioned book available for free.
While in the people category you'll also get a lot of recommendations for "Drive!" by Daniel Pink[2], which is a book about intrinsic motivators (autonomy, mastery, purpose) and how they are more important and effective than extrinsic motivators (e.g. money), particularly for knowledge workers. My personal advice, however, is to watch his TED talk[3] which is a great summary of basically the entire book. In this same category I could also recommend "The Great Jackass Fallacy" by Harry Levinson[5].
Now on the wall between people management and engineering/project management is "Slack" by Tom DeMarco[6], which is about how organizations and managers tend to run their staff at 100% capacity. As the book points out, however, this is a good way to not only burn people out, but it also sends response times through the roof (from queuing theory), and stifles change ("too busy to improve"). You can read this one on a plane. For some shameless self promotion, I've also written a tiny blog post relating Slack and the need for upkeep (software operations and maintenance)[7].
Next, fully in engineering/project management, I have to recommend "Waltzing with Bears" by Tom DeMarco and Anthony Lister[8], which is specifically about managing risk on software projects. The authors highlight the common practice of project/engineering managers communicating their "nano date", which they point out is typically the lowest point on the uncertainty curve. In other words, the project has the lowest possible chance of shipping by this date when you look at the possible timeline as a probability distribution. This book changed the way I talk about projects and the way I manage my team's various risks and I have been more successful as a result.
One final recommendation I'll make, since you're in the midst of a transition, is "The First 90 Days" by Michael Watkins[9]. It's a wonderful book that outlines how and why one should develop a transition plan in order to hit the ground running - and in the right direction. For my last engineering management opportunity, developing a preliminary 90 day plan as part of a "starter project," was a major factor in being given the job.
I believe that a subset of these will give you a great start. After that, you should read on the areas you feel the need for the most amount of help with or the areas that interest you. If you are avidly interested in project management, for example, you should read books on various methodologies, particularly the one that you or your organization practice.
[1]: http://www.amazon.com/Managing-Humans-Humorous-Software-Engi...
[2]: http://randsinrepose.com/
[3]: http://www.amazon.com/Drive-Surprising-Truth-About-Motivates...
[4]: http://www.ted.com/talks/dan_pink_on_motivation?language=en
[5]: http://www.amazon.com/Great-Jackass-Fallacy-Harry-Levinson/d...
[6]: http://www.amazon.com/Slack-Getting-Burnout-Busywork-Efficie...
[7]: http://www.charleshooper.net/blog/on-slack-and-upkeep/
[8]: http://www.amazon.com/Waltzing-Bears-Managing-Software-Proje...
[9]: http://www.amazon.com/The-First-90-Days-Strategies/dp/159139...
But on a general note: Hire an agency and spend your time growing your business in other areas and stop doing your own PR.
A good agency will help you create the proper messaging to use from your business' strategy, make sure that the announcement newsworthy, and, frankly, likely already has rapport with journalists and knows what they need to write a good story -- this will make them more effective and more efficient than you at pitching your announcement.
The opportunity cost on the pitching alone is insane so, seriously, don't do your own PR.
Generally, we'll start with new work on the left side of the board and completed work at the right side of the board; this roughly resembles a kanban board. The standard columns are:
Ready/Next (backlog) -> Doing -> Done
* Ready/Next are the top items from the backlog (usually a separate Trello board just so only active items are on the primary board) that are next in the queue
* Doing is work-in-progress
* Done is completed work (of course :))
Some teams also use additional columns for:
* Blocked - Work that is blocked on something else. In planning meetings and standups these are called out so we can unblock the items as quickly as possible
* Shepherding - Work that is mostly coordinating cross-team efforts. These items generally don't take up alot of active cycles of the "Shepherd" but they are an additional context switch throughout their work
* Interrupts - Usually this is called something else, but the gist is that some teams track operational items separately. For example, if support escalates a support ticket to an engineering team, the trello card referencing the ticket and any troubleshooting info will end up in one of these columns
As for ensuring that the Trello boards are up-to-date, many teams have standups and walk through their Trello board and confirm that it's consistent with reality.
* HipChat (sync and async chat with a variety of ChatOps functionality)
* Documentation: Google Drive for non-technical documentaton that might need feedback and some dynamic spreadsheets backed with dataclips: https://postgres.heroku.com/blog/past/2012/1/31/simple_data_...
* Video conferencing: Every single meeting has a corresponding Google Hangout. For some meetings we might use Fuze
* DCVS: git. Our repos are hosted on Github and we use all the usual stuff there: Pull Requests, Issues, in-line commenting, etc
* Project/task management: Trello trello trello - If it's not in Trello, it doesn't exist. This works great when you're widely distributed across geography and timezones. With the right workflow, we can at-a-glance know the status of all of our work-in-progress.
* Mailing lists! Every team has its mailing list and nearly every other thing of interest has its own mailing list. Interested in an upcoming project? There's a mailing list for that. Are you remote or based out of the SF bay area? There's a mailing list for that. Are you into Golang, functional programming, or want to chat about Linux? We have those covered too. Are you into biking or photography? Mailing lists!
P.S. - If you're interested in remote work, we're hiring! http://jobs.heroku.com/
FWIW, I did a free "starter project" for my current position and it was completely awesome. Worked with many teams, learned a few codebases, learned the metrics/logging systems, and learned the true size and formation of the production environment. A+++, would interview again.
On-Topic: Anything that good hackers would find interesting. That includes more than hacking and startups. If you had to reduce it to a sentence, the answer might be: anything that gratifies one's intellectual curiosity.
Off-Topic: Most stories about politics, or crime, or sports, unless they're evidence of some interesting new phenomenon. Videos of pratfalls or disasters, or cute animal pictures. If they'd cover it on TV news, it's probably off-topic.
The way I look at it, personally, is that we're not hardware providers. Instead, we provide some level of "ops as a service" -- partially through tooling and partially through highly skilled engineers.
Regarding their support, I've submitted tickets at sometimes very odd hours and gotten responses back within minutes.
What I really find amusing though is that you follow up your complaint of "completely absent during outage" with a reference to Amazon.
L1 cache reference..............................0.5ns
Branch mispredict.................................5ns
L2 cache reference................................7ns
Mutex lock/unlock................................25ns
Memory reference................................100ns
Compress 1K bytes with Zippy..................3,000ns
Send 2k bytes over 1Gbps network.............20,000ns
Read 1MB sequentially from memory...........250,000ns
Round trip within datacenter................500,000ns
Disk seek................................10,000,000ns
Read 1MB sequentially from disk..........20,000,000ns
Send packet CA->Netherlands->CA.........150,000,000ns
http://www.regexprn.com/2009/12/numbers-everyone-should-know...
http://engineering.twitter.com/2012/04/mysql-at-twitter.html
http://highscalability.com/blog/2011/12/19/how-twitter-store...
http://www.percona.com/live/mysql-conference-2012/sessions/g...
http://www.slideshare.net/yousukehara/introduction-of-twitte...
http://engineering.twitter.com/2010/05/introducing-flockdb.h...
http://www.infoq.com/news/2009/06/Twitter-Architecture
http://engineering.twitter.com/2010/05/introducing-flockdb.h...
http://www.charleshooper.net/blog/painless-instrumentation-o...
Implementation was easy. statsd is pretty simple to deploy and graphite wasn't too difficult either. To add statsd reporting to your code, it's essentially one line to create the statsd socket, another line of code to declare each timer or counter, and another one to increment. I think more time was spent determining what name to give each metric than it was implementing it in this project.
Now that I'm at dotCloud, I'm working with a much larger distributed system and we use it here also. We liked it enough to build some statsd hooks onto our RPC layer we use for just about everything. Now every time a component makes a remote procedure call, a counter for that call is incremented and the response time is sent to statsd. It's been very useful for troubleshooting odd behaviors and correlating events across the platform.
As people who work with complex distributed systems, we can't know exactly what they're doing. We'll think we know, and sometimes we'll be close. Other times we'll think we know, and then we'll wake up at 2AM because something failed horribly. By being able to monitor the system's behavior (sometimes in gross detail), we can get a little closer to knowing what's really going on.
More node.js
More C
Erlang (well, functional programming in general)
Calculus
Distributed Finite State Machine implementations
Public speaking (need to practice more than anything)
Reading/writing/speaking multiple languages
The blog post is supposed to be about "keeping your website up through catastrophic events" and the dominant theme seems to be "invest more heavily on Amazon." IMHO, the exact opposite needs to happen.
Sure, I understand that being in multiple regions means you supposedly have very autonomous deployments (to include Amazon's API endpoints), but nobody can prove to us that each zone or even each region are totally separate.
I'm not saying that Amazon is being dishonest about their engineering - I simply believe that by being 100% reliant on a single vendor you fail to mitigate any systemic risk that is present. That risk can be technical risk or business risk, as engineering at this level isn't strictly a technical profession.
I've seen alot of these types of services popping up lately. How does your service differentiate itself from competitors such as Cloudyn or UptimeCloud?