I'm impressed they were able to do this so quickly.
I'm impressed they were able to do this so quickly.
Something like that.
If it was malicious they would have made a bunch of them, not just one. I personally have seen many files with unreal amounts of whitespace at the end.
If time pressure is not a factor and the production server has the VStudio or WinDbg installed, attaching the debugger to the process can see the data related to the post. But the symbol file and source files might be needed; it's just more hassle to set up. For stressful situation, simple steps and simple tools are more useful. Whatever works.
Probably went like this: 'web server crashed. so did another. what page did they crash on? ok, let's take a look at the post on that page. what in the....'
> 10 minutes to roll out the fix
That seems very slow to me. 30% of their down time was because their deploy process is slow.
FWIW here's a write up on their process
Also I imagine that 10 minutes included dev and testing, not just the deployment part of "rolling out"
Maybe "deploy" means the "Deploy" section of this article: http://highscalability.com/blog/2014/7/21/stackoverflow-upda...
Seems to target only being able to "deploy 5 times a day". I guess maybe the build time is the limiting factor.
Its just interesting to me the implications of what folks optimize for and that this is considered fast. We have very minimal deploy testing and optimize to be able to revert quickly when there are problems because performance issues like this are very hard to predict. Probably means we create many smaller short hiccups though (that generally are not a full site crash).
- login to server
- make dump of all threads stack traces
- see that something like Regexp.match present in all stacks
- find function that called this regexp.