Homegrown DevOps Tools at Stack Exchange
blog.serverfault.com
blog.serverfault.com
[1] http://everythingsysadmin.com/2013/09/the-team-im-on-at-stac...
If you have passion, you will have plenty to offer.
If we rephrase the blog post as "we could not find any good tools in the Windows devops space so we wrote them" and add it to the departure of the only CEO willing to dance on stage chanting developers developers developers and Windows is not an ecosystem but a hub with a few brave outlying satellites.
I am impressed by the stacke change folks and their story and skills but it feels like amazing stories of software skill written for one company and never released in the open - it just leaves no legacy
MiniProfiler is my favorite one, and now exists in .NET, Ruby and Node.js due to my insanely smart co-workers.
SE has a very high reputation as far as I can tell so its not a quality issue - it is in this case a problem of quantity.
As someone commented the idea is to focus on Azure and mobile - but that is no substitute for lots and lots of developers in a culture of releasing stuff outside their immediate company and so driving each other to new heights
ASP.NET MVC has definitely been a big factor with this on the web development side.
For instance, a key problem I had in a prior gig was that I needed to automatically log into a Windows machine, run a job, and then log out. Pretty bog standard; didn't need careful error recovery or anything particularly sophisticated.
In Linux, you configure your SSH keys, then ssh automatedjob@server -c "./run-my-thing", and that's that. I literally could not find any identical analog in Windows besides telnet (if anyone knows of a solution here that approximates the Linux one for simplicity, I'd love to know about it). Today I'd probably just requisition a copy of a Windows SSH server and be done with the sorry mess. Better yet, throw Windows out and go full Linux. ;-)
Also, iirc, using the psexec interface from Linux in a Linux->Windows connection does not work (a wrinkle in the original story is that it was a Linux->Windows connection).
Supposing that a keyfile is copied, the multiplicity of ssh files limits the damage.
Now, I was pretty sure you could use WMI to do this, and looking around I found this -
http://4sysops.com/archives/three-ways-to-run-remote-windows...
and this
http://blog.commandlinekungfu.com/2009/05/episode-31-remote-...
Remote PowerShell is probably 'the way' Microsoft will be pushing now
http://msdn.microsoft.com/en-us/library/windows/desktop/ee70...
What was happening is that there was a job controller running on Linux that needed to contact the Windows machine and boot the job into action. One solution would have been to install a Jenkins slave instance, which turned out to be the solution eventually adopted.
[1] It's been a few years since I studied this, but it was a known issue in '08/09 or so.
We've developed approach using Chef that spans Linux and Windows boxes.
To log in and run some script on Windows boxes, we use WinRM, which is integrated with Chef knife tool and works almost as well as SSH.
I seem to remember doing clever things with WSI for just such a thing curiously I am pretty sure there will be examples of both on StackOverflow ;-)
Edit : oh yes WMI and psexec - amazing how quickly things drop out of your brain.
He is also priming out infrastructure for Desired State Configuration (configuration management (like Puppet/Chef) for Windows).
I'm glad its working for you, I use stackexchange daily and am generally quite happy with it!
That I was hired as a Windows specialist was so that we can go deeper on the OS side and the PowerShell side, just as we have a Linux expert to go deeper on the Linux side. Our sysadmin team was just more tilted with experience on the Linux side (though you wouldn't know it - as almost everyone I work with would qualify as a senior admin in any Windows shop in the world).
Thank you for your response.
I am glad to see that there are comparable capabilities in the modern Windows world and will dig into the WMI side of things next time Windows admin tasks come up.
* Set the firewall to only allow telnet access on a private subnet
* Set the vpn connection to use that subnet
* login from any computer that supports telnet and PPTP.
Windows will die, to make way for Azure and Windows Phone.
The Microsoft development ecosystem is strong (C# in top 10 of TIOBE index). It has open source released by Microsoft (ASP.NET MVC), 3rd parties (Mono by Xamarin), and the community (opensourcewindows.org). Microsoft includes open source libraries (jQuery) in their projects.
StackExchange's software will leave a legacy.
From the article: "In order to create more, and open source it we need help. So we are looking a full time developer with ops experience to join our SRE team."
As far as what we do at Stack, a lot of our code gets released. Not the core q&a engine, but our logging framework (StackExchange.Exceptional), a profiler (MiniProfiler), and some other stuff. On the sysadmin side, we are active in developing Desired State Configuration modules and contribute those to the PowerShell.org DSC repo on GitHub. Part of the job description for our SRE developer includes the fact that much of what we develop in house will be targeted to be open sourced.
LogStash, ElasticSearch and Kibana are a great open source stack for log management.
StatsD and Graphite are nice tools for metric tracking and visualization.
There are lots of open source dashboard offerings which combined with a bit of scripting can get you far.
You are also spoiled for choice with SAAS monitoring stuff such as NewRelic and Server Density, even if the OP isn't a fan of cloud based tools.
When I looked at StatsD and Graphite last time I didn't really see an API. I really like the model of data being queryable and returning nice serialized format like json (like OpenTSDB does). I'm also not that fond of the "many files" model and the automatic data summarization as it ages (it does save space, but makes forecasting difficult as it can skew data).
Graphite has a very simple and powerful JSON API [1]. Any graph URL can include &format=json and you'll get back the raw JSON values for the datapoints. I haven't used OpenTSDB yet - I'm curious how its API is better.
You get to choose the levels of summarization. If you want to keep 1 second intervals for a year, you're welcome to do so.
[1] http://graphite.readthedocs.org/en/0.9.12/render_api.html
Part of the reason for the open position linked to in the post ( http://careers.stackoverflow.com/jobs/39983/developer-site-r... ) is to make it so we have more manpower to get this stuff open sourced.
One of our major problems with existing monitoring and management systems is the lack of good APIs. We are a shop of developers and sysadmins who all understand the real management systems need to be composable. The system needs to understand that it won't solve every case out of the box and expose hooks into management and functionality, allowing us to tie disparate systems together and enhance their coverage. I'd rather take a bunch of existing products and put some cool dashboards on top, but most enterprise solutions (and some of the open source ones) don't offer a decent API to work with.
Nagios check plugins don't have an API per se, but they have a very simple standard for exit codes to be interpreted by the Orchestration component.
Nagios reporting/alerting plugins primarily use RRD, so you always have that API to do interesting things like trend analysis.
(One theory I have is that there is a such a large and growing culture of technologists that are thinking "integration tool" but only know how to say "HTTP API")
This makes that page much easier to read. Especially the headings, which are all mashed together for some reason...
#content {
margin-left: auto;
margin-right: auto;
float: none;
width: 40em;
font-size: 15pt;
line-height: 1.4em;
}
h1 {
font-size: 180%;
line-height: 1.2em;
margin-top: 2em;
margin-bottom: 0.5em;
}
h2 {
font-size: 160%;
line-height: 1.2em;
margin-top: 2em;
margin-bottom: 0.5em;
}
#wrap {
width: 100%;
}Basically:
- clean up code,
- make sure the infrastructure is sufficient,
- help with marketing and adoption,
- write documentation
A Kickstarter project, for instance.
We have definitely outgrown Orion, and a lot of stuff in Orion is very rough, sloppy, and not well integrated.
We don't really like the idea of cloud hosted monitoring (which is a lot of what more modern monitoring systems are). Alternatives also seem very expensive.
So if we are going to make a investment (cash or labor) I would rather we get a system that fulfills all of our needs (fit) and share it with everyone.
The logging tool is the only one I'm not sure about, if you're parsing 1200 events a second. I'm not familiar enough with Orion to understand why it was insufficient.
Would you say these in-house tools are primarily about integrating / presenting the data, or are there custom parts doing more heavy lifting too?
A lot of these tools are about integrating / presenting but that isn't entirely the case. Status does its own polling of things like Redis and SQL. Realog parses the data, structures it in Redis etc, and the patching dashboard can kick of updates.
One thing Nagios handles well, although it isn't exactly well polished, is distribution, scheduling, and aggregation of the polling. Also, anecdotally, "nobody" seems to know how easy it is to do simple-intermediate monitoring of MS-SQL in Nagios. I happen to be using this, at the moment:
https://github.com/scot0357/check_mssql_collection
edit: And thank you for the concise explanation of Hiera. I had entirely ignored it as just another add-on as Puppet "goes enterprise"
We use Puppet to manage our Linux infrastructure and are testing Desired State Configuration for our Windows systems. Puppet and DSC solve a different problem than monitoring.
Our tools span the gamut of aggregating and presenting data, to doing "heavy lifting" of things like installing patches, managing our load balancers, and removing bad query plans from our database servers.
Our in-house tooling fills in gaps where existing tooling wasn't responsive enough or made it difficult to deal with certain edge cases. General purpose tools like Orion satisfy 80% of our use cases, but we are fanatical about performance and functionality so we want to fill in that additional 20%. In fact, the majority of our Orion monitors are custom script monitors, which we have to create ourselves as it is.
If what we do can work for others (I've worked in several environments with Orion and had the same issues in each environment), that's a bonus.
We've found we are building a significant number of these projects and that's why we are looking for a dedicated developer for our team.
These are definitely distinct from monitoring, but monitoring configuration should be populated from the same configuration store as Puppet. (Sometimes Puppet is the configuration store.)
I've never evaluated Orion in particular, but it's slightly puzzling if you're creating a lot of custom monitoring scripts. In the low-touch Nagios deployments I've seen, this is often because someone didn't understand good places to use macros, and centralize more of the parameters.
Configuration systems have historically looked at config as bits-on-a-disk.
Supervisor systems look at bits-in-memory.
And they emerged and evolved independently. So there's pain points and impedance mismatches regardless of which one you start with.
What's needed is a system that sees that configuration and supervision are the same problem: you have a directed graph of what a system can look like, plus a compare-and-repair mechanism to drag the system to that state frequently.
I've personally looked at Chef, Puppet and Cfengine, none of which really does all of them the way I would like.
http://chester.id.au/2012/06/27/a-not-sobrief-aside-on-reign...
BTW, we made heavy use of setting and logging http headers too. One trick I liked was capturing performance timing metrics as a request was processed and stuffing it into a response header as the response went out. We then logged the response headers, which gave us the ability to report on the performance metrics. We also had a debug mode in the app on the browser side so we could see the performance metrics from the headers there too.
~70 events per second doesn't sound like much to capture and aggregate. How much of this parsing did you need to perform in real time? Creating a unique token to pair requests/responses shouldn't add much overhead at all.
- No real-time parsing; it's all nightly batch processing after devops rotates the Apache server logs to a storage volume. The logs sit there for a while then get compressed and moved to offline tape archives.
- No DB storage of the logs; space was too expensive and the Oracle database we had couldn't have kept up. It was already heavily burdened with a completely separate usage statistics system that fed into user-facing reporting and billing, which had a much higher event rate, about 100x higher, than the http logs.
- We had unique tokens, but they identified a particular user session that tied together all of the user's http requests from login to logoff/abandonment, and which also tied into the Oracle-based statistics for that user, that user's organization, and the customer responsible for the user (often multi-organization). My reports had breakdowns for individual user experiences, session-level metrics, and user type/organization/customer/region/etc metrics.
- I don't recall how long the analysis took; it was between half an hour to two hours I think. A lot of that time was spent on disk I/O reading the logs. I had optimized the parsing, analysis, and results recording about as much as I could.
- This stuff was written in Perl, and ran on Solaris servers from that time era... probably not a lot more powerful than a handful of smartphones today, though they did have lots of cpus. I don't think traffic has grown much since I left the company (we had pretty full market penetration already) so it's likely those servers haven't been upgraded.
I think I have a good idea of how businesses (at a high level) have failed to understand Moore's law from 2000-present. I'm curious what those failures of understanding were like from 1985-2000.
We all know that technology has been advancing rapidly, but these specific anecdotes of organizations paying a million dollars just for the backing storage of a system that you can essentially get for free from Google now...
Actually, they're probably still paying over $100/GB. The whole datacenter was outsourced to Perot Systems in the mid-2000s, and the storage fees were astronomical. We calculated that Perot must pay a separate tech to stare at each individual hard drive with a replacement in-hand in case any errors were reported. At least, they could afford to do that with what we were paying them for storage.
2. System lock-in, I like to have the data and be able to query it as I see fit
3. If an event happens where the facility is cut off from the Internet you won't be able to tell what happened inside the facility (unless maybe there is a store and forward agent, but even then it is only useful after the event.
4. Latency monitoring (again maybe an agent can help) but if I see changes in response time I can't tell if that is the WAN or not.
Basically it is an additional perspective and a redundant level of monitoring. But I view it mostly as an up/down layer of monitoring. Extremely important, but simple and not really meant to give insight into the complexity of our system.
We use Puppet to deploy the client to our Linux boxes and for Windows we deploy the task and scripts with Group Policy (soon to be replaced by Desired State Configuration).
This information isn't something we need to poll often for (and on Windows, there are difficulties interacting with the Windows Update apis remotely). We add and replace servers often, so the client adds themselves to the dashboard, as well as updating their status.
We have integrated some reporting into Orion, but that was a side effect of not having the dashboard before.