AI agents invade observability: snake oil or the future of SRE?
monitoring2.substack.com
monitoring2.substack.com
Compare this to codebase AI, where much of the data you need lies in your codebase or repo. Even then, most of these coding tools aren't even close to automating meaningful coding tasks in practice, and while that doesn't mean they can't in the future, it's a long ways off!
Now in the ops world, there's little to no guarantee that you'll have relevant diagnostic data coming out of a system that you need to diagnose it. That weird way you're using kafka right now? The reason for it is told via oral tradition on the team. Runbooks? Oh, those things that we don't bother looking at since they're out of date? ...and so on.
The challenge here is in effective collection of quality data and context, not the AI models, and that's precisely what's so hard about operations engineering in the first place.
Not related to your main point, but I've introduced Github Copilot to my teams, and, surprisingly, two of our strongest developers reached out to me independently, and told me it's been a huge boost to their productivity, one in refactoring legacy code, and another in writing some non-trivial components. I thought the primary use would be as a crutch for less capable developers, so I was surprised by this.
As a middle-manager whose day job previously robbed me of the opportunity to write code, I've used ChatGPT 4o to write complex log queries on legacy systems that would have been nearly impossible for me to do otherwise (and would have taken a lot of effort from my teams) and to turn out small (but meaningful) tasks, including learning Android dev from scratch to unblock another group and other worthwhile things that keep my team from being distracted and able to deliver.
I guess there's a "no true Scotsman" fallacy hiding there, about what constitutes "meaningful coding tasks in practice", but to me, investing in these tools has been money well spent.
However, these same coding assistants lack so much! For example, I can't have a CSV co-located in my directory as a jupyter notebook file, then start prompting+coding without having also done a call to df.head to get those results burned into the notebook file. The CSV is sitting right there! These tools should be able to detect that kind of context, but they can't right now. That's the sort of thing I mean when we have a long way to go.
PS: I hate how many "agents" already have to run on systems. Especially when the prod stuff is core starved already. I can't tell you how many times I've found an agent (like crowdstrike) causing some strange cascade issue!
The article talks about moving past “copilots” and moving right to “agents”. There’s probably some semantics to decipher there, but we haven’t even gotten copilots to work well! At their core, they are essentially the same problem, but I feel a lot safer with a chatbot suggesting a mitigation than just going and performing it.
IME benchmarks, though valuable, don't fully reflect the real world, often only reflecting the easily quantifiable. The best way is to be able to quickly try out an agent to see how it performs on your work environment. Sort of like having a private test set you can try different agents on to see how they perform in the real world quickly.
Disclaimer: I'm building MinusX, a data science agent (github.com/minusxai/minusx)
I mean is the AI going to read your sourcecode, read all your slack messages for context, login to all your observability tools, run repeated queries, come up with a hypothesis, test it against prod? Then run a blameles retrospective, institute new logging, modify the relevant processes with PRs, and create new alerts to proactively catch the problem?
As an aside - this is garbage attempt at an article, kinda saying nothing.
SRE as a term is effectively meaningless at this point. I can count on one hand the number of SRE coworkers I’ve had who were worth a damn. Most only know how to make Grafana dashboards and glue TF modules together.
It's glorified sysadmin work, and the role tbh in the industry should have stayed that way.
If your senior swes can't monitor, add metrics, design fault tolerant systems, debug the system and environment their stuff runs on and need to baby sat there's a problem
Most devs I’ve worked with didn’t have any interest in dealing with infra. I have no problem with that; while I think anyone working with computers should learn at least some basics about them, I completely understand and agree with specialization.
You need someone to be disconnected from getting birthday to display on the settings page and make sure Bingo, Papaya, Raccoon and Omega Star are just up and running. Esp since Omega Star can't get their shit together.
I'm all for a k8s / Terraform / etc focused GPT trained on my workspace to help me remember that one weird custom Terraform module we wrote three years ago, but I don't want it implementing system wide changes without feedback.
For all the grief people give Datadog, it is an incredibly good product. Unfortunately, they know it, and charge accordingly.
And I’ve worked on some of the world’s largest systems and in most cases simply looking for the words: error, exception etc is enough for parsing through the logs.
For everything else you need systems like Datadog to visually show you issues e.g. service connection failures.
Node is throwing EADDRINFO. Why? Well, it's DNS likely but cause of that DNS failure is pretty varied. I've seen DNS Server is offline, firewall is blocking TCP Port 53, bad hostname in config, team took service offline and so forth.
Yes. That's exactly where this stuff is going right now.
The whole point of them is for the team to understand what were wrong and how processes can be improved in order to prevent issues happening again. The idea of throwing away all the detail and nuance and having some AI generated summary will only make the SRE field worse.
Also really don’t understand the benefit of LLM for bringing in relevant data. Would much prefer that be statically defined i.e. when a database query takes 10x longer, bring me the dashboards for OS and hardware as well.
Disagree. It was a good survey of the current AI SRE agents landscape. I don't have my head in the sand, but there are new startups coming up that I hadn't heard about.
>I mean is the AI going to...
Yes. Not one AI but multiple AI agents, all working on different aspects of the problem autonomously and then coming together where needed to collaborate. Single LLMs doing things by themselves are like concrete, creating swarms of agents is like adding rebar to the mix.
I can't solve the article.