Jev-Driven SRE Diagnosis: What Worked and What Failed
sregym.com
sregym.com
Why? What’s the end goal? Lower cost? Faster analysis?
I configured an agent to respond to pages via Slack. It has access to ClickStack (for telemetry), Kubernetes, and GitHub. It runs one of the Sonnet models. The cost is so low, responding to less than a dozen pages a day (mostly from a very sensitive error count alarm) that swapping got Jev makes zero sense. This is especially true if the results are less trustworthy.
We can also add a confidence output to this pipeline, so we invoke a heavier reasoning agent (like Sonnet) when something suspicious deserves a closer look.
An SRE goes and fixes the cause of the recurring alerts.
And to answer the obvious question: because latency. This thing turns on the light or water valve in under 100ms on a laptop from 2001 with 8gb ram .. all running locally on the laptop.
The same way I'd not trust an SRE when they've "got a feeling", I don't intend to start trusting Jev. The audit of process doesn't allow audit of reasoning.
Labeling something as "don't escalate" that should have been isn't great. Paying 5x and having to wait 10s instead of 100ms or whatever for a reasoning model (with tools?) is very likely worth it.
We had an LLM based SRE product from a vendor I won't mention because we had an NDA as our management is sucking them off by trading whitepapers for discounts. It was like a drunk monkey with a wrecking ball. Think we had to pull it in under 2 weeks because it took out multiple production systems and lead to an entire cluster failover.
The meat sacks now know they have job security.
And a lot of "wishing up a feature", I was using it to explore solutions for a given problem in too I didn't knew 100% and it pretty much came to same solution I wanted to do but... the capabilities were not there in the tool so it just started making up probable config clauses, and of course, it didn't work.
Even on simpler stuff there were traps, for example in middle of debug session I asked it to modify Gitlab config to add request duration logging, so it added correct config format to a flag that didn't exist (option was there, just under different name), because it didn't bother to read the docs (since then I generally link it the docs first so it doesn't try to remember and get it wrong).
All of that is both very dangerous, and also easily fixed by just having competent operator there. And as a tool it's great, as replacement it is just AI bros delusion
Yeah, and the article puts the cost of all 105 diagnoses at about $0.15 in Jev calls. At a dozen pages a day, I wouldn't worry much about the bill for either model. I'd be more interested in the pass rate: 76.2% vs 77.8% for GPT-5.6 Sol (medium).
I can see trying a cheaper model first if you're handling lots of requests and it can resolve most of them without escalating. A dozen pages a day doesn't seem like a reason to add that extra step.