Real World Examples of GPT-3 Plain Language Root Cause Summaries
zebrium.com
zebrium.com
But good root cause analysis is a matter of establishing facts and building a chain of logic on top of those facts to get to the root cause. You cannot rely on models like GPT-3 to give you reliable baseline facts. Particularly when you are talking about a production issue that needs to be fixed ASAP. The key line that worries me from this blog post is "when results are suboptimal, they are mostly not misleading". 'Mostly not misleading' isn't going to cut it when I'm in the middle of an outage. I think that will prove to be a problem if this tool gets widespread usage.
That being said, I'm a huge fan of applying AI to human-in-the-loop problems and this was a cool idea for how modern language models can be applied.
Secondly, I imagine most cases of "root cause analysis" require you to be very, very clear in understanding the... root cause... So using generalized language models will probably lead to unacceptable errors, which means there are probably better ways of addressing this problem (as per the discussion here on error rates and unacceptable errors in ML-products: https://phaseai.com/resources/how-to-build-ml-products)
What constitutes a useful sequence of facts in root cause analysis is not just some platonic existing thing. It’s a complex problem involving mind-melting log sleuthing, correlating all kinds of disparate metrics, comparing against timestamps of merges and eventually synthesizing the results.
Even seasoned veterans who know systems inside and out struggle with the sheer volume of logs, metrics and facts to compile. And most of the time their approach is based purely on inductive experience with similar incidents combined with heuristics.
This is precisely the kind of problem that ML solutions excel at. It has many hallmarks of a good fit and almost none of the hallmarks of “solution in search of a problem” ML over-engineering.
First, you are conflating the underlying log relevance scoring ML system with the GPT-3 summarizing system. ML is a good fit for relevant log identification for the reasons you describe, although characterizing this software as root cause identification is not very accurate in my opinion, based on the examples you can find on their website. But the value of summarizing a log line into natural language is low, while the cost of misleadingly characterizing that log line is high. Whoever needs to debug this system and find the real root cause (e.g. why did the system go OOM?) probably needs certainty more than the convenience and in all likelihood, they are more likely to correctly summarize what the log line says than GPT-3 is (obviously we don't know since there is no evidence, but I don't work with any engineers whose ability to summarize the contents of a log line would be described as "mostly not misleading").
Secondly, I can't agree with this sentence:
> AI excels at long-tail problems where the cost of failure is high, precisely because human failure is such an expensive problem in those cases
Maybe it depends on domain and tech, but in my experience humans don't fail on out-of-sample data nearly as often as AI does. When they do fail, it is often more predictable to other humans and humans inherently have the ability to assign confidence levels to their conclusion which you don't see in many AI models such as GPT-3. Humans are also more effective at applying rules (e.g. common sense) to improve predictions on out-of-sample inputs. I think of "AI is worse than humans at generalizing to out-of-sample" as being a widely held, well-evidenced belief, but I would be interested if you disagree.
For me, the quintessential example is something like traffic light identification, where models generally struggle to identify unseen variants correctly while humans rarely struggle at it. What examples are you thinking of where AI excels at long-trail problems?
* The root cause of the issue was that the Jenkins master was not able to connect to the vCenter server. ==> Why was it not?
* The root cause was a drive failed error ==> Why did the drive fail?
* The root cause was that the Slack API was rate limited. => Why was it rate limited?
These exampled from the article may be human-readable errors, but that doesn't make them root causes.
To have a root cause analysis, try asking Why five times. https://en.wikipedia.org/wiki/Five_whys
(Just kidding. The space of Everett branches is not countable.)
I wish projects like the linux kernel would work on making log messages, at least those for common events, more readable to an engineer who isn't familiar with kernel internals.
Having detailed human-level descriptions of what's going on and how to fix it is great. But you also don't want to drown out any important details under waves of verbose text.
The solution, then, is to show the extra detail only when it's requested with the -x flag.
This works pretty well, all things considered. The detailed messages are fine, but they could be better—but that's probably always going to be true. It's a start, anyway.
Like most people, I love the idea of ingesting large amounts of data and making it readable. I guess what I would want, personally, is more like a gpt-3 powered stackoverflow search where I can put in an arbitrary cryptic error and get a human readable, root cause based explanation. This is a very interesting use case and I hope they continue to develop it.
I think a lot of commenters here are not quite getting the use case for these kind of summaries. As I see it, it's not that an automatic summary is more accurate, or more complete, than an investigation by a human engineer - it's that the summary lets you resolve issues much faster and with less of a requirement to remember obscure implementation details.
This post shares examples of real summaries generated during beta tests, as well as examples of some sub-optimal outcomes.