Looming Liability Machines (LLMs)
muratbuffalo.blogspot.com
muratbuffalo.blogspot.com
LLMs are good at predicting the next word in written language. They are generative; they make new text given a prompt. LLMs do not have base sets of facts about how complex systems work, and do not attempt to reason over a corpus of evidence and facts. as a result, I would expect that an LLM might concoct an interesting story about why such a failure occurred, and it might even be a convincing story if it happened to weave bits of context, accurately into the storyline. It might even, purely randomly, generate a story that actually correctly diagnosed the root cause of the failure, but that would be coincidental based on the similarity of the prompt to text of similar postmortem discussions that were part of its training set.
If you had an extremely detailed postmortem document, then I would expect LLM‘s to do a very good job of summarizing such document.
But I don’t see why an LLM is an appropriate tool for analyzing failures in complex systems; just as I don’t see a hammer being a very effective tool for tightening bolts.
Right now, I am concerned that the relative ease that modern frameworks provide to author LLM based applications, is leading many people to optimistically include LLM technology in attempts to solve problems that it doesn’t seem particularly well suited to solve.
[0] https://www.apmreports.org/episode/2019/08/22/whats-wrong-ho...
> The theory was first proposed in 1967, when an education professor named Ken Goodman presented a paper at the annual meeting of the American Educational Research Association in New York City.
> In the paper,5 Goodman rejected the idea that reading is a precise process that involves exact or detailed perception of letters or words. Instead, he argued that as people read, they make predictions about the words on the page using these three cues:
> graphic cues (what do the letters tell you about what the word might be?)
> syntactic cues (what kind of word could it be, for example, a noun or a verb?)
> semantic cues (what word would make sense here, based on the context?)
This is interesting because this is fairly similar to how LLMs do next-token prediction, but using exclusively backwards-facing clues from the text (although perhaps also the graphic cues if you are talking about multimodal models).
And those people should be fired.
That's because humans are stochastic parrots too. We use leaky abstractions we don't fully understand, or their edge cases. So we don't know what we are saying, but keep doing this as long as it doesn't break. When it breaks, we go to experts, another abstraction we don't really grok. What's the difference between using LLM and using a human expert? Not much if you have no clue about the topic. It's a matter of replacing real understanding with trust.
Even more, can we say we understand anything down to first principles? Probably not. It's a patchwork of people, each with their limited perspective, like the Elephant and the Blind Men parable. The world is based on functional understanding, combining partial perspectives, it's never fully grokked. And what you don't understand, you can't be conscious of, not in it's real meaning. So we might be unconscious operators of language as well, no better than hallucinating LLMs.
And then there is status.
Since you don;t understand anything anyway, you ascertain value and correctness based on status signals.
For you engineer that is going to be, first and foremost the way he looks. Whether he's tall, talks without stuttering, which school he went to.
For LLMs it might be a variant of:
> No one was ever fired for choosing Google/IBM/Microsoft
You should do this because AI is a jagged frontier where humans can't predict if AI is good at something or not.
>AI is weird. No one actually knows the full range of capabilities of the most advanced Large Language Models, like GPT-4. No one really knows the best ways to use them, or the conditions under which they fail. There is no instruction manual. On some tasks AI is immensely powerful, and on others it fails completely or subtly. And, unless you use AI a lot, you won’t know which is which.
https://www.oneusefulthing.org/p/centaurs-and-cyborgs-on-the...
I recognize that its still a robot and you have to give it stoic, stern guidance to keep it on point.
Also, having a good understanding for the domain youre seeking to build understanding in, but with the augmentation of an AI assist that can formulate the Output of the thought in a more complete and packaged manner than one is able to do without leveraging AI as a tool. And as someone said "And its not going to get any worse, its only going to get better" I think people are really underestimating, and under-utilizing AI tools.
But they are terrifying when you consider just that - a divide between those who have/use AI and those who are subjugated by those who do.
If all the people whinging on here took some of that time and actually ‘formally’ experimented with LLMs, measuring their reliability / correctness against a human in some task in their domain, they may be surprised by the results. And no, “I tried Copilot for an afternoon and hated it” doesn’t suffice.
At work, recently, I happened across an opportunity to do just this. There was a task that I thought that it was quite possible for an LLM to be good at. The task was such that we could run a bit of a ‘study’ to see how the LLM fared against a real-world meat-bag person. A skilled person at that. The person we would’ve had do the job in the first place. The LLM and the human agreed the vast majority of the time (>99.9%), and the LLM with its infinite ‘attention’ (heh) was on more than one occasion correct in cases where the human wasn’t, because it was a repetitive task that’d put someone to sleep.
It was a task that involved parsing language, but I’m sure one that the geniuses on HN would say requires “understanding semantics”, “intelligence” or whatever armchair philosophy nonsense they whip out in lieu of intelligent conversation. It wasn’t sentiment analysis, categorisation, or anything of that nature. Maybe it’s something I could’ve tackled without an LLM, with traditional ‘deep learning’, or whatever. I really don’t know. I couldn’t think of a way off the top of my head. It was beyond ‘throw linear regression at it’ anyway.
Software engineering isn’t engineering, but evidently computer science is increasingly not a science. This industry deserves all of the belittling pejoratives people throw at it. There’s a disappointingly large contingent of utterly unengaged, incurious, drones that let their entire professional skill set be guided by whatever some other incurious drone says on a social network.
Makes me wonder why that person wouldn't just write code to automate it instead of manually making the changes.
I have several dozen custom code generators in some of my projects, where I just have a spec file written in a DSL.
Many of the things people who work in the trades do is repetitive work, most of which is not currently possible to automate. It's the same for mental tasks.
If it requires so little thought to do, that it causes human error ... then it is probably nonsensical to NOT automate it.
This is a well-tred argument, and it is much stronger when it holds itself to, especially in the short term, augmentation is likely, not wholesale delegation. It gets weak when it tries to couple that to asserting that "all it does" is lightly rephrasing training data. It's somewhat trivial to demonstrate this is false, even society as a whole has noticed that it's beyond a parrot, and it's worrying.
Do you have a link? The most recent information (not a paper) I could find was back in June: https://engineering.fb.com/2024/06/24/data-infrastructure/le...
This is the HN discussion to which you may be referring: https://news.ycombinator.com/item?id=41326039
Now explain what humans are doing and why it is different. RCA can be done fully remotely, so we know it can be modelled as a word-prediction model with context and a "how is all this context made consistent?" prompt. There is every reason to expect that an advanced text prediction model would be excellent at predicting the most-technically-correct response to that prompt.
I've seen a few people make this argument as though text prediction is some specific sub-field that can be solved independently of having a world model and intelligence. That doesn't hold up at all, if we solve text prediction we've solved general intelligence - all aspects of human intelligence are less complex than being able to predict the most objectively correct next word in a sequence of words because all aspects of intelligence can be framed as a text-prediction problem. A system can't predict the next word in a sequence and fool a human without being at least as clever as a human.
> predict the most objectively correct next word in a sequence of words
Currently all LLMs are only determining the most probable next token, but this means they are not aware of the probability of the entire sequence of tokens they are emitting. That is, they can only build sentences by picking the most probable next word, but can never choose the most probable sentence. In practice, there are a great many very likely sentences that are composed of a fairly unlikely words. When we use the output of an LLM we're thinking it of a sequence sampled from the set of all possible sequences, but that's not really what we're getting (at least as far as probability is concerned).
There are approaches to address this: you can do multinomial sampling instead of greedy so that are casting a slightly large net or you can do beam search where your once again trying to search a broader set of possible sentences choosing by the most probable sequence. But all of these are fairly limited.
Which gets to your first remark:
> Now explain what humans are doing and why it is different.
There's very little we really know about how humans reason, but we are certainly building linguistic expressions at with a more abstract form of composition. This comment for example was planned out in parts, not even sequentially, and then reworked to the whole thing makes some sense. But at the very least humans are clearly reasoning at the level of entire sequences as their probability rather than individual tokens at a time.
The word "planning" almost tautologically implies thinking ahead of the next step. When humans write HN comments or code they're clearly planning rather than just thinking of the next most likely word over and over again with some noise to make it sound more interesting. No matter how powerful and sophisticated the mathematical models driving the core of LLMs are, we're fundamentally limited by the methods we use to sample from them.
But then your argument seems to drift into humans not producing text as a series of words. I'm not sure how you type your comments but you should upload a YouTube video of it as it sounds like it'd be quite a spectacle!
If your argument is that LLMs can't reason because they don't edit their comments, it'd be worth stopping and reflecting for a few moments about how weak a position that is. I wrote this comment linearly just to make a point with no editing except spellchecking.
As systems humans and LLMs behave in observably similar ways. You feed in some sort of prompt+context, there is a little bit of thinking done, a response is developed by some wildly black-box method, and then a series of words are generated as output. The major difference is that the black boxes presumably work differently but since they are both black boxes that doesn't matter much for which will do a better job at root cause analysis.
People seem to go a bit crazy on this topic at the idea that complex systems can be built from primitives. Just because the LLM primitives are simple doesn't mean the overall model isn't capable of complex responses.
Do you write the middle of the sentence first, come back to the start then finish it?
Am I the only one that does this?I'll have a central point I want to make that I jot down and then come back and fill in the text around it -- both before and after.
When writing long form, I'll block out whole sections and build up an outline before starting to fill it in. This approach allows better distribution on "points of interest" (and was how I was taught to write in the 90's).
Before the model arrives at the candidates for the next word it first computes vectors in high dimensional space that combine every combination of words in the context and extract semantics from it. When producing the next token the model effectively has already "decided" the direction where the answer will go and that is encoded as a high dimensional vector before being reduced to the next token (and the process repeated)
No it hasn't, if you tell it to write a random story and it starts with "A", it hasn't figured out what the next word should be, and you run it many times from that "A" you will get many different sentences.
It will do some adapting to future possibilities, but it doesn't calculate the sentence once like you suggest it will, it comes up with a new sentence for every token it generates.
If instead the prompt says it should emit 5 repetitions of the letter "A" unsurprisingly it will compete the output with " A A A A".
The task is performed by emitting tokens, but on order to correctly execute the task the model has to "understand" the prompt sufficiently well in order to choose the next token (and the next etc).
Now, obviously current LLM models have severe deficiencies in the ability to model the real world (which is revealed by their failures to handle common sense scenarios). This problem is completely compounded by a psychological factor in which we humans tend to ascribe more "intelligence" to an agent that "speaks well" so the dissonance of a model that sounds intelligent and yet sometimes is so hilariously stupid throws us off rails.
But there is clearly some modeling and processing going on. It's not a mere stochastic parrot. We have those (Markov chains of various sorts) and they cannot maintain a coherent text for long. LLMs OTOH are objectively a phase transition in that space.
All I'm trying to say is that whatever is lacking in LLMs is not just merely because they "just do next token prediction".
There are other things these models should do in order to go to the next level of reasoning. It's not clear if that can be achieved just by training the model on more and more data (hoping that the models learn the trick by themselves) or whether we need to improve the architecture in order to enable the next phase.
They're normally trained to output a probability distribution for the next token and _sample_ from that distribution. Doing so iteratively, if you work through the conditional probabilities, samples from the distribution of completed prompts (or similarly if you want to stop at a single sentence) with the same distribution as the base training data.
You're right that you can't pick the most likely sentence in general, but if there exists a sentence likely enough for you to care then you can just repeat the prompt a few times and take the most common output, adjusting the repetition count in line with your desired probability of failure. Most prompts don't have a "most likely" sentence for you to care about though. If you ask for meal suggestions with some context, you almost certainly want a different response each time, and the thing that matters is that the distribution of those responses is "good." LLMs, by design, can accomplish that so long as the training data has enough information and the task requires at most a small, bounded amount of computation.
Beam search.
Why should I believe, in the first place, that this is even a coherent concept?
They make language that sounds like what thinking people make, therefore they must be thinking like people do, duh! /s
Because an LLM is really good at convincing laymen it knows what it's doing. It's really a conman simulator :)
It is great at language based tasks but people use it as an oracle for everything without understanding the limitations. It's just that friendly convincing tone that makes them think they're taking to a super intelligent being.
https://engineering.fb.com/2024/06/24/data-infrastructure/le...
The real problem in cases like this and other applications, as you and many others have mentioned, is that LLMs are basically correlation machines. They can find very complex, immensely-multivariate correlations in large data sets, and reproduce these correlations very well. But they cannot reason (so far) beyond these correlations in other to find deeper, less obvious causal relationships. They’re simply not trained to do that, yet. But it’ll come…!
That is actually extremely correct. The only purpose for software is automation, which is the elimination of labor. Getting that wrong directly influences your quality of product more than any other downstream factor.
Dijkstra's quote succinctly summarized the arguments:
"The question of whether a computer can think is no more interesting than the question of whether a submarine can swim."
I wonder if his point was that the question didn't need to be answered for whatever they were discussing at the time?
I think I missed your point. Wasn’t that exactly what I said?
I take it you never played a game in your life?
Fighting games are just better versions of Rockem Sockem Robots.
And 4X games could be fancier versions of Settlers of Catan.
Could you point me to some of the places/articles where this is being shown? I'm definitely amongst those who have bought in to the common cop out you are rebutting here
In my experience, LLMs are word predictors, and the impacts of that fact are not immediately obvious.
LLMs are capable of "explaining" what code does. What it is doing under the hood is pattern matching: I've seen code that looks like X, with an explanation that looks like Y
LLMs are capable of formatting text. It has seen English written like X, that is reformatted to look like Y
One resounding fact my team has found over and over is that "the things we think are hard for LLMs aren't necessarily hard; the things we think are easy aren't necessarily easy"
LLM style pattern matching and rewriting does not preserve semantics, except accidentally due to an overwhelming amount of examples.
What we call reasoning can be the art of finding a chain of small correlations to connect ends of a big correlation. Some sort of quantum-powered DFS algorithm.
However a reasoning machine is just a machine. Someone needs to tell it what to reason about.
Not wanting to advocate for using LLMs for RCA, which is a dumb and dangerous idea for all the reasons the OP mentioned - but I am getting allergic to the phrase "it just predicts the next word".
Yes, that is how the "API" of an LLM works, but on itself, it says nothing about how the LLM does the prediction and how complex the internal model is that it uses for that task.
It's obvious that the task of predicting has a huge variance in complexity, depending on which word has to be predicted.
E.g. the sentence "The apple does not fall far from the" could be completed by a decently trained Markov chain, whereas (correctly!) completing "sqrt(153847)=" would either require an impossibly large training set or an internal model that can parse integers and perform square root calculations.
Yet both are on the surface "predict the next word" tasks.
The actual complexity of LLMs' internal models seems to still be poorly understood. That's not to say the model has superhuman ability or even reaches human abilities. But the point is that it's still really hard to make predictions how complex the reasoning is that an LLM performs for a task.
In the OP's example, we can't really say if the LLM "has base sets of facts about how complex systems work", because we don't really know if and how "facts" would be represented inside the model. If enough in-domain examples were in the trainset, there is no fundamental reason why it wouldn't have learned such a set of facts.
Just saying "it can't reason at all because it's just a next word predictor" is mixing up different layers of meaning and, I believe, does not lead to more insight.
1. Take your existing incident reporting / review docs (you have those, right?) that cover everything including 5-why incident reporting and analysis.
2. Fine-tune a Llama-3.1-70b [1] LoRA on the data associated with the outage as input, and the root cause analysis as the output
3. Tada! You have a state-of-the-art custom LLM that is good at analyzing your outages and guessing what the root causes might be.
It's a little shocking sometimes to me how underutilized fine-tuning is. Most of the "learning" happening in "machine learning" is in training — yes, it's definitely true that LLMs can exhibit a surprising amount of "in-context learning" via prompts, but it's a surprising because learning during training is so much more powerful, and it's surprising that in-context works at all. Honestly even just fine-tuning an 8b model — which is totally doable on a 3090/4090 — can yield SOTA results on task-specific performance. It's so much better than prompting!
1: I mean you could also finetune 405b, but you'll need a lot of GPUs both to train it and to run it. In my experience (public benchmarks be damned; the models all tend to saturate the public benchmarks, even though everyone claims not to train on them), 70b on internal evals tends to perform similarly to gpt-4o, and with OSS models you have somewhat more ownership and control.
But there are plenty of tasks that even large, well-trained models struggle with. If the OP is struggling to get useful root-cause analysis for cloud service incidents out of an existing large model, that seems exactly like a use case where a finetune would shine.
Also, finetunes don't have to be just for small models! Medium-sized models like Llama-3.1-70b can be finetuned, and if you want to burn a lot of GPUs you can finetune 405b as well.
Typically issues arise because they are novel and are unforeseen. If we did see these issues beforehand they'd be fixed! LLMs by definition are trained by example, so I fail to see how finetuning LLMs on things that have already happened to be helpful for determining the root cause of a novel issue.
LLMs seem to lack systemic modelling that humans do. I can see LLMs being practical for this if it is shown that LLMs are capable of modelling scenarios outside of their dataset, but thus far none such examples exist.
I'm not even sure that LLMs are even capable of solving standard bugs see: [1]. Hallucination seems to be a significant hurdle and any time spent validating the fixes of an LLM is wasted when it could be spent tackling the bug head on. The amount of energy spent espousing garbage requires an order of magnitude more effort to invalidate.
[1]. https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-f...
This shows so much faith in management!
Feel free to point out any error in my logic:
There are huge financial incentives--tens if not hundreds of billions of dollars--for developing an LLM which can solve novel bugs. So surely there exists AI companies developing an LLM capable of doing so. If an LLM capable of solving novel bugs exists, AI companies would rush to showing it off to capture tonnes of VC money. AI companies could show off their fancy bug-fixing LLM by closing issues on public Github repos using said LLMs.
No such mythical LLM exists. We are thus left with two choices:
1. My logic is flawed or there is an alternative possibility I haven't considered.
2. The LLM capable of doing what OP asserts doesn't exist and can't be made, despite their assertion that it is trivial to fine tune and put into application.
Your original wager was that it would be in production, not that it would work.
What I mean by successful is that the LLM can generate accurate (>80%) incident response reports and propose correct fixes.
I'm fairly certain anyone literate could have read the rest of my comment and parse out what a successful deployment means.
There is a way to disprove that the statement: "There is no god" by simply showing a counterfactual god.
There is however no way to disprove the statement: "There is a god."
Likewise, there is a way to disprove the statement: "LLMs cannot be successfully used for X application." By showing that LLMs have been used in X application.
Again, there is no way to disprove the statement: "LLMs can (eventually) be used in X application."
The meat of my question was meant to demonstrate a failure to apply the scientific method.
that both parties to the argument agree is a god.
>Likewise, there is a way to disprove the statement: "LLMs cannot be successfully used for X application." By showing that LLMs have been used in X application.
Again there the point of argumentation will be the word "successfully", the LLM would have to be such an overwhelming success at what it is trying to do that one cannot weasel out of it with "successfully".
To answer your question more directly, I would go with option 1.
It seems like your assumptions are unfalsifiable. As computer scientists, I believe it's important that our hypotheses are testable. If a hypothesis is unfalsifiable, then the hypothesis is no better than theology and should be discarded.
What's stopping you and other VCs just pouring endless money into an idea that won't work?
Your variant of the "scientific method" would've meant we never discovered electricity, or invented airplanes, or really anything else, because why bother trying? If it worked someone else would've done it.
Are you saying that if someone finetunes a current SOTA LLM with incident response data and demonstrates that it doesn't work that you'll say that LLMs are infeasible for this application? That would invalidate the hypothesis: "X application can be done on current LLMs."
Such a test could never invalidate the the hypothesis: "X application can (eventually) be done on LLMs."
If it's the former hypothesis you were asserting, then yes I agree that it is testable, but I'm fairly confident you were asserting the latter.
Earlier I had asked you: "I think we're at an epistemic impasse here. At what point would/could you be convinced that LLMs are incapable or unsuited here?"
And you have yet to provide a response.
Proving that LLMs can never do this would require extremely rigorous theoretical evaluation that even top ML labs are currently unable to do, given the problem of interpretability. In general proving a negative is typically harder than a positive, since a single experiment succeeding proves a positive, but a single experiment failing does not prove a negative; generally science does not demand that scientists attempt to prove a negative when running experiments, or else nearly every drug trial, for example, would be impossible to perform. Complaining that you have staked out a very difficult to defend position — that it's impossible for LLMs to generate good incident reports — does not mean your ideological opponents, who have simpler positions, must do your proof work for you.
Please refer to this: https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)
Just to make sure we both understand what burden of proof is:
Suppose two people are having a debate over whether or not a teapot exists in the orbit of Jupiter which is impossible to observe via telescope. Where does the burden of proof lie?
Just to reiterate plainly:
Does the burden of proof lie on the person making empirically impossible to falsify claim or the person making the empirically possible to falsify claim?
Which of the following two claims is impossible to empirically falsify?
1. "LLMs can eventually be used to produce good incident reports."
2. "LLMs can never eventually be used to produce good incident reports."
> That does not mean the hypothesis is true, it only means you haven't falsified it.
Correct. You can't prove a negative, you've only just figured this out? After I listed TWO simple examples that could be found in an introductory philosophy class?! In the Wikipedia article consisting of a couple of paragraphs I linked you?
PLEASE JUST READ. PLEASE JUST READ. PLEASE JUST READ. PLEASE JUST READ.
You can't prove negative statements. I want you to admit this so I know you understand, now repeat after me: "You can't prove negative statements."
I know this is likely wasted on you, but here goes:
You can't prove the hypothesis: "There does not exist an unobservable teapot in the orbit of Jupiter."
For the SAME reason I can't prove the hypothesis: "LLMs can never eventually be used for generating good incident reports."
For the SAME reason I can't prove the hypothesis: "God doesn't exist."
For the SAME reason I can't prove the hypothesis: "A unicorn does not exist at the center of the Earth."
I fully admit this in the comment you've supposedly read.
This is why the burden of proof lies on the person making the positive claim (this is you).
1: https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)#P...
I imagine it is because of the costs of building a dataset.
I think we will see more of this. There are some big privacy concerns, but "an internal RAG for every business" will probably find some customers.
lora does not work though, you need full parameter training if the knowledge isnt already present in pre training set.
Because it would be extremely profitable if they could. That’s how we’ve ended up in this mess, the promise is just so tantalizing even if it’s not based in reality.
Think about it - these models have been trained on mountains of technical docs, incident reports, and discussions about cloud systems. They've soaked up a ton of knowledge about how these systems work and what tends to go wrong.
You're right that they don't reason from first principles, but they're incredibly good at spotting patterns. When you feed an LLM details about an incident, it can quickly pick up on similarities to known issues and suggest potential causes. It's not just making stuff up - it's drawing on a vast pool of relevant information.
And these models aren't just spitting out random text. They've gotten really good at understanding context and applying the right knowledge to a given situation. They can take in technical details about an incident and connect them to possible root causes in ways that can be surprisingly insightful.
I'd argue that LLMs can actually be pretty effective tools for analyzing failures in complex systems. They can process tons of information quickly, spot connections humans might miss, and generate multiple plausible hypotheses for what went wrong. They're not replacing human experts, but they can definitely augment our capabilities and speed up the initial triage process.
There are already some success stories out there of LLMs being used effectively in IT ops and incident analysis. And as these models get fine-tuned on more specific cloud-related data, their performance in this area is only going to improve.
You're right to be cautious about applying LLMs everywhere just because we can. But in this case, I think they actually have a lot to offer when it comes to cloud incident analysis, especially when used alongside other tools and human expertise. They're not a silver bullet, but they're definitely more than just elaborate text predictors when it comes to tasks like this.
This comment was written by an LLM. If it can make a coherent argument, I guess it can process information about an incident, too.
Not that I am claiming they are great for post mortem analysis, I not really have an opinion about that.
*blink*
...why are they boasting about migrating to Java 17... in 2024?
...and from what versions were they migrating from? Java 17 didn't introduce any significant breaking-changes (IME) for users on Java 16 or even the next previous LTS version, Java 11 - I don't work at Amazon, but surely Amazon isn't in the habit of running on unsupported JVMs? - so assuming these Java projects were being competently maintained, then the only work actually required to migrate to 17 is changing your `org.gradle.java.home=` path to where JDK 17 followed by running your test suite. If Amazon was using a monorepo then 1 person could do this in 5 minutes with a 1-liner awk/sed command - whereas I expect it would likely take an AI far longer to do this, if it's even able to make sense of an Amazon-sized monorepo - they'd also likely re-prompt it for every separate project for reliability's sake. So after considering all that, the "50 days" number he gives, without any context either, is a nice shorthand to communicate his disconnection from what really goes-on inside his org... or he's lying - and he knows he's lying - but he also knows there won't be any negative consequences for him as a result of his lying, so why not lie if it gives you a good story to tell for LinkedIn?
In conclusion: these remarks by leadership unintentionally make the company look bad, not good, once you fill-in-the-blanks to make up for what they dind't say. (What's next...? Big Brother increasing our chocolate ration to 20 grammes per week?)
-----
One more thing: the comment-replies to the post on LinkedIn are utterly derranged and I genuinely can't put my sense of unease into words.
I don't want to dig up my old credentials to sign in to see the entire set, but from what's publicly visible...
Perhaps the sense that some comments are mostly desperate scrambling to self-promote, and a few of those also contain fawning ingratiation which they seem to think may be reciprocated?
A wise decision.
> ...but from what's publicly visible...mostly desperate scrambling to self-promote ... fawning ingratiation ... reciprocated
Well, yes - there were enough of those - and they were bad enough, but what threw me off was a screenful of rambling comment replies from a single person accusing Amazon of "being racist" against her (a white woman in the US, if her avatar is accurate) while also making references to some kind of lawsuit she was pursuing - with a surprisingly restrained sprinkling of emojis throughout.
I want to know more about this. What were the behaviors before and after the change? In what year or versions did this happen? Are there any write-ups I can read?
I think it was more like this:
“We need to upgrade Java at some point.”
“That would take thousands of hours and a billion dollars.”
“Ok, forget it for now then.”
“Well, we could try to automate it with an LLM…”
This means that in most cases, these RCAs are the output of a long and over engineered incident review process that was designed to impress the higher echelon.
The problem is, that in a decently sized corporation, you have tens to hundreds of daily fuck ups (also known as "incidents") that completely suck out the free time out of engineers that have to navigate the long game of post incident management process.
The utilisation of LLMs on these cases are just engineered solution to the problem of organisational bureaucracy.
I would take under process any day of the week. From my experience, adding a process is far easier than removing it.
Post-mortems are as short as possible because technical people usually write them and have better things to do. Getting an LLM to do them will only remove this feedback channel, as it is much easier to ignore a LLM suggesting more time or money is spent on quality than it is to ignore a human.
The rate at which llm/llm compound systems can produce output > the rate at which humans can verify the output
I think it follows that we should not use llms for anything critical.
The gunghoe adoption and hamfisting of llms into critical processes, like an AWs migration to Java 17, or root cause analysis is plainly premature, naive, and dangerous.
If the verification systems for LLMs are not built out of LLMs and they're somehow more robust than LLMs at human-language problem solving and analysis, then you should be using the technology the verification system uses instead of LLMs in the first place!
The issue is not in the verification system, but in putting quantifiable bounds on your answer set. If I ask an LLM to multiply large numbers together I can also very easily verify the generated answer by topping it with a deterministic function.
I.e. rather than hoping that an LLM can accurately multiply two 10 digit numbers, I have a much easier (and verified) solution by instead asking it to perform this calculation using python and reading me the output
I think automating verification generally might require general intelligence, not an expert though.
But that hasn't stopped the last 40 years from happening because computers made fewer mistakes than the next best alternative. The same needs to be true of LLMs.
There is nothing in the theory that prevents you creating a program that verifies a particular specific program.
There is an entire field dedicated to doing just that.
This is what gets swept under the rug whenever formal methods are brought up.
For example, many things can be proven about the following program without having to solve any general problem at all:
echo “hello world”
Similarly for quick sort, merge sort, and all sort of things. The degree of formality doesn’t have to go to formal methods which are only a very small part of the whole field
Congratulations, you just launched all the worlds nuclear missiles.
This is to spec since you didn't provide one and we just fed the teletype output into the 'arm and launch' module of the missiles.
I’m still not sure why some of us are so convinced there isn’t an answer to properly verifying LLM output. In so many circumstances, having output pushed 90-95% of the way is very easily pushed to 100% by topping off with a deterministic system.
Do I depend on an LLM to perform 8 digit multiplication? Absolutely not, because like you say, I can’t verify the correctness that would drive the statistics of whatever answer it spits out. But why can’t I ask an LLM to write the python code to perform the same calculation and read me its output?
> I think it follows that we should not use llms for anything critical.
While we are at it I think we should also institute an IQ threshold for employees to contribute to or operate around critical systems. If we can’t be sure to an absolute degree that they will not make a mistake, then there is no purpose to using them. All of their work will simply need to be double checked and verified anyway.
2. I believe this is falsely equating what llms do with human intelligence. There is a skill threshhold for interacting with critical systems, for humans it comes down to “will they screw this up?” And the human can do it because humans are generally intelligent. The human can make good decisions to predict and handle potential failure modes because of this.
That is, ignoring all the other myriad, multidimensional other nuances of human/social interactions that allow you to trust a person (and which are non-existent when you interact with an AI).
We have a project working on very large code-base in .NET Web Forms (and other old tech) that needs be updated to more modern tech so it can be in .NET 8 and run on linux to save hosting costs. I realize this is more complicated that just convert to later versions of Java, but it's roughly the same idea. The original estimate was for 5 devs for 5 years. C-types decide it's time to use LLMs to help this get done. We use both Co-Pilot and later others, Claude of which turned out to be the most useful. Senior devs create processes that offshore teams start using to convert code. Target tech can be varied based on updated requirements, so some went to Razor pages, some to JS with .NET API, some other stuff. Looks to be pretty good modernization at the start.
Then the Senior devs start trying to vet the changes. This turns out to be a monumental undertaking. Literally swamped code reviewing output from the offshore teams. Many, many subtle bugs were introduced. It was noted that the bugs were from the LLMs, not the offshore team.
A very real fatigue sets in among senior devs where all they're doing is vetting machine generate code. I can't tell you how mind numbing this becomes. You start to use the LLMs to help review, which seems good but really compounds the problem.
Due to the time this is taking, some parts of the code start to be vetted by just the offshore team, and only the "important things" get reviewed by Senior devs.
This works fine for exactly 5 weeks after the first live deploy. At that point the live system experiences a major meltdown and causes an outage affecting a large number of customers. All hands on deck, trying to find the problem. Days go by, system limps along on restarts and patches, until the actual primary culprit is found, which turns out to be a == for some reason being turned into a != in a particular gnarly set of boolean logic. There were other problems as well, but that particular one wreaked the most havoc.
Now they're back to formal, very careful code reviews, and I moved onto a different project on threat of leaving. If this is the future of programming, it's going to be a royal slog.
The year is 2034; after a surprisingly cutthroat economic trade-war fought between China and the US left the world in the throes of another great recession, a growing wave of anti-China sentiment captures the attention of domestic political leadership which cultivates the movement despite (or more likely: because of) the growing interest from xenophobic reactionaries and other populist movements looking to scapegoat their way out of a dip in GDP - eventually those same poltiical-actors win the presidency and use their democratic mandate to instigate a new McCarthy-era of anti-China paranoia leading to utterly deranged domestic policy, namely as the executive ordering, by-decree, that the State Department terminate the employment of anyone who even speaks Mandarin[1] - a few weeks later in the South China Sea another Filpino/Sino boat-ramming incident escalates into something serious - the US Navy urges the US civilian government to communicate with China over the D.C.-to-Beijing "red telephone" deescalation e-mail system, but no-one knows how to communicate to the Chinese in their own language, so the overworked federal employee manning the red-Outlook-inbox sees nothing wrong with simply having that Microosft Office 365 CoPilot translate it for him - the same AI bot that's somehow always on his screen with that distracting sidebar (despite the best efforts of the US Federal Gov's Active Directory Group Policy) - it wasn't long before the first warheads exploded over North America that the President learned the AI translated the polite request to China for them to "please stop ramming the fishing boats" was received by them as "I'll ram my fish into your Junk...boats". If there's any upside to this story, the collective mass of AI were wiped out first by the high-altitude EMP bursts, leaving us humans with the last-laugh before we were all incinerated moments later - while those not fortunate enough to die instantly instead suffered months of prolonged fatal radiaiton sickness while what little left of civilization collapsed around them[2].
[1]If you think that's too ridiculous to be realistic, consider the Japanese internment-camp policy or Trump's declared Muslim ban. Elsewhere, in the late-1970s (in the age of CT Scanners and VHS tapes), Pol Pot targeted people for wearing glasses.
[2]Blame James Burke's editorial slant in his documentary series' for turning me into a nhilist.
[edit] I herewith introduce a new shitcoin called NukeCoin. Everyone in China and America gets a NukeCoin that will go up in value every day you hold it that no one nukes anyone.
Even if this accuracy is 95% then in a complex system the probability of getting to the right answer diminishes with each new step being added. This is also the key tenet of an agentic system.
While the analysis in the blog is excellent but an answer needs to be found. A layer on top of LLMs for error control/check.
As an analogy, in the analog to digital transmission stack of OSI, an error correction mechanism such as frame check sequence (fcs) detects transmission errors in the data link layer.
If for example you have an LLM agent that's effectively "solved" every security flaw your software may encounter for the next 50 years, unless it can simultaneously impart 50 years of training to the people who rely on the software, it's done nothing but introduce us to more complex flaws that we would need approximately 49 more years of experience to tackle ourselves.
But you don't offload it in the sense that you expect the tool to completely take the wheel.
You ask it for suggestions to inform a human. If the suggestions turn out to only be a distraction in your environment then you abandon the tool.
For plenty of environments the suggestions will be hugely useful and save you valuable time during an ongoing outage.