I don't think this is true for LLMs. Their output is not deterministic (up for discussion). Their weights and the sources thereof are mostly unknown to us. We cannot really be confident that an LLM will produce correct output based on correct input.
I don't think this is true for LLMs. Their output is not deterministic (up for discussion). Their weights and the sources thereof are mostly unknown to us. We cannot really be confident that an LLM will produce correct output based on correct input.
It's not that LLMs aren't deterministic, because neither are many compilers.
It's also not that LLMs produce incorrect output, because compilers do that to, sometimes.
But when a compiler produces the wrong output, it's because either (1) there's a logic error in my code, or (2) there's a logic error in the compiler†, and I can drill down and figure out what's going on (or enlist someone to help me) to fix the problem.
Let's say I tell an LLM to write a algorithm, and it produces broken code. Why didn't my prompt work? How do I fix it? Can anyone ever actually know? And what did I learn from the experience?
---
† Or I guess there could be a hardware bug. Whatever. I'm going to blame the compiler because it needs to produce bytes that work on my silicon regardless of whether the silicon makes sense.
Sure, those collections of bits tend to do exactly the same thing when executed, but that's is in some sense a subjective evaluation.
---
Szundi said in a sibling comment that I was "completely [missing] the point on purpose" by bringing up compiler determinism. I think that's fair, but it's also why I opened my post by saying "I agree [with the parent], but I want to try to define the language better." Most compilers in use today are literally not deterministic, but they are deterministic in a different sense, which is useful as a comparison point to LLMs. Well, which sense? What is the fundamental quality that makes a compiler more predictable?
I'd like to try to find the correct words, because I don't think we have them yet.
When I’m calling ‘getFirstChar’ from a library, me and the author have a good understanding of what the function does based on a shared context of common solutions in the domain we’re working in.
When you ask ChatGPT to write a function that does the same, your social contract is between you and untold billions of documents that you hope the algorithm weights correctly according to your prompt (we should probably avoid programming by hope).
You could probably get around this by training on your codebase as the corpus, but until we answer all the questions about what that entails it remains, well, questionable.
I use Cursor at work, which is basically VSCode + LLM for code generation. It's a guess and check, basically. Plenty of people look up StackOverflow answers to their problem, then verify that the answer does what they want. (Some people don't verify but those people are probably not good programmers I guess.) Well, sometimes I get the LLM to complete something, then verify the code is completed is what I would have written (and correct it if not). This saves time/typing for me in the long run even if I have to correct it at times. And I don't see anything wrong with this. I'm not programming by hope, I'm just saving time.
However ... your colleagues just do the same.
We'll see how this unfolds. As for now the industry seems to be a bit stuck at this level. Big models too expensive to train for marginal gains, smaller are getting better but doesn't help this. Until some one new idea comes in how LLMs should work, we won't see the 99.95% anyway.
I find that I don't use it as much for generating code as I do for automating tedious operations. For example, moving a bunch of repeating-yourself into a function, then converting the repeating blocks into function calls. The LLM's really good at doing that quickly without requiring me to perform dozens of copy-paste operations, or a bunch of multi-cursor-fu.
Also, I don't use it to generate large blocks of code or complicated logic.
Recently got email about gcc 14.2, they fixed some bugs in it. Can we trust it now, these could be the last bugs. But before that it was probably a bad idea to trust. No, even compiler's output requires extensive testing. Usually it's done at once, just final result of coding and compilation.
> Their output is not deterministic
yes.
> Their weights and the sources thereof are mostly unknown to us
Some of them are known. Does it make you feel better. There are too many weights, so you are not able to track its 'thinking' anyway. There are some tools which sort of show something. Still doesn't help much.
> We cannot really be confident that an LLM will produce correct output based on correct input
No, we can't. But it's so useful when it works. I'm using it regularly for small utilities and fun pictures. Even though it can give outright wrong answers for relatively simple math questions. With explanations and full confidence.
There are 2 things at play here, one is LLM with human in the loop, in which it's just a tool for programmers to do the same thing they have been doing, and the other is LLM as black box automaton. For the former, it's not a problem that the tool is undeterministic, we are double checking the results and add our manual labour anyway. The fact that a tool can fail sometimes is an unsurprising fact of engineering.
I think the criticism in this chain of comment applies more to the latter, but even it always has values to non-tech people, just like how no-code approaches are, however shitty it looks to us software enfineers.
So other path forward could very well be LLMs as they can save lot of time with writing boilerplate code
They can't even perform basic arithmetic (which is not surprising since they operate at the syntactic level, oblivious to any semantic rules), yet people seem to think offloading more complex tasks with strict correctness requirements is a good idea. Boggles the mind tbh.