Are compilers deterministic?
blog.onepatchdown.net
blog.onepatchdown.net
To me that implies the input isn't deterministic, not the compiler itself
You might argue that this is redefining the question in a way that changes the answer, but I'd argue that's also an academic objection; pragmatically, the important thing isn't the exact language but the intent behind the question, and for an engineer being asked this question, it's a lot more likely that the person asking has context for asking that cares about more than just the literal phrasing of "are compilers deterministic?"
And that's just one really low hanging fruit type of example, there are many more for instance selecting a different optimization path when memory pressure is high and so on.
If we're not going to assume the input state is known then we definitely can't say what the intent behind the question is - for many engineering applications the compiler is deterministic. Debian has the whole reproducible builds thing going which has been a triumph of pragmatic engineering on a remarkable scale. And suggests that, pragmatically, compilers may be deterministic.
or the system upon which the compiler is built (as well as the compiler itself) has made some practical trade offs.
the source file contents are usually deterministic. the order in which they're read and combined and build-time metadata injections often are not (and can be quite difficult to make so).
Either way it's a nitpick though, a compiler hypothetically can be deterministic, an LLM just isn't? I don't think that's even a criticism of LLMs, it's just that comparing the output of a compiler to the output of an LLM is a bad analogy.
lol, should. i believe you have to control the clock as well and even then non-determinism can still be introduced by scheduler noise. maybe it's better now, but it used to be very painful.
> Either way it's a nitpick though, a compiler hypothetically can be deterministic, an LLM just isn't? I don't think that's even a criticism of LLMs, it's just that comparing the output of a compiler to the output of an LLM is a bad analogy.
llm inference is literally sampling a distribution. the core distinction is real though, llms are stochastic general computation where traditional programming is deterministic in spirit. llm inference can hypothetically be deterministic as well if you use a fixed seed, although, like non-trivial software builds on modern operating systems, squeezing out all the entropy is a non-trivial affair. (some research labs are focused on just that, deterministic llm inference.)
The point being that determinism of a particular form is expected and required in the instances where they do that.
(I'm not arguing for or against that, I'm simply saying I've seen it in real life projects over the years.)
Determinism would help you. With a bit of engineering, you could make LLMs deterministic: basically, fix the random seed for the PRNG and make sure none of the other sources of entropy mentioned earlier in the article contribute.
But that barely impact any of the issues people bring up with LLMs.
Determinism is a red herring. What matters is how rigorous the relationship is between the input and the output. Compilers can be used in automated pipelines because that relationship is rigorous.
Suppose you had one of those widely unstable compilers: concretely if you change formatting slightly, you get a totally different binary. It still does the same thing as per the language spec, but it goes about it in a completely different way.
This weak determinism is still useful, because you can still get reproducible builds. Eg volunteers can still audit eg debian binary packages by just re-running the compiler with the exact same input to check that the output matches. So they can verify that no supply chain attack has fiddled with the binaries: at least the binaries belong to the sources the are claimed to.
No clue whether that's a good argument.
The argument is that determinism in compilers isn't particularly important for building software because we did without it for a long time.
Your argument would be... that building software isn't particularly important for building software...?
The actual argument you'd be making would be something like, building software isn't particularly important for survival. Which is pretty obviously true, for the reason you state.
I think this is the more important property and I'm not sure if it has a well-known name. The article obliquely calls it reliability, but regardless it's the key difference from LLMs. Compilers mostly achieve it, ignoring an endless list of exceptions you learn with experience.
LLMs usually don't, even with 0 temperature and floating point determinism.
I wrote "We have not remotely solved the halting problem in the formal sense", which does not read like a claim that LLMs have solved the halting problem to me, but I'm open to rewording it. How would you put it?
I added in a bit about compiler contract, wdyt?
Attempting to get consistent results from floating-point code is another rabbit hole. GCC and clang have various flags for "fast math" which can enable different optimisations that reduce precision.
Before SSE, fp on x86 was done by the "x87" FPU which always had 80-bit precision, even if the type in the source code was 32 or 64 bits — and it used to be accepted to sometimes get more precision than asked for. Java got its "strictfp" mode mainly because of x87.
The real issue is prompt instability (chaos). A one word change to a prompt/spec will produce a drastically different program. Until that is solved there’s no world where we just check in the prompt and almost no one ever has to worry about the code.
You’re agreeing with my point though which is that LLMs aren’t higher level compilers. LLMs aren’t abstraction, they are delegation.
There are certainly useful things you can do with that.
But give that team to a PM who can’t read the output and anything beyond a todo app would collapse under the chaos.
They are much much better at writing plausible looking code or providing plausible sounding answers that look correct but aren’t than any junior developer. They are much better at writing subtle bugs that won’t show up for days and will take hours to debug than any junior developer is.
You can’t mentor them into being a senior developer.
Junior developers (human) can kind of get by without testing, at least in the short term. The LLM can never get by without testing unless it’s some simple one off logic.
They’re also very bad at understanding what edge cases are useful to test and what edge cases can be ignored.
I had a situation just the other day where the LLM produced tests for scores of edge cases, but a managed to leave out a specific sequence that caused a very difficult to diagnose bug. I’m assuming because it decided that sequence wasn’t possible.
But it did include dozens of redundant cases that no sane human ever would have.
I just went through this a couple of weeks ago. My test cases were ok but missing some basic edge cases. I asked it to enumerate the edge cases but that didn’t work well either. I then asked it to list near positive and near negative edge cases, and that worked well for my problem. For bugs I found in the code that weren’t covered, I asked the LLM to add them and/or mark them critical to cover (the case might had been there already but I only let it generate 3 tests with a limited size data set, so not all cases would be covered). That worked well, although I’m sure I could do better (probably by generating more tests, but many current approach is slow because each artifact is generated in a context isolated agent, so I have to speed that up first).
So in other words, determinism (or lack thereof) is the hard problem!
Granted that is not the rigorous definition of determinism. Also some existing compilers are non-deterministic with such a definition, namely when "undefined behaviour" would be triggered. That's precisely the problem with UB.
But lets all hope these are not vital systems we end up depending on.
Sure everything you have unit tests for might stay the same, but unless your unit tests are testing all observable behavior (and if they are they’ll be 100x longer than the code) users will notice incredibly confusing differences in every build.
My point was that "didn't look at the code, looked at the app and gave feature requests and tried the results" is a way of building applications that essentially anyone who doesn't code has been doing all along.
I have actually encountered multiple people proposing that. People in the camp of LLMs are the new high level compiler and code is assembly language.
That’s what prompted this entire are compilers deterministic article.
But ignoring whether it’s regenerated from scratch, when you ask an agent to add minor functionality, it will regularly change existing unrelated functionality.
Some of that is accidental, and could theoretically be solved. For example agents when refactoring won’t copy and paste existing code, they will rewrite it from scratch and frequently change it. Some of it is a required because systems are intertwined.
If the application is small enough you can get away with just trying the results to build confidence. However for any non-trivial application combinatorics explosion means that this stops working very quickly. This has been understood for decades.
> My point was that "didn't look at the code, looked at the app and gave feature requests and tried the results" is a way of building applications that essentially anyone who doesn't code has been doing all along.
That is not building applications. That is delegating building applications to a trusted team of humans. If that trusted team of humans are essentially junior programmers with no agency or ability to learn and improve, your app will fail.
Imagine being a product manager with a team of 100 off shore junior developers that rotate out every day. You could 100% build something useful with that setup. But anything more complex than a todo app would eventually collapse under the chaos.
What is included in the 'verify' step? Does it involve changing the generated code? If not, how do you ensure things like code quality, architectural constraints, efficiency and consistency? It's difficult, if not (economically) impossible, to write tests for these things. What if the LLM does not follow the guidelines outlined in your prompt? This is still happening. If this is not included, I would call it 'brute forcing'. How much do you pay for tokens?
Compilers aren't deterministic in small ways, timestamps, encoding paths into debug information, etc. These are trivial, annoyances to reproducible build people and little else.
You cannot take these trivial reproducibility issues and extrapolate out to "determinism doesn't matter therefore LLMs are fine". You cannot throw a ball in the air, determine it is trivial to launch an object a few feet, and thus conclude a trip the moon is similarly easy.
The magnitude matters, not merely the category. Handwaving magnitude is a massive red flag a speaker has no idea what they're talking about.
https://bazel.build/basics/hermeticity
Although I should add that a hermetic config for bazel is not something you just throw together in an afternoon. And yes, that's work you do before the first bit of source gets compiled. It's worth it though.
When did the girlfriend enter the discussion? Did I miss something?
The OP brings up testimony of someone other than himself who prefers when software drives their car rather than him.
Lots of engineering effort goes into making this be true.
TFA argues that you can't control the inputs perfectly, and so the behavior may differ if you fail to control the inputs. Yeah sure.
But the answer to the clickbaity question in the title is simply "Yes".
A compiler, making my job harder by being unpredictable? All the time.
So did other programmers, users with creative input, random parallel processes running at the same time.
LLMs are actually kind of tame in comparison.
Maybe you don't build or tinker with things enough to have warranted making a dartboard out of the gcc contributor graph, but damnit, some of us do. That compiler is not magic, continually floats around, and when you're just trying to get something from the stage of "doesn't exist at all" to "exists", does absolutely throw curve balls your way. I start wit -O0 -g and then crank up the optimization level once everything works. Otherwise come debug time, shit's missing, stuff happens at weird times, etc. If you don't treat the compiler as spooky, you haven't paid enough attention to it.
Also having an -O0 debug build is standard practice.
My point isn't that compilers are super easy to use and never frustrating, my point is that the notion that LLMs "compile" english to code is a bad analogy. Compilation is a translation from one formal representation to another. LLMs are an interpretation of informal language into a formal language. They just are not at all the same thing.