Dunno what you mean about having no rules imposed on your code -- you can define the semantics of the language formally and execute them on a machine, same as you can for C, JavaScript, or Prolog.
Dunno what you mean about having no rules imposed on your code -- you can define the semantics of the language formally and execute them on a machine, same as you can for C, JavaScript, or Prolog.
Yeah but I could add 3 billion to the pointer instead. Now what happens, if we're trying to define behavior?
> you can define the semantics of the language formally
If the semantics aren't a direct translation into CPU opcodes, then it's not machine code. And CPU opcodes can access all memory at all times including self-inspection so you can't change anything for optimization purposes without an additional framework on top.
The cell whose address is three billion greater, mod the size of the address space, now has a value of 20?
> If the semantics aren't a direct translation into CPU opcodes, then it's not machine code.
Just because the semantics are defined that way doesn't necessarily mean you couldn't make another implementation that obeys those semantics. For instance, QEMU can perform optimizations on the code it runs, and invalidate them if the program changes those instructions. The semantics of reading from and writing to memory are unchanged by these optimizations, and the implementation doesn't need to do anything special on read.
Cool. It sounds like you figured out a restricted form of pointers for stack variables, since there's no danger of corrupting the code or breaking important invariants elsewhere.
Now imaging tightening those restrictions a bit. Instead of wrapping pointers inside the entire stack, wrap/constrict pointers inside their original variable.
Now you have a language with no UB that allows the optimizer to replace y with 42!
TL;DR: If you let pointers go hog-wild, your language has UB even before you consider optimizing. You don't need to add UB to allow optimizations. If you don't let pointers go hog-wild, in a language without UB, then you can enhance those restrictions to allow the y->42 optimization while still not having UB.
> For instance, QEMU can perform optimizations on the code it runs, and invalidate them if the program changes those instructions.
The machine can take shortcuts, and when QEMU is acting as a virtual machine it can do that. But the compiler can't take any shortcuts, because it never knows when external code is going to examine random bytes and need them to be unchanged.
Yep, that's one approach, though that's just as hard to compile as one that is defined to trap and end execution on an out-of-bounds write.
As I said up-thread, it's possible to define a C dialect without UB, it's just not what most of its users actually want out of the language, since on current hardware it adds significant runtime overhead.
> The machine [...] the compiler
Isn't QEMU both the machine and the compiler (EDIT: perhaps better-put, what's the distinction? They're both just the implementation.) The bytes only need to be unchanged on read from inside the program; if you compile them to something else, that's fine. Prolog programs can introspect with clause/2 and it doesn't cause trouble there.
> As I said up-thread, it's possible to define a C dialect without UB, it's just not what most of its users actually want out of the language, since on current hardware it adds significant runtime overhead.
So my core point is this: It's not UB that enables this optimization. When you ask "how can that optimization be legal without UB?", the hard part is "without UB" all by itself. If you have a language with UB, the optimization is easy to enable. If you have a language without UB, the optimization is easy to enable. That optimization is not an example of why we need UB.
There can be significant runtime overhead to remove UB, but it's not in service of enabling that optimization.
> Isn't QEMU both the machine and the compiler (EDIT: perhaps better-put, what's the distinction? They're both just the implementation.) The bytes only need to be unchanged on read from inside the program; if you compile them to something else, that's fine. Prolog programs can introspect with clause/2 and it doesn't cause trouble there.
Basically, I don't think "The bytes only need to be unchanged on read from inside the program" is true. If you're compiling machine code you're not in charge of the entire computer. The code might depend on other code looking at the bytes, and you won't be able to intercept that unless you do some kind of ridiculous rootkit takeover when the compiled program launches, creating a virtual machine and moving everything that was already running into it. And I don't just mean that in a theoretical gotcha sense, real libraries sometimes need to alter function calls in other libraries.
I believe simply disallowing "+" and "-" from accepting pointer types is slighlty less work than writing logic for "integral + pointer", "pointer + integral", "pointer - integral", and "pointer - pointer". Although on the other hand it means that instead of literally rewriting "a[i]" to "*(a+i)" in the AST, one has to write a separate logic for array accesses so... okay, so it's exactly as easy to compile ("as hard" technically works as well, but since it's actually rather easy, not hard, I've chosen "as easy", hope you don't mind).