What’s the difference between an integer and a pointer?
blog.regehr.org
blog.regehr.org
Edit: a good explanation of that macro is at http://lists.linuxcoding.com/kernel/2006-q3/msg17979.html :
"The reason for it is that gcc assumes that if you add something on to the address of a symbol, the resulting address is still inside the bounds of the symbol, and do optimizations based on that. The RELOC_HIDE macro is designed to prevent gcc knowing that the resulting pointer is obtained by adding an offset to the address of a symbol. As far as gcc knows, the resulting pointer could point to anything."
How can one C object span two symbols in the first place for this assumption to be invalid?
char symbolA[10];
char symbolB[10];
symbolA[15] = 0; // actually writes to symbolBRight - I'm saying that if they need this macro then their program is not a defined C program in the first place. They should fix that, rather than trying to use a macro to work around the problem. The macro may break in the future if GCC adds other optimisations.
They call it a 'miscompilation' - how can you 'miscompile' something which is undefined!
That GCC may break your program in the future is one of the ways that people get locked into specific versions.
Today, if you were starting a kernel project from scratch, I'd think Rust would deserve a good, hard look.
Obviously at the application level it's frowned upon, but in the kernel? You do what you have to with the tools you have.
Think about Java. Yea sure, keep telling us sun.misc.Unsafe isn't legit. It's going to get use until that capability exists elsewhere.
I'd rather my kernel not have a security hole in it because a compiler decided to optimize out undefined code: https://lwn.net/Articles/342330/
I appreciate using software that is actually written with security first, performance second in mind.
1) you have rightly pointed out all the authors of the Linux kernel are total idiots who don't understand C.
2) It turns out writing a kernel isn't possible in pure C, and the kernel requires some carefully designed and considered methods of accessing all of raw memory, in various unusual ways.
This one is the case. Writing a kernel is not possible in pure C. I'm asking - given this, why are they trying to do that? Write the parts which aren't expressible in C using assembly, where they can get the semantics they want.
While technically this means your kernel is "gcc c" rather than "true c", this doesn't cause a practical problem. I know gcc, clang and visual studio's c++ standard libraries require a flat address space and 8 bit bytes, for example, but I've never seen anyone bothered by this.
How about this as a compromise: use GCC, with all its flags, to generate the assembly, and just use that? Skip the messy transition step.
Instead we could just mark some files as "These C files have to be run with GCC versions X to Y", problem solved. Then why not just compile everything with those versions of GCC?
Is it solved? You'd still have to check the generated assembly of the entire file after every change you make, to make sure it didn't have some knock-on effect that optimised something differently, even if you kept the version of the compiler the same.
And is GCC deterministic? I don't know - do you? Did you know it has a flag `-frandom-seed`?
[Edit: I am almost sure that the only part of Linux on i386 and amd64 that is written in assembly is the bootloader stub at the beginning of "Linux kernel image" file format. And it is somewhat funny and relevant that this part depends on abusing gas to generate 16bit code]
The "correct" way to do this is to write the bare minimum assembly needed to present this view to a programmer in C (so there's no need to worry about undefined behavior), then write standards compliant code from there. Writing all this code in C is just papering over the fact that what you're doing is architecture-specific and should be treated as such.
It simply implies that they consider themselves, rather than the C standard, as the authority on how C should behave.
> gcc assumes that if you add something on to the address of a symbol, the resulting address is still inside the bounds of the symbol
I'm asking - how can you write a valid C program which adds something on to the address of a valid symbol and get something outside the bounds of the symbol?
I am guessing here, but my guess is that they're using C but actually want to do something very low level. As in they believe (correctly?) that they know the exact memory layout relative to something.
char x[N],*p=x+sizeof x;Reading or writing the pointer p will not be a valid C program.
You have a bunch of objects, defined in different files, and no code is aware of what the complete list of objects is. You want to iterate over an array of all these objects. How to make the array, with minimal overhead?
A common technique is for each file defining on the these objects to to tell the linker to put it in a special section. The result of this is that there's then an "array" of objects, which according to the C language could be anywhere in memory, but because of how the linker has been told to lay them out, we know that actually they're all consecutive. So any code which wants to iterate over then can do so easily.
Except... this is isn't actually a valid C program, and is actually undefined behaviour, so the compiler might "optimise" things in a way which breaks it. Hence the RELOC_HIDE macro which obscures things from the compiler and prevents it from applying "optimisations" we don't want.
Because Linux isn’t written in C. It’s written in C, assembler, and linker script (.lds). In both assembly and linker script, one can lay out multiple symbols with a defined relationship. Unfortunately, C doesn’t have a way to say “give me a pointer to the object 10 bytes past the address p. Yes, I know it’s valid and it is not the same object as p.”
Maybe we can one day get a `--std=simple-world` flag that forgoes a few optimizations and makes UB code do the "obvious" thing.
Its more like we had this window of time when pointers were not weird and got too used to it.
It's strange that we put so much effort into learning all this arcane knowledge and now we're supposed to unlearn it...
I also don’t understand a reason behind dragging all that legacy into current standards. Does a real application of it make at least 0.1% of all usage? You cannot even buy a chip that implements segments, tagged pointers, etc etc. at least they could make a special mode where you can span across symbol whatever it means or see a pointer as a cpu sees it.
All this can be solved with simple ptr_untag(p), ptr_span(expr) and similar constructs, but instead we have resort to outsmarting the specific compiler logic or introducing ourself with complicated type systems. That went insane. Personally I just want my bytes in linear address space and a way to tell which bytes point to which and which do not. I liked asm and then C, they were two of my first three languages, but what C became today is just a horrible mess.
Would you mind expanding on this point a bit?
If not that, I’m also curious.
Below the compiler, I'm pretty familiar how things are laid out, since there is ABI, packing/alignment rules, documented hardware and 15 years of dealing with them. Please don't mock up magic and wisdom out of few conventions.
At higher levels of abstraction we pretend like it's all flat but we do need to allow for optimizations to occur.
I would however argue that the existence of hierarchy in physical memory is orthogonal to the existence of a linear address space. Caches work really hard to maintain the ability of programs to use a single linear address space.
(Otoh virtual memory can be used to create the appearance of single address space for any of the above, and sometimes has been - eg as/400)
Not only can you buy a chip that implements segments, the computer you wrote that statement on probably has such a chip.
Edit: though strictly speaking I was obviously wrong on that, clarifications are welcome.
There’s no universe where traversing a page table (actually a tree) in memory is faster than an offset and a bounds check.
We could imagine more segment registers and bigger descriptor tables, but that would be just poor man's manual TLB (manual always failed in cpus). If you concerned with constant checks, you may order yourself a cpu with BOUND instruction support and put it everywhere with the same result. Oh, it is already in x86-64, nevermind.
>actually a tree
It is 2 level "tree", afair. Directory and table. Please make your homework and research why segmented model was kicked off software arena by virtually everyone involved.
You can use huge pages but it has all of the drawbacks of segments and none of the benefits.
VMWare used to have a super fast 32-bit hypervisor based on segments long before special instructions were added. This of course had to be reworked completely for X64.
Also Intel’s bound check instructions are still extremely rare and don’t work that well in practice. I’ve used them.
>VMWare 32-bit before special instructions
DOS also was pretty fast, but that didn’t make it a good multi-user protected-mode OS. All these early emulations and monkey patching of guests cannot substitute hardware vt in the wild.
One way to implement my idea would be as pragmas that tell the compiler "don't optimize away this null check", "assume this variable can alias", "just emit the obvious ASM for memory access" or "put memory barriers around this code".
I don't see how this kind of "stupid memory access" would make e.g. SROA impossible. If I just access the memory around "this", sure, I would not see the layout I'd naively expect. But that's not what I want anyway. I want to see the memory as it is actually laid out.
I do personally wish the C standard was changed a bit to have less undefined behavior, leaning towards getting programmer intent correct and not worrying (as much) about getting every possible optimization into the generated code.
Agreed. There are some particularly egregious ones where the behavior should at the very least be implementation defined.
TFA argues that converting to integers first sidesteps the issue but I'm not convinced (although I'm not a C standard lawyer). Surely if you convert a pointer to an integer, add an offset that's greater than the size of the original object and then convert that to a pointer you trigger UB? How else could this be defined? You can't really expect the compiler to figure out what you're doing when you're casting "random" integers as pointers.
Also the article mentions Rust in its introduction but in Rust anything surrounding pointers (including pointer arithmetics) is unsafe so I don't think there's any conflict whatsoever here.
https://internals.rust-lang.org/t/pointers-are-complicated-o...
"This summer, I will again work (amongst other things) on a “memory model” for Rust/MIR. However, before I can talk about the ideas I have for this year, I have to finally take the time and dispel the myth that “pointers are simple: they are just integers”. Both parts of this statement are false, at least in languages with unsafe features like Rust or C: Pointers are neither simple nor (just) integers.
I also want to define a piece of the memory model that has to be fixed before we can even talk about some of the more complex parts: Just what is the data that is stored in memory? It is organized in bytes, the minimal addressable unit and the smallest piece that can be accessed (at least on most platforms), but what are the possible values of a byte? Again, it turns out “it’s just an 8-bit integer” does not actually work as the answer."
Ralf Jung is a coauthor of the "What’s the difference between an integer and a pointer?" paper.
So, in order to generate faster programs, compilers are getting more and more agressive about reusing data that has already been copied into registers, hiding under the shadows of 'undefined behavior.' The standard doesn't indicate a defined way to mess with this data so I'll just keep using the copy I have.
Writing through an untracked pointer probably convinced the compiler to dump its register caches.
C used to be thought of as kind of a portable assembly language, but when the compiler is doing all these optimizations, there's a lot of room for clashing.
I’m a bit confused by this statement. I don’t know the C & C++ standards, but isn’t pointer arithmetic defined and allowed within the bounds of the heap allocation of the base pointer? Like, on an array for example, aren’t pointers required to “follow the same rules as pointer-sized integer values”, as long as you’re not indexing outside your array allocation?
Wouldn’t the issue here be more clearly stated as “don’t use pointer arithmetic to index outside the bounds of the base pointer’s allocation” rather than “never use pointer artithmetic”?
Dunno what it means for stack pointers though, probably the same except now you've made it so they can't be registers.
The fact that an Iterator, a pointer and an integer all tend to be different ways of looking at same integer value is more or less an implementation detail.
*x will be 7 in the end.
Mac:tmp $ clang -O3 mem1d.c ; ./a.out
diff = -96
x = 0x7f9f4fc00300, *x = 7
y = 0x7f9f4fc00360, *y = 5
p = 0x7f9f4fc00300, *p = 7
Mac:tmp $ diff mem1d.c mem2d.c
13c13
< int *p = (int *)(yi + diff);
---
> int *p = (int *)(yi - 96);
Mac:tmp $ clang -O3 mem2d.c ; ./a.out
diff = -96
x = 0x7fc2b5c00300, *x = 7
y = 0x7fc2b5c00360, *y = 5
p = 0x7fc2b5c00300, *p = 7
Mac:tmp$ clang --version
Apple LLVM version 10.0.0 (clang-1000.11.45.2)
Target: x86_64-apple-darwin17.7.0
Thread model: posixDoesn't this depends on which allocations have been freed?