Careful: In C, memory isn't a huge array of bytes.
Careful: In C, memory isn't a huge array of bytes.
The physical memory isn't necessarily contiguous, but the model of memory is. Which is the entire point. arr[0] and arr[999] are 1000 virtually contiguous blocks.
You can confirm that in the spec here: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf
> it'll also give you something random should you go out of bounds.
Out-of-bounds array accesses are UB
> it'll also give you something random should you go out of bounds.
Are you just taking umbrage with the word I used? "random" clearly is meant to imply that the language makes no guarantees.
Uninitialized memory/data isn't "random", it's specifically null. Random could be something uninitialized, or it could be the kernel entrypoint or it could be a pixel data byte from some BMP.
Ooh that's very untrue [0].
Although, I guess I should make sure we're talking about the same thing. I'm saying this is undefined (the little program I've linked below):
#include <stdio.h>
int main() {
int a[4] = {0, 1, 2, 3};
int b = a[1000]; /* undefined behavior! */
printf("%d\n", b);
}
A lot of people think b should be something random (uninitialized memory), or NULL, but literally anything can happen: those options, your printer spits up all its ink onto the ceiling, you read data from a sensor on an embedded board, you read data from a different process' memory, you get a segfault, etc. It's dependent on how your OS (or whatever) manages memory beneath you. In this case w/ these compiler options and such, it's uninitialized memory.Suffice to say, my original statement clearly matches what you seem to be getting at; whether you like the choice of words or not.
But, maybe I've been too much of a jerk here! Maybe I jumped too hard into "someone's wrong on the internet!" mode. It's OK to be wrong, cool and rock and roll even. Lord knows I've been wrong and will be wrong again. If this is true then I've embarrassed myself and I apologize. Let's go forward and be excellent to each other.
But C lets you treat memory like a huge array of bytes. You can do things like:
char *p = (char*)0;
char value = p[address];
This will segfault if you touch the wrong address. Other than that minor detail, you can think of memory as a huge array of bytes.Of course it is; the platform you are used to might have additional constraints, but that is outside of C itself.
On your platform (hardware/OS combination) it may not not be an array of bytes, but rest assured that conforming C compilers exist on platforms, such as embedded ones, where you can literally get a pointer to address zero, and use all the RAM as a huge array of bytes.
In other words, location 0x0000ffff in your application does not map to system location 0x0000ffff, but instead a translated portion of said block. In addition, there are no guarantees as to how that memory will be ordered/allocated/segmented outside of specific requests for a contiguous block of memory via something like malloc. You can assume your array (static and dynamic) is contiguous, but that's the only assumption you can make.
If you were to write freestanding C code, this certainly changes. At which point the memory model becomes "what you decide to provide".
That doesn't make sense; the OS itself is written in C.
> If you were to write freestanding C code, this certainly changes. At which point the memory model becomes "what you decide to provide".
Okay, so we're back to, C doesn't impose anything on you, though the OS may.
Please point to a non-userspace memory allocation library/implementation that does not rely on lower level logic to function.
> Okay, so we're back to, C doesn't impose anything on you, though the OS may.
You're just reversing the original statement.
The statement was "you can't assume anything about memory in C" (paraphrasing). They then asked "why not?"
The explanation is that:
What you think is 'memory' in C isn't, and certainly doesn't map to what most people assume about memory; because C doesn't impose a memory model, it relies on the underlying OS/environment to do so. The only "memory model" is thus: a requested allocation (whether on the stack or heap), if provided to you, will match your request; all other assumptions are invalid.
If you want to move goalposts, argue about logical boundaries, etc, have at it. Or if that answer doesn't satisfy you, I don't know what to tell you; but simply rephrasing the original problem does even less.
[0] https://wiki.osdev.org/Global_Descriptor_Table
[1] https://wiki.osdev.org/GDT_Tutorial#Telling_the_CPU_Where_th...
> In addition, there are no guarantees as to how that memory will be ordered/allocated/segmented outside of specific requests for a contiguous block of memory via something like malloc. You can assume your array (static and dynamic) is contiguous...
The standard can only be referring to the abstract machine. In truth, your OS might give you 1,000,000 elements on 1 page, and 1,000,000 elements on a different page, which exist nowhere near each other in RAM (or have been swapped to disk, etc. etc.), or are being CoW'd into existence, and so on.
This is empirically true--from the days when programs like Chromium would try and malloc all the memory in a machine. The pointer that malloc returned could not have referred to a contiguous memory block. You can try it on your machine by malloc'ing more memory than you have and then reading from the blob.
[0]: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3054.pdf#s...
This is where the "model" portion of the statement comes into play and why OPs point is even more cogent. The memory will be virtually contiguous, even if it's physically disparate.
I don't really understand the value of extolling the fact that like, you can incrementally iterate through an array in C. The days of systems with segmented memory are long behind us, and even then it wasn't like you'd have an array that spanned segments--you couldn't! Honestly what language/platform exists that doesn't have this property and what would that even look like? Like you'd somehow have to know that elements 100-200 in an array are "no good" and you have to skip them?
People keep trying to put meat on the very basic bones you've got here, but you keep insisting that, yep, 2 comes after 1 and 717 comes after 716. Great! We know! And we say stuff about how trying to index off the end of an array leads to UB, or how the array may not actually be contiguous in physical memory, or it might be in various caches, or how sometimes your array goes from 1024 members to 1025 members and your FPS drops from 300 to 30 because you fell out of cache, and you're like, "sure but you still access arrays with consecutive indexes". Well, yeah! It's kind of the point of using C that you have access to or some control over these kinds of things. They're useful to the conversation. Continually bringing us back to indexing... I don't think is.
> This is where the "model" portion of the statement comes into play and why OPs point is even more cogent.
Eh, "memory model" is a specific phrase referring to how memory is defined to work in a threaded environment [0]. The original "Help us understand: what is the model of memory in C?" prompt is referring to the fact that unless you prolifically use the API in stdatomic.h (which very few things do), you just have undefined behavior all over the place if you ever dare to use anything related to threads. Also, it was only defined in C11--not a lot of things have updated, even now.
---
Overall I want to emphasize that this whole thread is doing a real good job of proving Ned Batchelder's point: even people who think they know or understand C don't (to be clear, I do not think I understand C), and things are generally alright. There's--clearly, reading through everything--a culture of "You need to be an expert in some low-level, 'real' language/platform before you can write meaningful software", but you don't, and my evidence is Facebook, maybe the most influential software ever written.
But I think the OP’s point is more about how when you allocate memory, you get an array of bytes to play with. And all higher level languages build on top of that and abstract it away as much as they can.
Will I get arrested? No.
Will the compiler stop me? Also no.
Will the program crash? Maybe. Almost certainly if I do it often, or without understanding.
You aren't guaranteed to be safe if you access memory or addresses outside of allocations you've made (with stack and static memory counting as "allocations you've made).
But on embedded systems with memory-mapped I/O, I have done things like
*(unsigned long *)0xFFFE1404 = 0x00011472;
in order to write values to the registers of a peripheral device. Those I/O registers were memory that I "owned", even though I never allocated it in any way.But of course an implementation is free to define additional behaviors beyond the C specification. That’s done all the time. But that’s really a “flavor” of C and not pure “vanilla” C.
An integer may be converted to any pointer type. Except as previously specified, the result is implementation-defined, might not be correctly aligned, might not point to an entity of the referenced type, and might be a trap representation.
Any pointer type may be converted to an integer type. Except as previously specified, the result is implementation-defined. If the result cannot be represented in the integer type, the behavior is undefined. The result need not be in the range of values of any integer type.
Also, the C standard merely codified existing practices and common extensions. Actual use of C has converted integers to pointers for a long time. If converting integer literals into pointers were undefined behavior, it would just show that the C standard isn't being practically useful in one area (since it's commonly done in practice).
Quite probably since the first C-based implementation of Unix on a PDP-11. So it's been known to be "a thing that C does" for quite a bit longer than the standard existed.
Say I'm doing my original example, working on an embedded system. My code isn't going to port to anything that doesn't have the same hardware, so architecture isn't an issue. Any compiler supporting that architecture for an embedded application is going to do the right thing with that kind of C statement, so that isn't an issue either.
So, implementation-defined means that you can't count on compilers doing the same thing. But there are some things that are implementation-defined where you can pretty much count on any compiler doing the same thing. And, as trealira said, compilers are supposed to clearly state what they do in such cases, so you can read the compiler's statement and see if there are any surprises.
This is less true if you're writing library code. There, you have to support all compilers, or at least all conforming ones, and you have to make fewer assumptions.
https://en.wikipedia.org/wiki/Far_pointer
https://en.wikipedia.org/wiki/X86_memory_models
However, this is not standard C.
If you want something constructive, what was the point of the post if you're wrapping it up with "learning C can be useful"? I love C, it's fantastic and it has taught me a lot about how and why some things are as they are. Just compiling with a simple lib helped me understand headers and the -devel packages some distros provide. The difficulties doing so efficiently shows why newer languages have package managers. Another point could be that the entire world runs on C, and knowing it can help either porting software safely or maintaining projects.
Basically what I get from it is: "you (probably) don't need a drivers license, unless you want to drive, then get one". Sometimes the journey can be extremely beneficial even if there's no direct use. Would you claim learning C hindered you in any way?
The point of the post is that abstractions are inevitable and you choose your own level. And: you can be a great programmer without knowing C.
Absolutely, but I get the feeling you're taking the common statement "learning C can be a great experience" and bastardazing it into a straw-man "you NEED to learn C" that you then argue against.
I just think it's an obvious statement and it reads like you're being nagged by people to learn, but you don't want to learn C and defend your position. For me nothing has taught me so much about development than a weekend hacking in C and stepping through the code with Valgrind or reading the binary output did. And I recommend all my peers do the same, most don't and that's okay. But I think it's a shame that people in general don't seem to care that much outside of work.
I did enjoy the read though, even if clickbaity.
Why C as opposed to other languages? It's portable and very simple on the surface, there's not much to it but difficult enough to be a teachable opportunity.