Pointers in C (2010)
boredzo.org
boredzo.org
It claims that, given a declaration
int array[] = { 45, 67, 89 };
the expressions "array", "&array", and "&array[0]" are all equivalent. They are not. "&array" and "&array[0]" both refer to the same memory location, but they're of different types.In the next section:
"By the way, though sizeof(void) is illegal, void pointers are incremented or decremented by 1 byte."
Arithmetic on void pointers is a gcc-specific extension (also supported by some other compilers). It's a constraint violation in standard C.
I don't think this is the "best" article on pointers in C.
I usually recommend section 6 of the comp.lang.c FAQ, http://www.c-faq.com/
As for the array types, I think maybe the article is really trying to say that they'll all compile down to the same thing in the end (probably...compilers can be weird sometimes).
If I could time travel and influence the design of C the first things I'd do is change switch to break by default. The 2nd thing I'd do is rename "char" into "byte" since that's effectively what it is (and it's seldom a character these days since we often use UTF-8 or other multi-byte encodings).
This isn't always true. See machines where CHAR_BIT > 8.
In French when talking about storage capacity we use "octet" instead of "byte", I always thought that made more literal sense (if you have a 1megabyte memory on a system where CHAR_BITS is 16, do you have 8 or 16 megabits?).
4-bit CPUs (like the one in your toothbrush or thermometer) can also of course address 4-bit "words", "bytes" or nibbles.
So 8 bits might not be the smallest addressable memory unit.
As a side note, there was a debate whether casting to `uint8_t * ` is a strict-aliasing issue (As `uint8_t` and `char` are not guarenteed to be the same type, and so the compiler isn't required to treat them as such).
That said, you're correct that you can always cast to `char * ` in this case as `char * ` is allowed to alias anything. I do agree pointer arithematic on `void * ` a bit messy, I think the intention was just as a convenience thing, but it can definitely be abused.
> As for the array types, I think maybe the article is really trying to say that they'll all compile down to the same thing in the end (probably...compilers can be weird sometimes).
I donno, I've read this page before and think the author genuinely doesn't know about pointers to arrays - if they do they've done a good job of pretending they don't exist. They are not mentioned anywhere on the page (Besides the part that says they're are all the same), and near the top the author talks about parentheses in variable declarations and says:
> This is not useful for anything, except to declare function pointers (described later).
Which anybody who has used a pointer to an array will know is not true, as you have to use parentheses to differentiate between a pointer to an array from an array of pointers.
Remember that `char`, `signed char` and `unsigned char` are the different types, even though `char` takes the same range of values as either `signed char` or `unsigned char`.
The typedef `uint8_t` is usually set to `unsigned char`, not `char`, even on systems where `char` is unsigned. Partly this reflects the fact that `char` is usually used to represent actual characters, while the other two types are usually used to represent integers that take the same amount of memory as `char`. The standard technically does not allow `uint8_t` to be `char` [1], although this requires an extremely pedantic reading.
Anyway, if you replace `char` with `unsigned char` in your comment then it's correct. I believe all current major implementations typedef `uint8_t` to `unsigned char`, but that's not guaranteed and even old implementations of GCC had a different type. `unsigned char` satisfies the same relaxed aliasing rule as `char` [2] but `uint8_t` may not.
That can't be right, because CHAR_BIT is not always 8.
In contrast I really like the posts definition.
> A pointer is a memory address.
Not perfect but for concision and accuracy it cannot be beat.
ptr++;
it increments the address stored in ptr, not by 1, but by 8 (bytes) - so as to now point to the next struct in the array. Same if you do: ptr--;
except in that case it decrements the address by 8.And if you did:
ptr += 2;
it would increment the address in ptr to point to 16 bytes further ahead in memory than it was earlier, for the same reason. ptr will now point to the struct which is two items further ahead in the array. So you can access that struct with the expression: *ptrNot necessarily a strong belief, just an amount of ignorance perhaps.
Indeed. And isn't this what makes the classic "countof" macro possible? The macro returns the number of elements in an array, calculating this by dividing the size of the array by the size of the first element:
#define countof( array ) ( sizeof(array) / sizeof((array)[0]) )"&array[0]" also does a dereference, which may cause undefined behaviour if array is null or invalid. You can get fun stuff like the optimizer omitting null checks later on.
In particular, it supplies a useful algorithm for decoding all pointer declarations such as functions that return function pointers.
// +---------+
// | +--+ |
// | ^ | |
int /*|*/ aa[2][3];
// ^ | | |
// | +------+ |
// +--------------+
The correct result is "aa is a (2-element) array of a (3-element) array of ints".To get the correct interpretation, you have to know that the spiral has to avoid the "int" element after passing through "[2]". This defeats the purpose of the spiral, since the line sometimes goes through the element and sometimes it doesn't. For example if it were "int * paa[2][3]" instead, the correct order is { [2]; [3];* ; int }. Note how the star is first avoided and comes after [3]. A 2-array of a 3-array of pointer to int.
How would you know when the spiral "avoids" the element on the left and when it doesn't? Well, you need to know the declaration grammar to know that [] and () bind stronger than the thing on the left, so you need to process those first. But if you know this, drawing a spiral is redundant, because you already parsed the thing.
I think the spiral rule is inherently wrong and should not be reposted as a helpful cheat-sheet for parsing C/C++ declaration syntax.
The reason is that some types that are syntactically valid are forbidden by C/C++. You cannot have a function returning a function. A function returning an array. An array of functions. You can only have a pointer to these (function returning a pointer to function/array or an array of pointer to functions), and then they need to be correctly parenthesized. Then the spiral rule works because you only have two elements in a parenthesis and you can just go right and then go left... unless you have arrays of arrays which are legal, and then it doesn't work.
But the more correct rule would be to go right and parse all [] and () inside the current parenthesis level, then go parse the * -s (including const/volatile) on the left. Then repeat for outer parenthesis levels.
The simplest correct algorithm is to 1. From the identifier, go right one by one in the current parenthesis level and process every element. 2. From the identifier again, go left one by one in the current parenthesis level and process every element. 3. Repeat from step 1. except the "identifier" is the part you already processed.
Within step 1, you will encounter arrays and function parameter lists. Within step 2, you will encounter pointers (in C++ also references), const and volatile modifiers, and named types. Neither algorithm covers it, but if there is no identifier to start with, then you start at the most nested level between the elements that may occur in step 1 and 2 and if you find a comma, you were just processing the type of a function parameter.
Spiraling is unnecessary. Within step 1, you DON'T spiral because of arrays of arrays or you turn back because of a closing parenthesis so a spiral doesn't need to guide you. Within step 2, the elements on the right are all processed already, so a spiral going back right will hit nothing. For simple but common cases like "const int * ptr * const_ptr_to_const_int" you're spiraling around nothing on the right-hand side.
Let's rework and simplify the examples from the http://c-faq.com/decl/spiral.anderson.html page. Comments show what the result of parsing an element is. Process >> >> left to right and << << right to left.
char * str [10];
// XXX >>>> "str is", "an array of 10..."
// <<<< < XXX "pointer to", "char"
char * ( * fp ) ( int, float * );
// XX "fp is"
// < XX "a pointer to"
// XXXXXXXX >>>>>>>>>>>>>>>> "a function taking (int, float*) and returning"
// <<<< < XXXXXXXX "a pointer to", "char"
void (* signal (int, void (* fp) (int))) (int);
// < XX "fp is", "a pointer to"
// XXXXXX >>>>> "a function taking (int) and returning"
// <<<< XXXXXX "void", "and it is a function parameter to something else"
// XXXXXX >>>>>>>>>>>>>>>>>>>>>>>> "signal is", "a function taking (int, void(*fp)(int)) and returning"
// < XXXXXX "a pointer to"
// XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX >>>>> "a function taking (int) and returning"
// <<<< XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX "void"
const char * chptr;
// <<<<< <<<< < XXXXX "chptr is", "a pointer to", "char", "that is const"
// Notice how a spiral would make this worse and the page omits drawing it for a reason
char * const chptr;
// <<<< < <<<<< XXXXX "chptr is", "a constant", "pointer to", "char"
volatile char * const chptr;
// <<<<<<<< <<<< < <<<<< XXXXX "chptr is", "a constant", "pointer to", "char", "that is volatile"
Notice that in the examples where Anderson drew any spirals, it only made sense because there was a single "<<"/">>" element to the left and to the right of the XXXX part. In all other cases, spirals don't make sense because you might have to avoid the element on the left or there is no element on the right to spiral through, you're just going right-to-left. Why fit a spiral when a straight-line arrow will do?The declared identifier is, first and foremost, the clump of high precedence postfix things that are on it:
a[3][4][5]; // a is an array of array of array
b(int); // b is a function
c(int)[3]; // c is an array of functions: nonexistent
Then we consider the unaries: *** whatever; // whatever of pointer to pointer to pointer
Then the declaration specifiers: int whatever; // whatever of/to intTo apply the spiral rule in these cases, we have to collapse together the chained unary and postfix operators so that the spiral traverses them as one clump:
// +--------+
// | |
// | |
int *** aa[2][3][4];
// ^ | v |
// | | | |
// | | | |
// +----+ +----+
"aa is an (array of array of array) of (pointer to pointer to pointer) to int".The spiral only makes additional iterations when precedence parentheses are present. Each level of parentheses has its own postfix-unary round trip:
// +----------+
// | +----+ |
// | | | |
// | | | |
int **(* aa[2])[3][4];
// ^ | | v | |
// | | | | | |
// | | | | | |
// +----+ | +--+ |
// +--------+For me, all confusion about C's pointer-happiness cleared up when I finally realized that C (and Asm, I guess) works with heap memory as a big blob of bytes, and it's programmer's job to keep the blob's contents from getting messed up―with some thinly-veiled help from the language and the compiler. Everything else, including variables, is just syntactic sugar when it points to the heap.
(With the clarification that afaik variables are different when they're on the stack or registers, becoming first-class 'indivisible' entities).
char theLetterC = "ABC"[2];
char theLetterC = *("ABC" + 2);
char theLetterC = *(2 + "ABC");
char theLetterC = 2["ABC"];
"You can't prove anything about a program written in C or FORTRAN. It's really just Peek and Poke with some syntactic sugar." -Bill JoyNow that's the words I haven't heard in a long time.
char theLetterC = 2["ABC"];
Wait, what? And I thought I was good with pointers...Edit: I think I got it, is it this?:
"ABC" + 2 == 2["ABC"] == *(2 + "ABC")
Right?
???
In Objective C, that would be [2 getNthCharacterOfCString: "ABC"];
Just joking! If you try to think about C in object oriented terms, you're misunderstanding what's really going on. It doesn't actually have arrays or strings as objects, they're just syntactic sugar (more like syntactic syrup of ipecac), and [] doesn't send a message. It's all just pointer arithmetic.
arr[i] (i'th element of array arr)
is the same as *(arr + i) (the value at the memory address arr + i)
which is the same (by commutativity) as *(i + arr) (the value at the memory address i + arr)
which in turn is the same as i[arr]
That last bit is what seems non-intuitive, but it is true.I first read about this, maybe in the K&R book or some other C book, many years ago, at first didn't believe it, and remember trying it out on at least a Windows C compiler (MS C, likely) and maybe GCC on Linux as well. Worked on both, no compiler error.
This works because the expression arr, while we think of it as the name of the declared array, is also a synonym for the address of the first memory location of the array.
Note: all that I wrote above is what I rememeber from using ANSI C (a lot, but some years earlier). Things may have changed some with C99, etc.
remember
It's why you can have stack overflow, heap overflow, etc via pointer bugs.
Also, C/C++/etc runtimes don't have garbage collection but that's a runtime issue not a language specific issue/type system issue.
I think you have a good understanding of pointers, but you are mistakenly confining it to the heap or the runtime. Pointers part of the language and the type system.
If the learner's mental model of what is happening has any flaws whatsoever, it will break under that strain. There are a few complicated concepts that are very similar but not quite the same.
I have no memory of pointers being a hurdle like that. But I guess the concepts are somewhat similar so there probably was a time like that with pointers also.
For me, everything became much more clear when I realized that the heap memory stands on its own as a concept, being a big pile of bytes instead of ‘boxes’—and implicit allocation/deallocation of variables is just a thin veil on top (muddying the matter somewhat since vars can themselves be stored in memory and point to it).
Guess I had to deal with OOP before I grokked all the pointer-juggling going on behind the scenes.
I guess it is a problem for those who C is the first language where they see pointers in action.
I used to teach C, found pointers a bit tough, then understood them after playing around a bit. General teachers of C indeed didn't have any real world exposure to programming.
https://unix.stackexchange.com/questions/11402/why-does-esc-...
>Why does `ESC` move the cursor back in vim?
>In insert mode, the cursor is between characters, or before the first or after the last character. In normal mode, the cursor is over a character (newlines are not characters for this purpose). This is somewhat unusual: most editors always put the cursor between characters, and have most commands act on the character after (not, strictly speaking, under) the cursor. This is perhaps partly due to the fact that before GUIs, text terminals always showed the cursor on a character (underline or block, perhaps blinking). This abstraction fails in insert mode because that requires one more position (posts vs fences).
>Switching between modes has to move the cursor by a half-character, so to speak. The i command moves left, to put the cursor before the character it was over. The a command moves right. Going out of insert mode (by pressing Esc) moves the cursor left if possible (if it's at the beginning of the line, it's moved right instead).
>I suppose the Esc behavior sort of makes sense. Often, you're typing at the end of the line, and there Esc can only go left. So the general behavior is the most common behavior.
>Think of the character under the cursor as the last interesting character, and of the insert command as a. You can repeat a Esc without moving the cursor, except that you'll be bumped one position right if you start at the beginning of a non-empty line.
Also:
https://superuser.com/questions/242156/make-vim-normal-mode-...
>Make VIM normal-mode cursor sit between characters instead of on them
>I would really like it if the VIM cursor in normal mode could act like it does in insert mode: a line between two characters. So for example:
>- Typing vd would have no effect because nothing was selected
>- p and P would be the same
>- i and a would be the same
>Has anything like this been done? I haven't been able to find it.
>Answers:
>The idea that the cursor is always on a line and on a character position or column is inherent in Vim's design. If you were to try to change that, many of Vim's operations would behave differently or would not work at all. It's not a good idea. My advice would be that you learn and become accustomed to Vim's basic behavior and not try to make it behave like some other editor. – garyjohn Feb 5 '11 at 23:55
>What you want is not Vim, I'm afraid. – romainl Feb 6 '11 at 7:15
And people wonder why I still use Emacs...
That is, this would be for the purposes of visual selection with 'v'. It now includes the character under the cursor, which makes no sense when the cursor is a vertical bar that is to the left of the character.
The code is so contorted and poorly modularized, it would have to change in numerous places; I estimated it at a minimum of two weeks of full time work to ramp up on the internals sufficiently to be able to add the feature with reasonable confidence.
We do int ptr instead of int ptr and talk about memory instead of just saying: "It is simply a box containing 1 value, we can also make a box containing a box containing the value."
Luckily, I don't do C.
Normally we see such guiding images with 2d boxes. There's nothing wrong with 3d because clearly, integers are also not 2d (they are not even 1d). However, don't associate the size of theses boxes with the memory size of an element. This overstresses the analogy and suggests the whole thing is a vector space, but it isn't.
everything?
var lowercaseIds = ids.Select(ToLower) var lowercaseIds = ids.stream().map(String::toLowerCase);
and an additional .collect(Collectors.toList())
if you don't want a Stream instance. (var is Java 10, though I believe it should not be used in production code, ever)Is that your opinion of var in general, or something about Java's implementation of it? If the former, I'm really curious why, as a C# dev who's used it for many years.
But I am a c# dev, and in the early days when var was introduced a lot of people avoided var because of misunderstanding of how it works. People thought it was dynamic rather than inferred, is it the same here?
See my reason for not using it above.
As someone who holds no mastery over C#: Is it appropriate to ever use dynamic anywhere besides horrifying interop code?
Sometimes I explore source code on GitHub, excessive use of var makes reading it very uncomfortable.
Turning it around, what are the benefits of var? Slightly increased typing speed. More time is spent reading source code than writing it and var makes reading harder and writing easier, so it's not worth the trade in my opinion.
RE your sibling comment about dynamic: I primarily use it with json. It's useful for things like one-off error responses where you just need to grab a message or code to return/throw and there's really no benefit from introducing a class for that one usage. I did once work on a project where that was attempting to re-use some awful legacy code in a new app, and they leaned on dynamic a lot to make the legacy code work without to re-architect it correctly. It was about as terrible as you're imagining.