C uses "&" for the address-operator because 'ampersand sounds like "address"'
softwareengineering.stackexchange.com
softwareengineering.stackexchange.com
When C has added structures, which did not exist in B, it has taken the keyword "struct" and also both "." and "->" from the IBM PL/I language, from which C has also taken some other features.
In general, almost any feature added by C to B was taken either from PL/I or from Algol 68. The exceptions are "continue" and the generalized "for", which did not exist in any previous language.
(However the generalized "for" of C was a mistake, because it complicates the frequent use cases in order to simplify seldom encountered use cases. The right way to generalize "for", i.e. with iterators, was introduced by Alphard in the same year with C, i.e. in 1974.) (Compare "for (I=0;I<N;I+=5) {" of C with "for I from 0 to N by 5 do" or "for I in 0:N:5 do" of previous languages. C requires typing a lot of redundant characters in the most frequent cases.)
The oldest symbol for indirection through a pointer (in the language Euler, in January 1966) was a raised middle dot (i.e. a point). This was before ASCII and ASCII did not include the raised middle dot (U+00B7), so it was replaced by the most similar ASCII character, "*".
Euler had used "@" for "address of" and indirection was a postfix operator, as it should. Making "*" a prefix operator in B and C was a mistake, which forced the importing of "->" from PL/I, to avoid an excessive number of parentheses. Otherwise "(*x).y" would have been needed, instead of "x->y". With a postfix "*", that would have been "x*.y", and "->", would not have been needed.
In CPL, the ancestor of BCPL and B, indirection was implicit, like with the C++ references. Instead of having an "address of" operator, CPL had a distinct symbol for an assignment variant that assigns the address of a variable, instead of assigning its value.
I mean technically the compiler could just have been less stupid and made `.` auto-deref, it’s not like C has operator overloading so the LHS is either a struct or a pointer, not both.
It is frequent to have multiple "*", "[]", "." and "->".
Now you have the simple rule of executing first the postfix operators from left to right, then the prefix operators from right to left.
If some special rules had to be invented to avoid writing the parentheses in "(*x).y", after adding several "*", "[]" and "." it would become impossible to understand the right evaluation order.
When all of "*", "[]" and "." are postfix, they are just executed in order from left to right, which is easy to understand.
While sometimes auto-dereferencing a pointer would be convenient, in C you want frequently to get the value of a link, or to do pointer arithmetic. With auto-dereferencing, some new ways of using an "address of" operator would have to be invented, which are unlikely to be more simple than the current rules.
Like I have said, the 2 languages that have introduced the concept of pointers were Euler, with explicit dereferencing, and CPL, with implicit auto-dereferencing.
Both are valid ways to design a programming language, but in both case a lot of rules must be added to cover all the use cases, which would have to be more complex in the languages with auto-dereferencing, so these languages usually solve this problem lazily, by just prohibiting some uses, e.g. by prohibiting pointer arithmetic.
As I have written above, a language with implicit dereferencing needs additional syntactic means for denoting the cases when the values of the pointers are needed, e.g. for doing pointer arithmetic, which is frequent in C.
If dereferencing would have been made implicit, that would have required a large number of changes in the language. C++ has introduced pointers with implicit dereferencing, i.e. what C++ calls references, but due to the fact that the other syntax changes that are needed have not been made, because they would break compatibility, the C++ references can replace C pointers only in a subset of their uses.
You're misunderstanding the comment. They aren't asking for implicit derefs. They're asking for explicit derefs, using * or . instead of * or ->
Which is ironic when discussing overloading!
Excuse my confusion, this possibility didn't occur to me.
And I think escaping emphasis was only introduced somewhat recently? I do remember that for the longest time you basically had to trick HN into not breaking your comments by using a different character in stead.
*
for italics, which sometimes causes a comment to go haywire because someone uses it to reference a footnote or otherwise drops it in mid-sentence, with a matching closing one, not intending italics.(Up-thread I think it probably was like that when they commented, but has since been fixed.)
> "a.b" should be enough for pointers as well,
I have interpreted this to mean implicit deref, as there is no "*" (could be a formatting problem).
The proposal is literally: "If you see . and the left operand is a pointer, pretend the . was -> instead, because otherwise the code is already invalid."
However, I assume that this would have been a too complex solution for compilers that had to work in a few tens of kilobytes of memory, while a postfix "*", as already used a decade before C, would have been a trivial solution.
It's literally the same complexity as the type checking compilers already do to tell you that your `.` does not work because the LHS is a pointer not a struct.
To determine if implicit dereferencing may be applied, more analysis has to be done, because there it may be not only a pointer to a structure, but a pointer to a pointer to a structure and so on, so multiple implicit dereferencing may be needed to obtain a valid expression.
However I agree that the difference in complexity is not big.
And no, this would not perform multiple levels of dereferencing anymore than -> does today. You could have literally find&replaced every use of -> with . and every C program would have had the exact same semantics. `struct point **a; a.x = 1` would throw the exact same compilation error that `struct point **a; a->x = 1` throws today. The only difference would be that `struct point *a; a.x = 1;` would write 1 to the field x of the object pointed to by a, instead of throwing an error that says "object of type struct point* has no field named x".
No it doesn’t. Let’s take this example in valid C:
struct my_struct { int field };
struct my_struct s = { .field = 42 };
struct my_struct *p = &s;
printf("direct : %d", s.field);
printf("pointer: %d", p->field);
printf("address: %x", p);
What is being proposed here is to make that code valid: struct my_struct { int field };
struct my_struct s = { .field = 42 };
struct my_struct *p = &s;
printf("direct : %d", s.field);
printf("pointer: %d", p.field); // note the use of a dot here
printf("address: %x", p);
That is, the naked p is still to be interpreted as what it is: a pointer. It’s just that when we write `a.b`, the language would first check the type of `a`, then dereference it as many times as necessary to get to the underlying struct, and then access its field. For instance: struct my_struct { int field };
struct my_struct s = { .field = 42 };
struct my_struct *p = &s;
struct my_struct **pp = &p;
struct my_struct ***ppp = &pp;
Now let’s see how this automatic indirection would work: // All would print the same value
printf("%d", ppp.field);
printf("%d", pp .field);
printf("%d", p .field);
printf("%d", s .field);
// We can still use explicit indirections
printf("%d", (*ppp ).field);
printf("%d", (**ppp ).field);
printf("%d", (***ppp).field);
printf("%d", (*pp ).field);
printf("%d", (**pp ).field);
printf("%d", (*p ).field);
We can still get to the actual addresses no problem: printf("p: %x", p);
printf("p: %x", *pp);
printf("p: %x", **ppp);
printf("pp: %x", pp);
printf("pp: %x", *ppp);
printf("ppp: %x", ppp);
Note that we can play with the & operator too. In valid C we can do this already: printf("s.field: %d", s.field);
printf("s.field: %d", (*&s).field);
printf("s.field: %d", (**&&s).field);
printf("s.field: %d", (***&&&s).field);
With automatic indirection the following would be valid too: printf("s.field: %d", (&s).field);
printf("s.field: %d", (&&s).field);
printf("s.field: %d", (&&&s).field);
---The kicker here is that the decision on whether an access to a struct member requires dereferencing the pointer or not, is not done at parsing time. It’s done at type checking time. And by the way, in standard C the decision to give you an error or not is already done at type checking time. All this to say, this would be a fairly benign change to compilers.
Now would users get confused? Possibly. With the conflation of pointers and arrays, the following would be equivalent:
array.field
(*array).field
array[0].field
Looks nifty to some perhaps, but some people really meant: array[i].field
and forgot to write the index.Since in all situations where pointer->field is valid pointer.valid is an explicit error in C today, this would have been very much doable without major changes in compiler or language implementation.
Of course, at C's level of abstraction, and given the speed limitations of the day, the syntax difference between -> and . may well be argued to have been helpful instead of harmful.
And yes, the same could not be said for C++, where you can implement dereferencing for your own type and get a variable where both `a.b` and `a->b` are valid and have different meanings. This is anyway not a proposal for changing C (or C++) today, merely a "what if" discussion about how C could have been designed differently.
x = a->b->c->d->e;
...is a pointer-hunting-nightmare / potential cache-miss-galore. x = a.b.c.d.e;
...no problem, the whole chain is resolved at compile time into a single offset and results in at most one memory access. x = a.b.c->d.e;
...makes it immediately clear that c->d is a pointer access and everything else isn't.PS: ...and of course C++ messed this simple rule up with the introduction of references.
That's why I would have used a postfix @ for dereferencing.
I’ll entertain this when C fixes (or at least removes) integer promotion, which is a source of far more bugs and misunderstandings than this could ever be: `.` auto-deref-ing does not actively undermine what little type system C has.
You assert that you do not want this on grounds of clarity. It is not an unrelated problem to point out that C has numerous obscurity issues significantly worse than this could ever be in the language right now.
> Implicit operator overload is just a bad design.
That's at best a bunch of words arranged nonsensically, and at worst an assertion that you want to remove arithmetic operators from the language?
> Integer promotion has advantages and disadvantages, especially in a language like C where you often use 8bit or 16bit types for memory optimization/alignment reasons.
It's mostly a major actual source of obscurity and bugs.
1)Implicit operator overload is a bad design in case of pointers and objects because a.b and a->b are both readable, easy to type and short.
2)Integer promotion is not as simple because requiring explicit promotion makes a holy unreadable mess out of your code in the simplest of cases.
2) is demonstrated many times by mongo code lines in Java with all the explicit casts. This is also the reason Python has implicit promotion for its number type. You sacrifice something to get something unlike in 1) where you only lose for no gain what so over.
You're just not being honest.
> 1)Implicit operator overload is a bad design in case of pointers and objects because a.b and a->b are both readable, easy to type and short.
That's just a baseless assertion. Here I can do the same: implicit overloading is a great design in case of pointers and objects because a.b is always unambiguous and uniform, and -> is a harder to read and type extra operator which has no justification.
And then obviously your assertion can be used exactly the same way to similarly assert that every numeric type should have its own set of arithmetic operators, after all u4+ is also readable, easy to type, and short.
> 2)Integer promotion is not as simple because requiring explicit promotion makes a holy unreadable mess out of your code in the simplest of cases.
Integer promotion is much simpler because it's literally a never-ending source of bugs.
> This is also the reason Python has implicit promotion for its number type.
Python does not have implicit promotion for its number type, at best Python 2 had that for performance reasons (definitely not "mongo code lines with all the explicit casts" which it could not care less about), and in reality that was not even the case because it does not corrupt your data upfront as C does.
Oh let me assure you, I have no actual expectation that C would ever change in such a way.
Although it should be noted that both suggestions are entirely syntactic and could be gated behind a per-file (or even per function) stricture a la strict mode.
Um, "for (I=0;I<N;I+=5) {" is fewer characters than "for I from 0 to N by 5 do".
The syntactically equivalent, but non-verbose, is "for I ∈ 0:N:5 ⟨", which beats easily C.
for i 0:N:5 {
// loop body
}
By the way, the same should have been done with if and while: optional parentheses, mandatory braces: if a == b {
// code
}
while idx < end {
// code
}Still... when did that syntax come out? Was it in Algol 68, or PL/I? It seems a bit unfair to complain about C being verbose if it was shorter than anything else available at the time.
When later implemented, including in an Algol 68 variant that inspired the UNIX Bourne shell, "∈" was replaced with "in", which could be written with ASCII.
Algol 68 could use either keyword pairs, like "do" and "od" (which became "do" and "done" in the UNIX Bourne shell) or, optionally, various kinds of parentheses in their place, for conciseness.
The notation "0:N:5" or "0:N", when the step is 1, for arithmetic progressions is ancient. Even Fortran used it, but with the mistake of using comma instead of colon, which introduced a syntactic ambiguity, because comma was also used for other purposes. Using colon started with Algol 60, but that one preferred keywords instead of symbols in the arithmetic progressions used inside "for", so it did not use the notation consistently.
Fortran in 1954 or the language of Heinz Rutishauser, in 1951 (the first one with a "for" statement), were already much more concise than C.
for (I=0;I<N;I+=1)
versus
for I from 0 to N do
The argument “is more keystrokes” in this context is beyound what i can comprehend.
Arguably that syntax would have been prone to a lot of unintended pointer multiplication due to syntax error.
The worst offender is the quasi-obligatory "break" statement in "case" blocks. Fall-through cases are useful for things like lexers but probably not needed in 99.999% of other programs. I wonder how many millions of (wo)man hours have been wasted on debugging missing breaks. (yes, I know that linters exist)
-Wimplicit-fallthrough
From the compiler docs. It's just not a problem. If you waste time once and learn about warning options in the compiler you will be better off anyway.
It's like with assignment being an expression panic - something that maybe bites you once and then you learn to read the compiler warnings and it never causes problems again.
If you don't want to read then there is: -Werror
Available.Pity the (wo)man who does not use -Wall -Wextra for (s)he is truly mistaken
https://www.godbolt.org/z/99ejvY9Pc
It requires the separate option -Wimplicit-fallthrough:
https://www.godbolt.org/z/GP576zG9h
(which tbh sounds like a bug, because usually Clang tries to emulate GCC behaviour)
https://en.cppreference.com/w/cpp/language/attributes/fallth...
(For backwards compatibility, it still must fallthrough with or without the attribute; the attribute just signals programmer intention and silences the warning.)
GCC has a warning in the '-Wextra' warning set, Clang requires the explicit option `-Wimplicit-fallthrough', MSVC is completely silent (apparently it's in the CppCoreCheck rules though).
This should really be in the default warning set with an annotation that the fallthrough is intended (unless the case-branch is completely empty).
This was (almost certainly) done to simplify the compiler. CASE in C is not actually a structured control construct despite its syntactic appearance, it's just a computed GOTO. This is what makes things like Duff's device [1] possible.
My mental picture of what C-compilers are up to turned out to be a lot more sophisticated than reality.
But having written a few Forth-interpreters, it makes total sense to me. Treating the input as a mostly unstructured stream of tokens is very convenient.
Choices/options for the same result, tend to make code less readable.
case 1:
case 2:
... code ...
the fallthrough is allowed. case 1:
... code ...
case 2:
... code ...
Is an error. But `goto case` can be used: case 1:
... code ...
goto case;
case 2:
... code ...
which makes it clear. No new keywords are needed. C could adopt this easily. All those /* fallthrough */ comments and warnings and compiler switches will just go away.Oh, and obviously:
case 1:
... code ...
goto default;
default:
... code ...
also works!BCPL has infix ! as well as prefix ! so you write an array index expression like array!index instead of array[index]. Infix ! is commutative like addition and [] in C.
You declare structures in BCPL by defining a constant for the offset of each member, so you can write object!MEMBER somewhat like C object->member. The semantics of -> in early C were very similar to BCPL infix ! except that members had types as well as offsets, but like BCPL there was nothing to tie a particular member to a particular structure as there is in modern C.
There’s a certain elegance to BCPL’s syntax that you don’t get from prefix-only or postfix-only indirection operators. C might have been better if it had stuck closer to BCPL in this respect, but sadly * conflicts with infix multiplication.
Some readers might also have used BCPL-style syntax for WIMP programming in BBC BASIC on RISC OS.
I don't think it does. There's no ambiguity in parsing any expression that uses it in both se se.
It does conflict with division, though, because /* starts a comment!
Really? How do you do the C's equivalent of `for(i = 0; i < N && j < M; i++, j++)` with your `for I from 0 to N` syntax? The for loop in C is extremely flexible and it captures the idea of the loop perfectly: there is the initialization block, the exit condition and the iteration. In other languages at that time the for loop was written in terms of a range, but without a strong range abstraction such loop is really primitive and limited in applicability.
That's a good thing: it gives you a clear syntactic marker for "this loop definitely terminates". There's already a full power anything-goes looping construct: `while`.
while(int i = 0; i < 100) { ... }
as a valid syntax to C++ (which would restrict the scope of i to the loop), as you can already do if(int i = f(); i > 0) { ... }
but it didn't go anywhere. for (int i = 0; i < 100; ) { ... }That is, the commenter was not asserting the supremacy of "from 0 to N" as a for-loop construct, but rather asserting the supremacy of iterators, and pointing out how poor the C-style syntax is even compared to its less-generalised cousins.
Even today, such simple loops with arithmetic progressions include an overwhelming majority of all "for" loops.
The C syntax made writing these simple "for" loops more difficult, by having to write a lot of redundant symbols, instead of writing the minimum number of separators. Also for reading, the redundant symbols obscure the meaningful text.
The C syntax allows the writing of "for" loops where the values of the control variables are not taken from a progression, but they are for instance the values of the links of a linked list.
This kind of "for" loops can be written in a simple way by using iterators, without complicating the syntax of the loops with simple arithmetic progressions.
In modern C++ there is no longer any case when you would want to use the C kind of "for", but the kinds of loops used by languages like C++ already existed in languages introduced at about the same time with C, e.g. Alphard and Clu.
for (bit = 1; bit <= 128; bit <<= 1)
Personally, I like the flexibility.For example, if the separator for geometric progressions would be ":>", your loop example would become "for bit in 1:>128:>2". Another example of (non-ASCII) separator would be "for bit in 1⋮128⋮2"
For example, it is normal practice to do some of initialization outside of a loop. It is normal to deal with finishing loop by using goto or if-branching after the loop (think of a search, that can be succesful or not). It is normal to ignore update counter part of loop and do updates in a body of the loop, because either updates are too big to squeese them between into for, or you need to do something between updating your counter and checking for the necessity of running another iteration.
With all this said I know 2 another approaches to the problem.
1. Common Lisp approach: create really flexible loop clause, allowing to represent any loop in a structural way. I'd recommend to look at cl-iterate package for CL, to see what happens to maniacs who had chosen this way.
2. A less ambitious approach defining some simple loops for most frequent cases (while, range iteration, iteration by iterator) and a loop for a general case looking like "loop { iteration }".
I personally prefer the second approach. I loved C-way, then I was a big fun of a Lisp-way, but now I believe that it is silly to create a whole new language just to write iteration, and a half-baked attempt to cover with for-loop more cases then just range iteration has more downsides than upsides.
Equivalent syntax was used in Fortran for writing cycles in a simpler way than in C already 20 years before C. ("do 10 i=0,N do 20 j=0,M")
The only case when the C "for" is useful is when the third operation is neither an addition nor a subtraction. The most frequent such case is when the operation is a link dereferencing, for accessing a linked list.
Such cases, for visiting all members of some non-array aggregate data, e.g. linked lists or trees, are solved more clearly with iterators.
https://www.kylheku.com/cgit/cppawk/tree/cppawk-iter.1
The above manual page includes an example of how to define an alpha_range clause that iterates over string ranges like alpha_range(var, "000", "999") or alpha_range(var, "AAA", "ZZZ").
There is a conditional clause if(...) which takes a condition and another clause as arguments. The iteration of the other clause is suspended while the condition is false.
At the shell prompt: add the values of the odd integers in the range 1 to 50.
$ cppawk '
> #include <iter.h>
>
> BEGIN {
> loop (range (i, 1, 50),
> if (i % 2 == 1, summing (sum, i)))
> ;
> print sum
> }
> '
625
Though if is an Awk keyword, this isn't a problem because if is a clause, not an ordinary expression. Moreover, though the clause is defined by macros, none of them are called if; they have if embedded in their name.Putting complex logic in the three places of a for loop is something we all frown on now because of readability and bugs, but back then it was almost more common than just iterating from 1 to N. Same for dropping through case statements - frowned on now (rightly) because gotcha bugs, but used heavily.
>Making "*" a prefix operator in B and C was a mistake
C dominates for the same reason it is terrible: it gets shit done for some value of "shit". In this light, there are no mistakes. There are other languages that don't make these "mistakes" and, behold, nobody wrote the majority of the world's operating systems and software in them. I've used C when the viable options were hand-crafted assembly or C. There are no mistakes here, only practicalities.
The result makes it much easier to refactor code. Ever try replacing a value with a pointer in C? Arrghh.
* Performance. Pointer access is an order of magnitude more expensive.
* Errors. Pointer access can fail (in particular, the NULL pointer).
It checks array access bounds, but segfaults for null pointers. Seems an odd choice, and there are multiple discussions of people saying the same.
Yes, I know, adding a huge fixed offset to the pointer can push it past the null protected pages.
It's super common to define a struct on the stack and then pass a pointer to an init function to initialize the value that is on the stack.
Sometimes pointers are faster because you don't have to copy the entire struct.
What largely determines the performance of values vs. pointers is whether that location in RAM is cached, and whether you need to allocate memory on the heap.
e.g. for(i=0,j=N.length;i<j;i++)
for i from 0 to N.length [by 1]At least that's what German Wikipedia Claims while the corresponding paragraph in english Wikipedia is short.
[1] https://vitrinelinguistique.oqlf.gouv.qc.ca/fiche-gdt/fiche/... [2] https://www.noslangues-ourlanguages.gc.ca/fr/cles-de-la-reda...
EDIT: "Creation of Computer Input in an Expanded Character Set" at https://ejournals.bc.edu/index.php/ital/article/download/292... from 1968, p112, describes it simply as "at".
EDIT #2: "TRANSLATION FROM MONOTYPE TAPE TO GRADE 2 BRAILLE" at p83 of https://archive.org/details/researchbulletin05lesl/page/82/m... from 1964 describes that symbol as the '"at" sign'.
I think that is early enough to say that email did not influence the "at" meaning but rather the other way around.
What I can't find is when it was first used in the US for indicating home-team for sports. E.g. [1] where there is vs. or @ depending on whether it is a home or away game. I suspect it's relatively modern, but not sure how far back it goes.
Also "pointer to integer" is written as "^integer" which is better than "int".
Syntax also doesn't allow confusing stuff like like "int a, b;".
I agree that it reads a little better. But as a small-handed person, it is unfortunately much more uncomfortable to type.
Another curious thing: Unix came shortly after and was entirely typed with lower case. Because those same terminals had only one font (font ROMs were tiny back then, and expensive). It looked like UPPERCASE but in fact the keyboard produced lowercase. Once better terminals were used to edit the code it was noticed for the first time(?) that it's all be edited as lowercase.
Or so the story went, back in those days, when it was all new.
The DEC PDP-11 assembler used the @ symbol for dereferences of registers ("@Rn" or "(Rn)" dereferenced a register, for example).
However, in Unix, the terminal convention was to use the @ symbol to delete the current line, and they didn't have a DEC assembler yet, so they used the asterisk (*) instead in B, as well as in the Unix PDP-11 assembler, which was written in B.
I'm paraphrasing this Quora post (which lists primary and secondary sources): https://www.quora.com/Why-did-the-developers-of-C-decide-to-...
Prior to IBM, it was on at least the Underwood typewriters[2], so it was never quite absent.
1: https://en.wikipedia.org/wiki/IBM_Selectric#/media/File:IBM_... Note that the number-row is only different from modern keyboards with the cent-sign being replaced with a caret (the cent-sign was never included in ASCII)
2: https://upload.wikimedia.org/wikipedia/commons/a/a4/Ernest_H...
Where’s that value in memory? It’s @ the address.
For reference, it's above the ' and between the ;: and #~ keys on my current board, rather than on shift-2 (where " lives).
Maybe this would have stayed consistent worldwide, if it was used in c.
That's just how they talked in 1972 in New Jersey where Bell Labs is located.
*
star
start
start of
;)It's where the C comes from!
Funnily enough, the guy that made B is also behind Go, where * and & are still a thing.
Ampersand begins with æ.
Address can use æ or the reduced vowel ə.
AH-dress vs eh-DRESS
Also, the "near on the keyboard" idea doesn't make sense anyway because C was developed on a bit-paired tty33 where the two symbols are not that close to each other.
And C++/CLI uses ^ for "managed pointers" (pointer to .net objects) and % for "managed references" which means there are all-together 4 ways to declare various types of pointers which is super-fun.
&
and
anddressI will remember it as the "Ursula" operator from here on
No it doesn't.
So, do people use consistent naming standards? Comment every line? Look at the assembly language produced by the compiler?
There are usually patterns that you match the code you're reading to. For example, the simple pattern of int f(int x, int y, int len, int *z);
and you immediately think "oh ok, they're doing something with x and y, and z is the out parameter. The fn is also likely going to allocate". Usually it matches what you see in the function body and it's all good. In (hopefully) rare cases you view it with such a lens and something doesn't make sense and you have to scrutinise the code for what it's actually doing. This is all without variable names. Proper variable names makes it much more straightforward to read e.g calling the function vec_add