The Clockwise/Spiral Rule of C declarations
c-faq.com
c-faq.com
The correct rule is "follow the C grammar". An easier to remember and also correct rule is "start at the identifier being declared; work outwards from that point, reading right until you hit a closing parenthesis, then left until you hit the corresponding open parenthesis, then resume reading right..." (this is sometimes called the "right-left rule"[2]).
The "spiral rule" dances around the truth without actually being precise enough to be useful.
[1] https://news.ycombinator.com/item?id=5079787 [2] http://ieng9.ucsd.edu/~cs30x/rt_lt.rule.html
Thanks!
const char *foo[][50]
the following expressions have the following types: foo -> const char *[][50]
foo[0] -> const char * [50]
foo[0][0] -> const char *
*foo[0][0] -> const char
Another example: int (*const bar)[restrict]
bar -> int (*const)[restrict]
*bar -> int [restrict]
(*bar)[0] -> int
One more: int (*(*f)(int))(void)
f -> int (*(*)(int))(void)
(*f) -> int (* (int))(void)
(*f)(0) -> int (* )(void)
*(*f)(0) -> int (void)
(*(*f)(0))() -> int
(In this case, the last expression is more readily written as f(0)(), since pointers to functions and functions are called using the same syntax.) const *
with * const int *const p;
Pointer to const value: const int *p
int const *p
>const is read as "the thing to the right of me is const"const is one of storage classes and is read at its order, not just "to the right".
bar -> const pointer to mutable array of ints
*bar -> mutable array of ints
And const pointers are dereferenced with * , not (* const), so the rule needs an exception for const pointers (as well as volatile pointers).To follow the declaration you make use of the fact that postfix operators in have a higher precedence than unary, and that of course unary operators are right-associative, whereas postfix are left-associative (necessarily so, since both have to "bind" with their operand).
So given
int ***p[3][4][5];
we follow the higher precedence, in right to left associativity: [3], [4], [5]. Then we run out of that, and follow the lower-precedence * * * in right-to-left order.If there are parentheses present, they split this process. We go through the postfixes, and then the unaries within the parens. Then we do the same outside those parens (perhaps inside the next level of parens):
int ****(***p[3][4][5])[6][7];
1 2 3 4
765
8 9
1111
3210
Start at p, follow postfixes, then unaries within parens. Then the postfixes outside the parens and remaining unaries.The result is in fact a spiral just from going root postfix unary out postfix unary out. We just don't have to focus on the spiral aspect of it.
int* arr[][10];
Spiral rule would state "arr is an array of pointers to arrays of 10 ints", where actually it would be "arr is an array of array of 10 pointers to int".Instead, when you write declarations, do it from right-to-left, e.g.:
char const* argv[];
"argv is an array of pointers to constant characters"It doesn't help with reading, unfortunately.
int* arr[][10];
If you index twice into arr and then dereference, you'll get an int. So arr must be an array of array of pointer to int.The other part is `int* x, y`.
char const
rather than const char
Is there a reason to prefer the second version? It's a lot more popular in my experience.Also your argument about which modifies which is strongly anglocentric: there are plenty of people whose native language puts modifiers after the things they modify.
str [10]*byte
which reads exactly as it is declared: "str is an array of length 10 of pointers to byte" (byte is Go equivalent of C char (mostly)).On the other hand, IMHO the whole "make declarations read left-to-right" idea is misguided --- plenty of other constructs exist in programming languages which simply can't be read left-to-right, but are nested according to precedence. I mean, you might as well make 3+4*3 evaluate to 21 if you want to try making everything consistently left-to-right, but I don't really see anyone complaining about not being able to understand operator precedence...
The point here is that type declarations are regular to read, and those tend to be the tricky ones. Expressions tend not to be so difficult, and are more commonly factored if they become complex. For various reason, type declarations are not so practically factorable.
Declaring `v * T` means we can write `* v` as an expression, so the use of token * is synchronized for both these uses, but I must vocalize the * in my head differently:
`*T` vocalizes as "pointer to something of type T"
`*v` vocalizes as "that pointed to by variable v"
`&v` vocalizes as "pointer to variable v"
So my thought process when I see * goes: If it's in a type, say "pointer to", otherwise say the opposite of "pointer to", i.e. "that pointed to by". It feels like an inconsistent use of * whenever I'm writing Go code -- even though I know it's a natural result of Go using Pascal-style declaration syntax but C-style tokens. var str: array[10, ptr byte]
(and much richer types)Edit: and while I'm here, Nim has other sensible syntax for this low level stuff...
var b: byte = 10
str[0] = addr b
echo $str[0][] std::array<std::pointer_to<byte>, 10> str;
certainly does not look any more readable to me than byte *str[10];
. (Disclaimer: I mainly work with C, but find some C++ features genuinely useful, although the majority of the time they seem more like absurd complexity for the sake of complexity.)std::array is useful for letting the compiler avoid array-to-pointer decaying, value semantics, and also actually putting array length type info in a function parameter.
let string: [&u8; 10];
string is an array of references to unsigned integers of 8 bits of length tenFor example, D also uses a similar type syntax, so in D if you declare:
int[10][20] x;
x[19][9] // is legal
In C: int x[10][20];
x[9][19] // is legal
I think the correct solution would have been to make pointer syntax post-fix like the arrays and functions, so that you get the best of both worlds. Go-like declarations and C-like matchup between use and declarations.(PS. Golang has the right idea, since its developed by the guys who contributed to C)...
Have you ever had to write a C parser?
If you don't have the available type names then it becomes ambiguous.
The alternative would be for the type syntax to mirror the expression syntax used to construct values of the type. Functional languages tend to do this, particularly ones which prefer pattern matching over destructors.
Cdecl (and c++decl) is a program for encoding and decoding C (or C++) type declarations.
foo(*baz(bing,boff(*bratz)(biff)))(buff);Even with typedefs, that declaration means “when you call baz with a bing and a pointer (named bratz) to a function of type boff(biff), then you get back a pointer to a function of type foo(buff).”
It’s an extremely concise notation for expressing type information without (much) special type syntax, and I think it’s quite elegant in that way.
foo (*baz(bing, boff (*bratz)(biff)))(buff);
“foo” is the type, and the rest is the declarator. Then you just break it down according to the usual precedence rules: baz(…)
“baz” is a function… baz(bing, …)
…which takes a “bing”, and… *bratz
…a pointer (arbitrarily named “bratz”)… (*bratz)(biff)
…to a function which takes a “biff”… boff(*bratz)(biff)
…and returns a “boff”… *baz(…)
…and “baz” returns a pointer… (*baz(…))(buff)
…to a function taking a “buff”… foo (*baz(…))(buff)
…and returning a “foo”.With typedefs for function pointer types:
typedef boff (*bratz_t)(biff);
typedef foo (*baz_ret_t)(buff);
baz_ret_t baz(bing, bratz_t);
Or for function types: typedef boff bratz_t(biff);
typedef foo baz_ret_t(buff);
baz_ret_t *baz(bing, bratz_t *);It's been 20years(!) Why is this incorrect advise still up at c-faq?
(())()(()())(())
these count different arrangement of parentheses for function application. this guy is describing something like contour integration for computer programsYeah, I know, I'm not good enough, I didn't study enough, I'm not enlightened enough. But why make things so overly comples in the first place?
"complex" is subjective. It reminds me of stupid "rules" like "don't use the ternary operator", "every function must be less than 20 lines" (I am not exaggerating --- this was on a Java project, however); and you could easily extend that to "every statement must have a maximum of one operator", "you must not use parentheses", "you must not use more than one level of indirection", etc. Where do you stop? To borrow a saying from UI, "if you write code that even an idiot can understand, only idiots will want to work on it." I don't think we should be forcing programmers to dumb-down code at all.
That said, I'm not advocating for overly complex solutions, and will definitely prefer a simpler solution, but you should know and use the language fully to your benefit.
To me, that makes about as much sense as when Ricky Bobby in Talladega Nights says "If you ain't first, you're last."
And people wonder why there are so many broken C programs out there...
If the complexity can be avoided, why not avoid it. Removing complexity is not the same as dumb-downing code. It will improve readability and maintainability.
This mindset is defintitely applicable to declaration as well as code construct.
edit: clarity.