On leading underscores and names reserved by the C and C++ languages
devblogs.microsoft.com
devblogs.microsoft.com
OK, I know it's only an offhand remark in a blog post, but now I'm going to have to spend significant energy on:
1. Finding where these identifiers are reserved, exactly, because obvious sources like https://pubs.opengroup.org/onlinepubs/9699919799/functions/V... don't seem to include them?
2. Trying to come up with the intended usage for these identifiers (`island` I imagine, might be a scope: only accessible to other identifiers on the same island? oh my...)
It's not that the word "island" has any particular significance; it's that all names starting with "is" followed by a lowercase letter are reserved (in external linkage and in the global scope of files that #include <ctype.h>), so that future versions of the language can add more standard library functions along the lines of isalpha, isdigit, etc., without making the current standards committee guess in advance which specific names following that pattern their successors might want to add in the future, and without breaking existing code. (Unless that existing code ignores these rules, which seems to be fairly common in practice.)
isisland(x);
Or to know if an element is a member of a particular terror group (or an Egyptian goddess): isisis(y); isisisland(z);
In this case you'd typically then construct a memory fence.Naturally, this would imply the validity of the optimization:
x = toilet(y); -> if (!isilet(y)) x = toilet(y);
[1] https://www.open-std.org/jtc1/sc22/WG14/www/docs/n2625.pdf (I hope that’s the right version)
Yes, that means that an implementation will warn you that your code might break in the future, but on the day when it actually breaks, it will just stop complaining. Why they could not just mandate detecting invalid uses of reserved identifiers is beyond me.
> Each header declares or defines all identifiers listed in its associated subclause, and optionally declares or defines identifiers listed in its associated future library directions subclause and identifiers which are always reserved either for any use or for use as file scope identifiers.
> Function names that begin with str and a lowercase letter may be added to the declarations in the <stdlib.h> header.
Or 7.31.13:
> Function names that begin with str, mem, or wcs and a lowercase letter may be added to the declarations in the <string.h> header.
So according to the standard, if your code #includes string.h, it may conflict with a future version of the standard if you use names like "strong" or "memorize". This is extraordinarily unlikely, though.
Because it's not the C compiler's business to care about POSIX.
POSIX and the C standard are different things. If you're writing code against the Windows APIs, POSIX isn't relevant for instance, and you need to be more concerned about colliding with type names or defines from the Windows API headers, and those are all over the place anyway.
...and besides: that rule is entirely pointless in reality, either your type names collide with POSIX types from headers you are including (in that case you're getting a compiler error anyway), or they don't collide, and in that case all is good.
That is, most of the point of that was for some standards to give a "follow these rules, and you will have an easier time integrating with this standard if that is in your plans."
Obviously, if you have no plans to head to posix, they accomplish nothing for you. Similarly, not following the rules doesn't prevent that direction, just adds some extra work. Potentially.
I think it means: a library author cannot name some type epoch_t, because POSIX may introduce an epoch_t in the next revision, and suddenly some code may fail to compile.
It is when people write POSIX-compliant code and their compiler doesn’t let them.
It is the POSIX platform owners business to ensure POSIX compliance on their own C compiler.
It’s just a way to state “if you’re using this API/library, be prepared that future versions may introduce additional symbols matching these patterns, and/or that the current version may define undocumented symbols matching these patterns”. You then have the choice to take care to not define symbols conflicting with that.
The alternative would be for libraries to just add arbitrary new symbols, with no way for client code to proactively prevent conflicts.
Love that tone
They also allow for registration free COM components.
To this day many Microsoft teams still haven't gotten the memo.
Or maybe they did - back when MSDN was actually well-organized and somewhat complete-ish (and shipped on optical disks or otherwise downloadable!), I was too young to make much use of it. Now, that wealth of knowledge is mostly gone, and external links to it don't resolve.
These are the current locations for them,
https://learn.microsoft.com/en-us/windows/win32/sbscs/creati...
https://learn.microsoft.com/en-us/windows/win32/sbscs/manife...
https://learn.microsoft.com/en-us/windows/win32/sbscs/isolat...
https://learn.microsoft.com/en-us/windows/win32/sbscs/author... (see last bullet points regarding not to use the registry as best practice)
The problem as usual, is lack of education on the matter, and willingness to change how people work.
I stand corrected on one point though, it has been a long time and actually it was only introduced on Windows XP and Server 2003, not Windows 2000.
At least that's my guess, given recent trends.
I have very slowly come to appreciate it as a language, but visually it will always be ugly to me because of clashes like this with other languages.
> const __m256i in = _mm256_loadu_si256((const __m256i*)ptr);
It's like complaining you have to use _Bool all over the code. Include stdbool...
And memo appears to be one of the disallowed variable names in C11, given the mem[a-z].* pattern
C++
foo = bar \* 2;
Are foo and bar local variables or members of some instance?vs
Python
self.foo = self.bar \* 2
100% clear. No naming convention needed.I bring this up because `_foo` for members is a naming convention that wouldn't be needed in a language that required `self` or `this`
That said, I get that maybe refactoring some code from standalone function to class method is easier if you don't have to change the code as much but I'd be curious how often that's a net win.
void add(const this, int x)
Is more readable than (IMO) void add(int x) const(Scroll down to the "proposed syntax" section.)
I agree that I like when functions are not special so
instance.add(10)
Is just sugar for add(instance, 10)
and you pass anything that fits as the first argument.But, following the "syntactic sugar" is okay rule
class Foo {
add(int v);
}
Is just syntactic sugar for void Foo.add(Foo this, int v);
... or something along those lines... ?The pimpl pattern is sort of a weird artifact of how the compiler works, but it generally works out as a smart way to structure your code.
So if you disregard the fact that those are reserved, it might build, but even a minor update of your implementation that e.g. adds a new local to some obscure standard function can break it, never mind a new version of the C++ standard.
https://gist.github.com/chjj/d0c1218e473bbb6d8f9e2224c583e2d...
Yes, I would guess that a large proportion of C programmers are not familiar with those rules. Is there a way of getting GCC or LLVM to warn about it?
The issue here is not GCC/Clang. It's developers
You risk having a name collision if you define a strfoom in your code and then later the C committee decides to add that to string.h
It's a stupid decision to reserve those and it ruins the usability of the language by arbitrarily disallowing many common words that have nothing to do with the intended feature. Even if you don't export names like this directly, having the name of your library or company (as prefix for all exported names) start with those common combinations is likely as well, so now you got to avoid certain library or company names.
If this is only since C11, then I'm even more baffled, since I could somewhat see how such thing could happen in the 1970s when there were limits and the language was brand new and not known to be popular, but in 2011 doing this makes no sense whatsoever.
I hope this is limited to C11 and will never happen in C++.
-use more obscure letter combinations than "is", "to" and "str" which are very common beginnings of words. How do "is", "to" and "str" help anyway if they want to add functionality that has no sensible name starting with those?
-use two underscores at the beginning since as per the article that is already reserved
-use stdc_ as prefix
stc_ or similar prefixes is the way to go.
I'd argue that a compiler update that breaks existing code for something as innocuous as using the "wrong" variable name is a compiler bug, even if the code technically violates the spec. It is simply too late now for the spec to begin using the keywords they reserved 40 years ago.
There was a similar thing for _Complex.
I'd like to see isualpha() isudigit() isupunct() isuspace() isugraph() etc. plus isucombining(). They'd supplement iswxxx() but (a) taking a `signed unicode int` (for C2x defined as at least 32 bits wide) and (b) being locale-independent.
> This check does not (yet) check for other reserved names, e.g. macro names identical to language keywords, and names specifically reserved by language standards, e.g. C++ ‘zombie names’ and C future library directions.
#define _POSIX_SOURCE
which is sometimes needed before the #include's to get access to nice, modern library functions that are, say, thread-safe.(Dave, however, has a really annoying presentation style, and is so overrated.)
I was surprised by C++11 having additional prefixes it has reserved. I'm now also wondering whether there is a clang/gcc option to warn about such things, as although I know we don't currently have any issues in our code base (as in, it compiles and works) I don't really want to publish a public API and have to revisit it because of such a conflict in C++29 or whatever
I'm not going to make a big deal about it, but it makes more sense to me to put the scope of a variable at the front. The front is where you put foo. and foo-> and foo[], after all.
Since we read code far more than we write it, sprinkling little "usage hints" like this across symbol names removes a lot more cognitive overhead than I would have thought.
I wonder if that paid off in some cases? Or to ask differently, if one would design a new programming language (now), would you consider reserving words in advance?
I'm currently making a language, and yes, I'm reserving words for it because I want an easy C ABI. But I'm also making them easier to avoid.
Here's my list of reserved words:
* Anything that begins with `y_`.
* Anything that begins with `yc_`.
* Anything that begins with `YC_`.
* Anything with three or more consecutive underscores.
* Edit: Anything that begins with an underscore. This is because my language will be able to transpile to C if necessary.
The first is for types and items in the standard library (the language's name is Yao, so a `y` makes sense). The second and third are for the C ABI (hence, `yc`) and for historical reasons. `YC_` in particular is for C macros.
The last is for name "mangling." I put it in quotes because my language's standard name mangling (it will be the same across every implementation) will not really mangle the names. Instead, it will concatenate them, using five underscores between package names, four underscores between packages and items in a package, and three underscores between an item and its suffix. (For overloaded functions, the programmer has to define a suffix for each one. That suffix is how their names will be different in the C ABI.)
My hope is that these rules will not be too onerous. I don't think the reserved prefixes are used much, and I haven't seen anyone use more than one consecutive underscore, though I'm allowing one more, just in case.
Declaring variables with those names actually will not conflict if they are done in pure Yao. There's a special way to access the C symbols of functions and types, and it's deliberately different, for this very purpose.
The restriction isn't on `y[c]_`. It's actually on `y[c]___`.
This is because `y` is the package name for the standard library, and the standard library name will always be separated from the rest of the name by 3 or more underscores.
Really, the restriction is on 3 or more consecutive underscores in C names, and you can't have have a package name that is either `y`, `yc`, or `YC`.
I apologize for the confusion. I was in a hurry and on mobile with the original post.
As for payoff, new versions of the C standard usually introduce new functions and macros with those name patterns, without breaking client code that respects the reserved-name rules.
I think the idea of reserving keywords in advance is a good one, but not something that interferes with vocabulary words so easily. The idea is to be able to add a new keyword in the future without breaking code that uses that keyword itself currently.
class Example {
#privateFieldName
}
Rather than: class Example {
private privateFieldName
}
The debate on that was pretty interesting.Also, a leading underscore followed by a capital letter, or other reserved name, is absolutely allowed in implementation headers -- even in non-standard headers. It is bad practice only because somebody else's compiler (e.g. Clang) might be obliged to read those headers someday, and have invented its own meaning for it.
And, only C++ is of any interest, here. Microsoft never gave a damn about C, and any name reserved in C is also reserved in C++.
I suspect it’s not that people can’t do the work and more that Apple’s silos mean their already employed engineers are not allowed to try or even talk about it.
I wonder if it was possibly not that you compared the kernels but rather how you chose to express the comparison. For example this bit makes me think you possibly expressed your opinion combatively:
> a dogshit IO scheduler
Perhaps you were more polite in your day-to-day, I don't know. But your language here makes me wonder.
So I'm sorry I offended you with my salty language but the evidence suggests I'm far from the only one that went ignored at Apple.
the way you are responding to some one who asked you a pretty mild question (assuming they were personally offended rather than asking about a potential communication issue) suggests there is some validity here.
Sure the phrasing was gentle, but the actual statement was quite offensive, presumptuous, and prudish.
You didn't.
https://www.collinsdictionary.com/us/dictionary/english/some...
For things like Final Cut/iMovie with lots of video/audio/misc tracks, it was trivial to saturate the disk with dumb seeks when reading otherwise linear data streams due to a lack of knobs.
Which is probably the root cause of the GUI civil war happening between all desktop frameworks.
Inside Windows NT from Microsoft Press
First, Candy Crush was never preinstalled. A link to purchase it in the Microsoft store was preinstalled, and it took all of two clicks to get rid of it if you wanted to.
Second, preinstalling that link reduced malware infections of windows systems by a visible percentage worldwide.
Microsoft has a duty to protect users who need it. Finding a compromise where the worst case for other users is that they need to click twice is pretty good.
Also as an aside, that link _kept coming back_ on my laptop (but not my desktop oddly).
Why? Because people would pirate it? Or Just download virus laden games in general?
https://devblogs.microsoft.com/oldnewthing/20041217-00/?p=36...
Unfortunately, Microsoft didn't exactly treat him kindly [1] in the end. And to add insult to injury, his blog was completely wiped. There are some archives around, but I haven't seen one that retained the images, which often makes the posts incomprehensible, unfortunately.
[1] https://vsubhash.wordpress.com/2017/04/17/rip-michael-j-kapl...