C23 is finished: Here is what is on the menu
thephd.dev
thephd.dev
This creates a problem. How do you portably printf() an integer type that you don’t (and can’t) know the size of, like, say, a uid_t?
Before, when intmax_t was guaranteed to be the largest type, you could reliably and portably cast a uid_t userid to an intmax_t and printf() that: printf(PRIdMAX "\n", (intmax_t)userid);
But now, when any type might be larger than an intmax_t? What do you do?
(The same problem exists when reading an integer string; how do you read, say, a port number, when you don’t know what type in_port_t is? Previously, you could call strtoimax() on the string, detect overflow, and then cast the resulting intmax_t to in_port_t and finally compare for equality to make sure it wasn’t truncated.)
> There is active work in this area to allow us to transition to a better ABI and let these two types live up to their promises
Yes please. Introducing new wider integer types in your CPU, but pretending you haven’t created a new architecture with an associated new ABI, is silly.
Complain to POSIX so they add printf macro defines for all their implementation-sized integer types by the year 2100.
Similarly for the second problem you can read digit by digit in reverse order and create the number using x = 10x + d while detecting overflow.
I'm not saying that this is a great thing to need to do, but the mentioned problems aren't particularly fundamental and the solutions are CS101 stuff.
Sure, someone might come up with an architecture where that's not true. But even in an intmax_t-is-really-the-largest-integer-type world [which hasn't been true for the major C compilers for the past decade or so!], I can almost assuredly come up with some compilation modes that still conform with whichever C version you use that still causes your code to fail horribly. I don't think there's any architecture where uid_t et al would exceed uint64_t, and until such an architecture comes to exist, it's quite frankly not worth caring about.
But the "physical" size of the number as represented in memory can change across platforms, no?
- POSIX defines that in_port_t is equal to uint16_t, so for that type there's no problem at all ;)
- Don't use printf/scanf for this, roll your own
- Refuse to print/read uid_t that are outside intmax_t range
- Use an autoconf test
Until then I guess I’ll have to do something like
#if sizeof(pid_t) > sizeof(uintmax_t)
#error Nope
#endif
…repeated for all the types I need to parse and/or print.Also this doesn't just apply "in the future" -- integer types larger than intmax_t already exist, C23 just updates the standard to match reality.
#if SIZEOF_PID_T > SIZEOF_UINTMAX_T
...
#endif
Where you detect these sizes from the toolchain and deposit them as #define constants in some "config.h" header.There are ways to detect sizes by compiling a source file to an object file, and then analyzing the object file (no execution), so things work under cross-compiling.
I've used a number of tricks over the years, and settled on this one:
https://www.kylheku.com/cgit/txr/tree/configure?h=txr-278#n1...
The basic idea is that we can take a value like sizeof(long), which is a constant, and using nothing but constant arithmetic, we can convert it to the characters " 8". These characters can be placed into a character array where they are delimited by some prefix that we can look for in the compiled object file with some common utilities.
This is quite durable in the sense that it can be reasonably expected to work through a wide variety of object formats.
In that test program, I have a structure with some character arrays. The larger character arrays hold the informative prefixes. The two-byte arrays hold decimal digits for sizes. Two digits go up to 99 bytes so we are safe for a number of years.
The DEC macro calculates the ASCII decimal value of its argument, expanding to a two-byte initializer for a character array:
#define D(N, Z) ((N) ? (N) + '0' : Z)
#define UD(S) D((S) / 10, ' ')
#define LD(S) D((S) % 10, '0')
#define DEC(S) { UD(S), LD(S) }
E.g. DEC(42) -> /* the equivalent of */ { '4', '2' }
But actually DEC(42) -> { UD(42), LD(42) }
-> { D((42) / 10, ' '), D((42) % 10, '0') }
-> { ((42) / 10) ? ((42 / 10)) + '0' : ' ',
((42) % 10) ? ((42 % 10)) + '0' : '0' }
The space instead of a leading zero is so that if we pull this into shell arithmetic, it isn't confused for an octal number.Well, that code is not, and was not ever, standards compliant.
After all, if predictable behavior in the future across all possible configurations were a high priority requirement for the tools people use to make computers do things, neither C nor C++ would have gotten off the ground. Instead, people don't memorize the entire standard. They write code that works on their machine, publish it, fix the bugs when people say it doesn't work on their machine, and maybe converge towards something that is standards compliant but, more likely, they converge towards something that works on the standard implementation on 90% of the operating systems available.
Would you be able to name three examples?
My intention was to make a good case for old code being compiled today by third parties.
But to my understanding intmax_t was introduced in C99, before widespread dominance of 64 bits platforms. So it would not be surprising that there exists code with lower assumptions about its size.
But actually I made also another mistake, the major problem had little to do with faulty logic and more to do with dynamic linking [0]
I will see myself out today.
If you want to know "what is the widest signed integer type available", you will get some kind of answer one way or another.
Obviously, when someone says “type X is size Y on architecture Z”, they’re not talking about the C standard, they’re talking about some particular ABI.
There are multiple ABIs for x86_64. I think the major compilers use 32 bits for int and 64 for long, but there could very well be a compiler that used different sizes, with a different ABI.
> A compiler vendor can set the sizes to whatever they want regardless of the target architecture and claim conformance if they meet the minimum.
Compilers claim conformance with other standards besides just the C standard.
Anyone who uses intmax_t in a durable API (public function argument, member of public structure) is just a goof.
How would you compile that to machine code which would, when run, still give you the maximum integer on next year’s CPU with all-new super-duper wide integers?
It is more reasonable to think of a binary to be compiled to a specific architecture triplet, and if the CPU changes, the architecture must change, and therefore the ABI, and if the OS wants to run a binary from an older architecture, some emulation layer is needed.
Of course, this would be a lot of work for the operating system people if they want to run binaries from lots of what are now different architectures, and it was apparently easier to just force the C standard to abandon the entire concept of intmax_t being the largest.
intmax_t by intent is the maximum integer available on a given platform. Abandoning that idea is a disservice to compiler writers, programmers -- especially scientific programmers -- everywhere.
That would add a lot of friction to changing CPUs, and making distributing software more complicated.
"Portably" across what? You need to narrow down the problem space and better define it.
If it's portably across a myriad of processor architectures, hardware profiles and operating systems then you would need to take the intermediate representation route (like Java). There's very little else that can be done in the face of infinite, non-compatible ABIs at every layer.
> you would need to take the intermediate representation route (like Java). There's very little else that can be done in the face of infinite, non-compatible ABIs at every layer.
No, I don’t think so. There’s an obvious solution, dictated by the architectual choices made by C (i.e. every architecture which introduces a new larger integer type, by necessity redefines intmax_t, and creates a new ABI), but the community (for reasons of their own) did not want to do it that way, and instead broke the promise of intmax_t always being the largest type.
https://thephd.dev/c-the-improvements-june-september-virtual...
Someone's got to write a paper to get it in, hopefully based on some sort of existing practice.
That makes parsing the unsigned variant relatively simple. Undefined behavior would be end-of-game, but unsigned addition and multiplication will not invoke that (if you need he,p, look at your favorite multiple-precision library for ideas about detecting wraparound)
For parsing signed integers, derive the min and max values of the signed variant from the max value of the unsigned one ((whatever_max - 1) / 2 or something like that
Then parse your input to the unsigned variant first, then check whether it fits in the unsigned version.
Here is the direct link: “Use #embed from C23 today with the Cedro pre-processor” https://sentido-labs.com/en/library/cedro/202106171400/use-e...
As I write there, “The advantage of using Cedro is that the source code is the same and this way it is very easy to use different compilers, only adding or removing the #pragma line.”
I also added a command-line option to embed the bytes as strings instead of byte literals; embedding an 8 MB file, the strings variant compiled between 28 and 72 times faster depending on which compiler I used, gcc or clang. It is more efficient, but less compatible, that’s why it’s optional.
I’m considering adding some low-hanging fruit, like the number literal separators: 1'384'849 → 1384849
Maybe that’s a better idea, implement it with underscores in Cedro as it does not need C++ compatibility.
Edit: alright, it works. I also extended the parser to accept the apostrophe as part of number tokens. The only thing it does is to remove the underscores when writing.
I might need another command line option to specify whether the output should be C23:
1_234_567 → 1'234'567
or pre-C23: 1_234_567 → 1234567
1'234'567 → 1234567They should have formally defined NULL as void* instead.
> Someone recently challenged me, however: they said this change is not necessary and bollocks, and we should simply force everyone to define NULL to be void*.
I'm glad someone brought it up. The scenarios nullptr "fixes" are so minor adding a new null concept to fix them is the sort of overkill solution I expect from the C++ committee, not the C committee. It's a shame this wasn't thought-out more.
> I said that if they’d like that, then they should go to those vendors themselves and ask them to change and see how it goes.
The C committee has already broken compatibility on numerous occasions: forcing 2's complement representation, changing the type of u8"" string literals from char to unsigned char, removing mixed wide string literal concatenation (allowed in C11, disallowed in C23). It's difficult to be convinced by the argument of backwards compatibility when the committee has already broken it on multiple occasions. Furthermore, I seriously doubt compiler vendors who define NULL as 0 would take issue with redefining it as void*. This change is so minor and inconsequential there should be no push back.
>Introduce the nullptr constant
>They should have formally defined NULL as void* instead.
That doesn't make any sense to me. void* is a type isn't it? nullptr is a value. I don't see the overlap.
The C standard permits NULL to be defined as (void*)0 or 0. The issue is this creates an inconstancy when NULL is used with _Generic selection: a compiler defining NULL as (void*)0 will cause NULL to be caught by a pointer case whereas a compiler defining NULL as 0 will cause it to be caught by an integer case. This inconsistently could be solved by mandating NULL be defined as (void*)0 so it will always be caught by the pointer case.
The issue is instead of mandating (void*)0 the C committee decided to introduce a whole new nullptr concept. This duplicates the concept of null in the language and requires all C compiler vendors to modify the languages type system to account for it. So it creates more work for vendors and duplicates the concept of null all to solve this minor use case.
Exciting to see this in the standard. Once this is in all compilers, this will probably have security effects beyond C.
Fun fact: Miguel Ojeda linked to @cperciva's article and its following HN discussion in the note (N2897) to propose this: https://news.ycombinator.com/item?id=8270136
Had a disscussion about this here:
I tried the following code in a few compilers, it works perfectly for overwriting buffer pointers that goes out of scope, also tested removing the volatile to ensure that the memory write was ignored and I reproduced the code vulnerability where the previous message was still there the second time I tried to read a message, and from the way I read the spec this shouldn't be optimized away by any compiler:
void memset_explicit(char* p, char c, size_t n) {
volatile char* vp = p;
while (n--) *vp++ = c;
}
Casting other data types to a pointer to volatile data is within the spec, and the compiler should then treat it like any other volatile data.Up until now, safely implementing explicit_bzero() has depended on various hacks that a new compiler version might drag out from under you. This should end the "arms race".
(if it isn't, hopefully somebody who knows more will post a better link)
memset_explicit is just to overcome stubborn optimizer people who insist to optimize functions away (without warnings!), even when they have no idea about side-effects.
Wow, was that a requirement for memset_s? I'd never heard of timing being a something guaranteed by the *_s functions, regardless of security being in their name.
Many of those open source projects happen to only compile with GCC, exactly because they rely on GCC C.
Google has spent several man years making the Linux kernel compilable with clang, and not all of it has reached upstream to this day.
Here's nbdkit compiled with clang 14:
https://gitlab.com/nbdkit/nbdkit/-/jobs/2814022460
checking whether clang is Clang... yes
[...]
checking if __attribute__((cleanup(...))) works with this compiler... yesWhich parts haven't? Asking as the project lead.
At least that is my impression from occasional Linux talks.
By the way what's up with NDK, Android packages, and that roadmap with better C++ development?
Thankfully no longer a problem of mine.
$ git clone git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
$ cd linux
$ make LLVM=1 -j$(nproc) defconfig all -s
$ echo $?
0
$
See also the official kernel docs which describe more which architectures can be expected to compile cleanly out of the box. There are something like 10^6000 possible kernel configurations though, so no promises (regardless of toolchain). https://www.kernel.org/doc/html/latest/kbuild/llvm.html#supp...
We support clang-11 and newer.
Linus Torvalds has been using clang to build the kernels he runs personally, for ~2 years now.
https://lore.kernel.org/lkml/CAHk-=wiN1ujyVTgyt1GuZiyWAPfpLw...
>> I build the kernel I actually _use_ with clang, and make sure it's clean in sane configurations
Not sure what's going on these days in Android land, thankfully no longer a problem of mine.
I just read the actual proposal on it. They describe an example implementation using macros and _Generic.
What’s really bad about this is that the arguments provided to these methods are used more than once in the macro body. This will cause:
- gibberish compiler errors in case of typos in the arguments,
- worse compilation times, as the compiler has to process more code.
Constructs like these may lead to an exponential explosion in preprocessed source code size. I hope that nobody implements it that way.
#define IS_POINTER_CONST(P) _Generic(1 ? (P) : (void *)(P) \
, void const *: 1 \
, default : 0)
#define STATIC_IF(P, T, E) _Generic (&(char [!!(P) + 1]) {0} \
, char (*) [2] : T \
, char (*) [1] : E)
#define _STRING_SEARCH_QP(T, F, S, ...) \
STATIC_IF (IS_POINTER_CONST ((S)) \
, (T const *) (F) ((S), __VA_ARGS__) \
, (T *) (F) ((S), __VA_ARGS__))
#define memchr(S, C, N) _STRING_SEARCH_QP(void, memchr, (S), (C), (N))
#define strchr(S, C) _STRING_SEARCH_QP(char, strchr, (S), (C))
#define strpbrk(S1, S2) _STRING_SEARCH_QP(char, strpbrk, (S1), (S2))
#define strrchr(S, C) _STRING_SEARCH_QP(char, strrchr, (S), (C))
#define strstr(S1, S2) _STRING_SEARCH_QP(char, strstr, (S1), (S2))
Here `IS_POINTER_COST` expands `P` twice so `_STRING_SEARCH_QP` expands `T`, `F`, `S` and each variadic argument 2, 2, 4 and 2 times respectively. Note that only `S` and variadic arguments are expressions and evaluated only once, but the number of expanded tokens can blow up exponentially.The real issue, in my eyes is, that "const" simply doesn't transfer well across function call boundaries. "const" is always only a local judgement - data that is const for one function is not necessarily const for another function. I've even read an old conversation where Dennis Ritchie was uttering concerns about const when it was initially designed.
Typically I use the const keyword only for global static data, because there it makes a real technical difference. Sometimes I declare function pointer parameters as pointer-to-const, partly for reasons of documentation, but I'm always aware of the strstr() issue.
field *my_struct_get_field(my_struct *s)
const field *my_struct_get_field(const my_struct *s)
In practice you will choose the first instance, so you simply won't use const where you could. Which is unfortunate.It would be nice to have a shorthand for this. The nested macro shown in the paper looks like it can fix the problem, but is quite complicated. I would like something like:
autoconst field *my_struct_get_field(autoconst my_struct *s)
The returned pointer inherits the qualifiers of the argument. Inside the getter function the pointer is const, and there is only one implementation of the function.An exponential code size increase even with the example implementation would only happen if people are nesting these functions inside each other. And while that could happen, it would be unusual to nest them very deeply. So for compilers that just want an easy implementation, and don't really care about quality of error messages or maximizing compile speed, the shown approach is probably fine.
It was a bit silly that stuff like that wasn't in either standard until this recently.
Damn, so close. Well I hope it's included next time.
Has anyone seriously proposed adding namespace to C? This feels like an obviously good addition to me. What’s the argument against it?
std::opinion::bad_idea()
;)
namespace foo { void x(); }
and namespace bar { void x(); }
and then you have to rely on compiler vendors to use the same mangling everywhere otherwise you end up exactly at the C++ position where there are multiple incompatible name mangling schemes, thus C code compiled with e.g. cl.exe would not be able to call a C function compiled with gcc (and FFIs wouldn't be able to either so you loose the "easy language bindings" "feature" of the operating system ABIs)Name mangling is needed if you want to overload functions and qualify them only by the types (not names) of their in and out parameters. This is where it gets ugly on the binary level.
how do you do when you want to access your C function from a language which is binary-compatible with C but uses :: for something else? [a-zA-Z_][a-zA-Z0-9_]* identifiers are the only thing that the whole world more-or-less standardized around.
e.g. in fortran you can directly import a C function and call it. But "::" in the middle of a function name would very likely fail (I don't know enough fortran to tell for sure but given how the syntax looks...):
subroutine foo (a, n) bind(c)
import
real(kind=c_double), dimension(*) :: a
integer (kind=c_int), value :: n
end subroutine foo
allows to call a "void foo(double *a, int n);" function defined in C. I imagine that subroutine ns::foo (a, n) bind(c)
would likely not workTypically you'd do this by picking a different internal name for the function, and putting the external symbol in quotes.
Take String::contains(). The equivalent feature in C++ is an overloaded function, so there are I believe it's three, versions of this function which take different parameters: A string, a char and a pointer to chars, they do similar things in practice but the compiler has no idea, there are just three functions with the same name. However the Rust feature is polymorphic, there are N versions of this function depending on what monomorphisations are chosen at compile time. The compiler knows these are all the same function, but the parameters have different types in practice at runtime and so the generated machine code is different. If your program can String::contains(cat_photo_jpg) then the compiler will produce the code to call it with your Jpeg type or whatever and it will need to keep it distinct from the version where it takes a String or a char or whatever.
Rust does this more often than C++ because it cannot choose ad hoc polymorphism. So if there should be a foo function which can take parameters bar, baz or quux, we need to decide either that bar, baz and quux all implement some trait which foo takes, or we need three separate functions foo_with_bar, foo_with_baz and foo_with_quux.
What does the term ad-hoc polymorphism mean in your comment? Are you saying Rust does not have ad-hoc polymorphism because the syntax acts like the function in parametrically polymorphic and only once monomorphisation occurs are the different functions per argument type are generated is it N different functions. Does this line of thinking also say Haskell does not have ad-hoc polymorphism?
And in case any reader wonders, I am genuinely interested in your response to these questions (i.e. this isn’t bait nor rhetorical).
The C++ standard library defines contains(x) so that it'll work for a string x, a single char x or a pointer to chars x. Nothing else can work, those are the arbitrary list of types which work, that's ad hoc polymorphism.
A separate implementation is provided for each of those three cases, which is why if I've got a JPEG, I can't instead ask if the string contains the JPEG, that's not one of the three implementations provided.
The Rust standard library just defines the type of x in terms of a trait, Pattern. So, my Jpeg type can just implement Pattern and now I can ask if Strings contain the Jpeg and that works because contains delegates the matching problem to Pattern and Jpeg implements Pattern -- only the specific idea of "contains" as distinct from "begins with" or "split" or a dozen other functions is handled by the contains function.
I am not a (serious) Haskell programmer but I would argue that Haskell also lacks ad hoc polymorphism here as I understand it (and to be clear: I think this is in general a good or at worst reasonable choice)
The place where ad hoc really shines is when you're got a function that say, makes complete sense with exactly two (or maybe three, fewer than two is irrelevant and more than three seems unlikely to provide reasonable ergonomics) specific types, which otherwise have nothing useful in common.
For example suppose I've got a whole variety of bird types, Goose, Chicken, Ostrich, Penguin, Sparrow and so on, and I've got this function thunk() and I realise that, oddly it makes sense to thunk a Sparrow or an Ostrich, but literally no other birds at all. I think hard about it, but the best I can come up with to describe this property is "Thunkable" since all it really means is you can thunk() them. And there's no "content", there's no special implementation work in "Thunkable" that shouldn't live in thunk() for maintenance anyway. In this case ad hoc polymorphism is great because it saves needing to make this stupid "Thunkable" trait / type class / type-of-types / interface just to group together Sparrow and Ostrich for this single purpose.
But I'd argue the "Thunkable" case is rare, and C++ has a lot of cases where ad hoc polymorphism was the wrong choice and they fell into it.
I mentioned three types, really C++ contains() does only two things but it spells one of them two ways for 20+ year old reasons. It can do strings (as std::string, but also via C's char * type) and it can do a single character char (one code unit, so, no poop emoji).
Rust provides four Patterns, they are a string reference, a single char (a Unicode scalar, so yes a poop emoji works), a slice of chars (any of the chars matches), or a predicate which matches characters.
Now, do any of those four feel like things you'd definitely never want in C++? Because if C++ wanted all of them that's now five overloads for contains. And five overloads for find, and for every other matching function, on every string or string-like type...
I believe ad hoc polymorphism is so rarely what you really want, and yet it so often detracts from the rest of the language facilities that "We don't have that" is a sensible language design choice, same as for multiple inheritance.
Typeclasses (in a Haskell context) we’re formalized in the Wadler and Blott paper “How to make ad-hoc polymorphism less ad hoc”. And for Rust, in [1] Traits are explicitly stated to be the method in which Rust achieves ad-hoc polymorphism.
1: Klabnik, Steve; Nichols, Carol (2019-08-12). "Chapter 10: Generic Types, Traits, and Lifetimes". The Rust Programming Language (Covers Rust 2018)
Edited to add: Hmm. Actually though, surely that second reference is just "The Book" as it's called, did it actually say it's about ad hoc polymorphism? Because I've read this section of The Book, although I hadn't in 2018, and it doesn't mention "Ad hoc polymorphism".
What's there now (I just checked) is a description of how you'd approach this problem in Rust, using traits, but doesn't claim this is ad hoc polymorphism, and sure enough doesn't involve an arbitrary set of types which is the sort of the point of why "ad hoc" is there in the name.
Strachey chose the adjectives ad-hoc and parametric to distinguish two varieties of polymorphism [Str67]. Ad-hoc polymorphism occurs when a function is defined over several different types, acting in a dif- ferent way for each type. A typical example is overloaded multiplication: the same symbol may be used to denote multiplication of integers (as in 33) and multiplication of floating point values (as in 3.143.14).
That is from the Wadler paper where typeclasses are formalized. Typeclasses and Traits are the implementation details for those function symbols that vary in implementation for each type. The restrictions on the types (like which types implement a trait or have a type class defined for it) are the types the symbol can be used on.
You seem to be focusing on the ‘arbitrary set of types’ point, but the only connection between the types accepted by Rust’s generic functions (which are functions that accept a type provided it has some trait) are that they take types which have an impl for that trait.
I think there is a bit of ambiguity regarding the term ad-hoc polymorphism at play. You seem to think the trait/typeclass implementation of ad-hoc polymorphism (which was invented to formalize a well behaved class of ad-hoc polymorphic functions) makes it no longer ad-hoc. My position echos Wadler, it’s still ad-hoc but just less ‘ad-hoc’ (i.e. more formalized).
In C++ there are literally three separate implementations of std::string's contains method for three type signatures. This is pretty clearly what is being discussed as "acting in a different way".
In Rust there's just one, here's the entire function body of contains: pat.is_contained_in(self)
OK, well that's just buck passing right? Clearly this is_contained_in() method on Pattern is really just the contains() implementation, we're passing the work to this function that as you point out needs to be implemented by each of the matching types for Pattern.
Except, wait, Pattern actually defines is_contained_in(haystack), thus: self.into_searcher(haystack).next_match().is_some()
Sure enough Pattern implementations although they're not forbidden from implementing is_contained_in themselves, do not in fact do that, they just implement into_searcher. Our hypothetical Jpeg type can provide a suitable into_searcher implementation which results in a Searcher for the Jpeg somehow, without knowing what contains() or split_once() or trim_start_matches() do, and now they will work on Jpegs.
So the "acting in a different way for each type" for contains() ends up only being because of details about the inner behaviour of that type, which is exactly parametric polymorphism so far as I can see.
- What is the benefit of using A::B::C versus A_B_C, really, in terms of avoiding name clashes? I can point out one immediate benefit of doing the latter - code is easier to read because line noise is reduced (perhaps subjective but I think it will be hard arguing the other way around).
- Is it a good thing that half the code will now refer to A::B::C using only B::C or even only C? Without assistance of a program with good semantic insight (a solid IDE, etc) this only makes identifier search harder.
- Namespaces "enforce" discipline in one way - you can be sure that the symbols at the binary level will be properly prefixed. But with only a little discipline programmers can do the prefixing themselves, and in return they can move code between files (/namespaces) more freely, which is good for refactoring.
And as a sibling commenter pointed out, one factor why namespaces get called for could be lack of understanding that we can hide the "guts" of any function (or variable) using "static" linkage. There is less discipline (prefixes..) needed for these functions because they are only visible in the current translation unit.
I suppose aliasing is useful for widely used libraries, for example if a big software projects wants to include two different versions of the same library. Or Team A wants to rename their module but make it easier for other teams that use their module to follow suit.
All that can provide some ease of use in the short term but produces more mess to clean up in the long term. IMHO.
Real name clashes between two different libraries are quite unlikely, and namespaces would only solve that problem on the source code level, not on the binary / symbol level.
And that brings you what benefit exactly? I've seen a number of projects that are preoccupied about "proper nesting", while it is 99% bureaucratics and all those projects are still a mess.
The benefits of "living under this or that" are technically zero, and with respect to human factors are minimal given that 1) you can get most of the organizational benefits of namespaces by doing A_B_C, and 2) you can also add new members to proper namespaces from external files in most languages, including C++, so there really isn't a difference w.r.t "knowing for sure".
You could say that the C/C++ systems already has "modules" if you look at object files and header files. Of course, they are a limited kind of modules because they are a bit low-level and C has the preprocessor problem, making it slow to "import" modules. But none of that has to do with namespaces.
I think not having namespaces may be an effective way to prevent enterprise programming from happening.
Node.js and Rust do module systems just fine.
Many C programs do not use hierarchical names like `A_B_C` as they ideally should. They instead pick a short identifier that is (often incorrectly) believed to be distinct enough.
> Is it a good thing that half the code will now refer to A::B::C using only B::C or even only C? Without assistance of a program with good semantic insight (a solid IDE, etc) this only makes identifier search harder.
If you meant that the identifier search should be possible with just grep, yeah namespace makes it harder but it's not the only cause and this doesn't explain why many other languages with less overall IDE support than C/C++ have namespaces. And I believe it is possible to design a namespace system that only needs a ctags-level automation for proper identifier search, though I don't know how you feel about ctags.
What is your list of name clashes that you experienced? How many real headaches did they give you? Or is it all a non-problem? Not a rhetorical question - I'm mostly working on smaller projects < 100KLOC.
> I don't know how you feel about ctags.
I use ctags from time to time when I have to, but I still don't like it when there are multiple namespaces that contain the same set of names which confuses navigation operations, even vim with ctags I think. Maybe even IDEs like Visual Studio, I'd have to check though what works and what doesn't.
I too refrained from using C for that large software so I don't have many examples either, but in one case I was using TweetNaCl where you have to supply `extern void randombytes(unsigned char*, unsigned long long)` for CSPRNG and I had to rename it for some reason I can no longer recall.
However this is not a reason to add namespaces. (In fact the bug was fixed using symbol versioning, an already existing feature of ELF.)
Here are a dozen examples from Xlib, where Windows happened to choose the same short common names for many things: https://gitlab.freedesktop.org/xorg/proto/xorgproto/-/blob/m...
... and IIRC this list is no longer sufficient; you need to add to it in order to compile with the current Windows SDK. (Upon closer inspection, it was updated two weeks ago, so maybe it's fine... until the next Windows SDK update.)
I'll admit that this happens less often on Linux, where the system headers are smaller (and everybody uses X11 so there's historic reason to avoid all the short names that were gobbled up by Xlib in the '80s), but I've still run into occasional clashes between different library headers, or between legacy code and updated headers. eg. bool/Bool/BOOL are common collisions among pre-C99 libraries (including libraries that require C99 but don't remove the old names for backwards compatibility reasons), as well as min/MIN/max/MAX which still aren't in standard C as far as I can tell.
The headaches it gives are real, but not large in the grand scheme of things. The lack of defer (or otherwise standardized and cross-platform __attribute__(cleanup) ) is a bigger headache, for example.
My experience is that sharing libraries is so wildly difficult In both C and C++ that code is not shared and wheels are reinvented. This has more to do with build systems than namespaces. But namespaces are a factor.
Because they are not C-style language. In Python or Java, for example, you need a separate file to create a package. In C++ you can add namespaces anywhere you want. This makes C++ namespaces harder to maintain even with automated tools.
I was asked during the meeting why I think that name spaces are a bad idea, so here is my thoughts. I must warn you that the following may sound like an anti C++ rant, because well .... it is.
Name spaces have several bad aspects. The majority of them are that they make the code a lot less readable.
First, If I see a function call in source code, I cant assume what function is called. There may be a namespace somewhere above or in a header-file somewhere that changes the meaning of the code. It matters a lot for instance when you copy code from one file to another, and all of a sudden the code just means something completely different. Name spaces add state that changes the meaning of the code around it. When your code isn't doing what it should and you are trying to figure out why, you need to be able to trust that the code you are reading on screen is actually what you think it is. I dislike it for the same reason i dislike overloading: it makes it less clear what things you are reading really are.
(It may be argued that you already cant trust your eyes in C, because you can do lots of devious things with the per-processor, and to that I say, yes its bad enough, so lets not add more of it)
Name spaces splits the identifier in half, but its still a unique identifier.
my_module::create();
is and has to be as unique as:
my_module_create();
So what are we accomplishing? Do we have the same level of collisions? No, in fact we have more collisions because the parts can collide individually! if there are multiple my_module they will collide, and if there are multiple create in different namespaces they will collide!
my_module::create(); // one name space my_module::destroy(); // another name space
Collision! how about:
using namespace my_module; // namespace with a create using namespace my_other_module; // namespace with a create
crate();
Collision! If we instead write plain old C:
my_module_create(); my_module_destroy();
No collision.
my_module_create(); my_other_module_create();
No collision.
Beyond creating new opportunities for collisions, we create confusion by making it possible to hide half the identifier somewhere else. Yes, it can save some typing, but programing is never hard because of typing, programing is hard because its hard to understand what is going on. Namespaces solves the easy stuff, by making the hard stuff harder.
Its the age old, trap of C++ of adding something clever out of convenience, that turns out to be unclear and something that the programmer has to manage.
Further, namespaces encourage people to use short common names even further, because they think it saves them typing, and they think if it goes wrong they can always manage it with namespaces, and then they end-up having to manage something they shouldn't have to manage in the first place.
I have never had a namespace collision in 20+ years of C programming. Why? Because I use long descriptive names, that always start with where the functionality resides. Its readable, straight forward and works. Namespaces is a complicated system for managing a problem that should never happen in the first place, unless the user is very careless. We should not encourage carelessness.
I think C, should take some blame for this being a problem because the standard library have far shorter names then is advisable. It has made people think that names like: "jn" or "erf" are good examples of unique, clear and descriptive naming. The added wrinkle of "significant characters" has made people think that C mandates short names, something that no major implementation requires. There also seems to be a persistent belief that display technology has not yet evolved to the level that we can display more then 80 characters per line and that we therefor need to use short cryptic names for everything. This is also an argument from a different age.
Namespace collision in C almost never happens between 3rd party libraries, it is almost the exclusive problem of the standard library because it is so poorly named. If we want to fix this, by hiding the standard library behind a namespace, we might as well just add a c_standard_lib_ prefix to all functions, and keep the garbage that is namespaces out of C. Not that we would do either, since both world break backwards compatibility. So why would we add namespaces if it wouldn't be used by the standard lib, the very library we have name collision problems with? In fact if we added an optional namespace for the standard lib, all we would accomplish is to pollute the namespace with one more identifier.
By now if a "module prefix" is bigger than say 4-5 characters, I get unhappy and know I need to improve, to find a short mnemnonic. Same goes for local variables names, which I try keep at 1-5, rarely they get 10+ characters long. There is always this tension between names being "self-documenting" and "just long enough to remind of the purpose that was explained in a not too distant context". Variables that are frequently used should be shorter. Variables that have a clear intuitive use (like "x" or "i" which most often have very clear meaning with only little context added) should be shorter. Module names should _always_ be very short because (I assume) there are only few modules, and it's better to remember the purposes of a few modules together with their abbreviated names, than to have to read a long repetitive module name each other line. Function name suffixes (without the module prefix) should often be long because there are many different functions and not all of them can be cached with their meanings in the programmer brain - so I allow function names to be a little self-documenting typically.
I agree that the names in POSIX and C are too short and cryptic but think they made more sense in the context of Unixes of the 1970s when projects weren't as big as today.
I still try to stay within 80 or 100 columns because line length a readability / eye strain concern as well, but given that I'm in the 8 spaces camp I don't freak out anymore if there is the occasional 140 characters line and I'm to lazy to trim it down. Judicious insertion of linebreaks is most often useless busywork. Then again, some function signatures take too many arguments (glVertexAttribPointer/glBufferData... or the Win32 API come to mind) and inserting linebreaks in calls can improve readability sometimes.
I have that problem, repetitive prefixes getting in the way of reading the code, and I’ve been doing some experiments with what could be described as “local namespacing”, only inside an expression or statement, with the prefix/suffix feature of the backstitch macro in my C preprocessor:
https://sentido-labs.com/en/library/cedro/202106171400/#back...
For instance, writing a program with libxml2, I write:
Next(reader) @xmlTextReader...;
which gets translated to xmlTextReaderNext(reader);
and I find the first easier to read.Whether that’s the case for others, I don’t know.
A longer example from the link I wrote above, this time for libuv:
@uv_thread_...
t hare_id,
t tortoise_id,
create(&hare_id, hare, &tracklen),
create(&tortoise_id, tortoise, &tracklen),
join(&hare_id),
join(&tortoise_id);
Result: uv_thread_t hare_id;
uv_thread_t tortoise_id;
uv_thread_create(&hare_id, hare, &tracklen);
uv_thread_create(&tortoise_id, tortoise, &tracklen);
uv_thread_join(&hare_id);
uv_thread_join(&tortoise_id);In C++, it is intentionally not straightforward to know what code gets called when a statement is executed - even without namespaces. We have:
* Function overloads
* Non-trivial (and not-built-in) implicit conversions
* Non-trivial copy and move constructors
* Template and concept resolution, which can be tricky
* Operator overloads, so that foo[n] or bar() may do something very different than what you expect.
And if you add in macros to the mix, you can really go crazy. For example, in this SO answer: https://stackoverflow.com/a/70328397/1593077 I explain how to implement an infix operator which lets us write:
int x = whatever();
if ( x is_one_of {1, 3, 42, 7, 69, 550123} ) {
do_stuff(x);
}
C is simply not that kind of language (well, macros notwithstanding I suppose).With your example, both a switch and simple if-else chains would be perfectly readable and also easy to write.
using namespace my_better_name = my_module;
//....
my_better_name::create(); // colision gone
Problem solved, better improve your C++ knowledge.The C standard library should have (maybe as early as C11 or C99) picked stdc_ as a "namespace", reserved it, made it a warning to put things in it, and used it for everything going forwards.
What C needs is not namespaces, it's a module system, so that symbol visibility can be more tightly controlled.
I agree that optimizing number of keystrokes is a bad goal, but in my opinion that isn't the selling point of namespaces. The benefit is better readability. Code written with namespaces lets your eye focus the semantic name of the method, not extra character noise that was added as a form of manual name-mangling.
This is highly opinionated, of course, because there is a tradeoff with the ability to uniquely identify a method at a glance, as you mentioned. Which camp you fall into is likely to depend on what sort of software you write and whether you use an IDE or a text editor.
Annex K tanking badly was just sad. At least they finally standardized strdup().
Some of the string functions, like strtok() or strcat(), and I would also include strdup(), are broken beyond repair and should simply not be used. The only way to "fix" the library would be to remove these functions, but I assume this would create more trouble with legacy code than it's worth.
C strings are mostly a useful storage format for small strings, including string literals, and as a "default" string type for the little basic functionality that is expected to be part of C to support a small program growing up - including printf() and friends.
For longer strings, you are expected to build datatypes that fit your usecase, and that includes decisions about memory management. It should not be the responsibility of the language to make these decisions for you.
It's something that I use. Maybe not frequently, but I definitely use it.
Also, it is in no way equivalent to malloc(strlen(x)). You're off by one byte and you've not copied any of the string.
At least, glob() has its own globfree() function, leaving room for slightly more mindful organization.
Perhaps they're just to give you a better understanding of the more minor details/intricacies of the language, but I couldn't help but think with a lot of their code, if I saw a coworker write something like that I'd roll my eyes that they didn't do it in a more obvious way.
So, it's not crazy to argue that because idiomatic C isn't acceptable therefore rule out C as a language for an organisation to use entirely. Clearly if you don't have an alternative this highlights a problem rather than itself being a solution. But it might be a reason not to do whatever it was you were thinking of doing at all. Do you really need to write micro-controller firmware for your project? Isn't there an off-the-shelf component where it's not your problem to maintain the horrible C code that makes it work?
[ Or you could go all Oxide Computer Company and decide no, C is just terrible so we are going to use Rust and if we need to rewrite the firmware on this switch that's what we'll do. That might make total sense, if you go into it clear eyed ]
> struct T_slice { T* ptr; size_t len; };
with all the syntactic sugar and bounds checks added in. This could catch so many out of bounds issues and doesn't seem that hard to do.
int arr[..];
syntax for such things. But, our plates are full and the proposal wasn't written: it'll be up for next release, for sure, and it's one of the many things I would like to put in the language.
I always thought it was silly that the best way to embed a resource in a binary using CMake was to either convert it to hexadecimal C array manually or use specific flags with ld to convert it into an object (and good luck getting that to work reliably with cross-compiling!).
Unfortunate characterisation of constexpr in C++ over revisions as feature creep. It's more that any C++ should be executable at compile time with constexpr ultimately going the way of the register keyword, but selling people on things like compile file I/O is a long game.
Ultimately though it's hard to be too excited. I don't have a use case for C any more, beyond keeping some old code alive (with fstrict-aliasing et al and a sense that I really should port it to something else before I can no longer get any compiler to build a program out of it).
If you want a better C, do yourself a favor and use D
Same syntax, more secure and with modules
Every language has downsides/upsides and time. There is a nice quote by Titus Winters: “Engineering is programming integrated over time”[0], this helped me a lot in learning that context matter for tools/languages/frameworks, and you can learn with their mistakes.
I'm still waiting for tagged unions, tuple and pattern matching
auto and #embed is nice to have though
You'll have a hard time pushing for those: C will likely never get them in those forms. At least, not for another 10-20 years, since there's no existing C compiler that implements any of those things yet.