Considering C99 for Curl
daniel.haxx.se
daniel.haxx.se
A much more interesting question is if you were writing curl today, what would you do. If the answer is still 'C89' then we as a profession have to wonder why - did we get it exactly right, and there are no lessons from the last 30 years, or is the fact there are no better alternatives deeply depressing.
C89 is still a reasonable choice today. I don't think it's depressing. It's a good language, and hitting the sweet spot of both language design and implementation is really hard, so you'd expect to have very few options.
Personal opinions of Rust aside, it has gained so much fanfare that “rewrite it Rust” is now practically a meme.
edit: It is also worth exploring why and how a systems programming language has generated this much excitement in folks that are "rewriting it in Rust". These people are also in their way making it that much easier for everyone else to transition to Rust, proving these projects work just fine in Rust. I do agree it is a meme, but it is a good one for us all. As an infosec practitioner nothing could make me happier than seeing people excited about a language that eradicates one of the worst and most pernicious classes of C/C++ bugs.
https://www.schneier.com/blog/archives/2007/09/the_multics_o...
> The combination of BASED and REFER leaves the compiler to do the error prone pointer arithmetic while having the same innate efficiency as the clumsy equivalent in C. Add to this that PL/1 (like most contemporary languages) included bounds checking and the result is significantly superior to C.
One data point more backing my theory that we're in the middle(?) of "computing dark ages", where the biggest crap and nonsense dominates everything (and people not even know how crappy everything is).
Do you think we will ever leave the dark age?
I mean before our AI overlords get in charge and kill and replace all the nonsense we've built, of course.
This is a matter of quality, and like everything in computing, quality only matters when money or law is involved, so in a way returning digital goods with refund is one way to make companies take quality more seriously, other are stricter liability laws for when exploits ocurr in the wild.
There are languages out there with much more advanced features than Rust, like e.g. Scala. But the IDE experience in Scala is not worse than with Java.
There is no fundamental problem with IDE support. It just takes some work with an advanced language.
(The only issue are languages which you need to write "backwards", like Haskell. But that's another story.)
I recently started working with Rust for contributing to projects like Rome/tools [1] and deno_lint [2]. My first impression with Rust is a bit frustrating: compilation takes tens of seconds or minutes. I am waiting in front of my IDE to get type hints / go-to def (often to fall-back to a text search). When I am launching unit tests, I am waiting that rust-analyzer terminates its indexing, and then I am waiting again that the tests compile…
The tools are now mature, and a lot of engineering work is done on both rust compiler and rust-analyzer. I am afraid that the slow compilation of Rust is rooted to its inherent complexity.
AFAIK that's not the case.
The problem here is that Rust has "issues" with separate compilation due to some language design decisions (which affect incremental compilation than obviously). It's not built for that and it is, and continuously will be, quite difficult to make this work somehow.
But that's less a problem with the complexity of the language as such.
My Scala example stands: Scala is also quite complex and not the fastest to compile. But after the build system and the compiler crunched the sources once (which may take many minutes on a larger code base) the IDE is very responsive. Things like type hints or go-to def are more or less instant. Code completion is fast enough to be used fluently. Edit-compile-test cycles are fast thanks to the fact that separate compilation considerations were part of the language design decisions. (That's for example why Scala has orphan type-class instances; which are a feature and a wart at the same time).
As I understand Rust's "compilation units" are actually crates. This is not very fine granular and I guess the source of the issues.
I would guess splitting code into a few (more) crates (which than need to depend on each other) may improve the incremental build times. Also things like not building optimized code during development of course apply, but I think cargo does this automatically anyway.
But I'm not an expert on this. Would need to look things up myself.
Maybe someone else has some proven tricks to share?
…
OK, a quick search yielded some useful results, so I share:
https://www.pingcap.com/blog/rust-huge-compilation-units/
https://fasterthanli.me/articles/why-is-my-rust-build-so-slo...
In fact, I included the lack of compilation locality in "inherent complexity of Rust". However, I agree that this could be considered apart.
In my experience with TypeScript (quite different, I admitted), splitting in distinct compilation unit may help. However, this does not solve the issue.
This could be great if Rust could deprecate some features in order to improve its compilation speed. I am not sure if it is feasible…
But... I'll admit there is one major stumbling block so far. Debugging iterator chains can be cumbersome because of the disconnect between the language and the compiled code. I've found myself stepping in and out of assembly more than I'd like. I assume this is the kind of problem that can be overcome with a nicer debugger though.
I think this would be something that modern debuggers need to solve somehow in general.
There are more and more languages with high amount of syntax sugar, where the output to be debugged doesn't have much in common anymore with the code written.
Debuggers need to be aware of desugarings somehow.
But it makes no sense to implement this on a case by case basis for every language. We need next generation debuggers! (But I have no clue how "a sugar aware debugger" could be implemented; something in the direction of "source maps" maybe?)
Particularly the MSVC ecosystem is identified in TFA as being a late adopter.
Once built c99 code can run where it wants.
Is that true on the MSVC ecosystem?
Don't you have to compile separately for each `msvcrt` environment, as I thought they aren't binary compatible? And would a non-C99 msvcrt necessarily have an `snprintf()` implementation in its libc-equivalent dll?
This is possible because, unlike on Linux where function names are global, on Windows function names are scoped to the DLL, so you can have MSVCR71.DLL and MSVCR81.DLL loaded at the same time in the same process and they won't interfere with each other.
I don't see how that fixes the problem of possibly not having an snprintf() implementation on a system that doesn't have a C99-compatible MSVC runtime environment.
Did I miss an implication of your comment somehow?
All supported MSVCRTs are installable on all supported Windows SKUs.
...and by "forget", I mean "block out due to trauma, because surely it can't be that stupid".
Windows 7 and 8 haven't been supported in years, but are still pretty common in the wild.
Windows 7 hasn’t been supported for two years now, but Windows 8 EOS isn’t until January 2023.
This is also possible on Linux with linker scripts AFAIK.
What you can use, to a limited extent, is dlmopen().
Shared objects in Linux are just really late linked static objects, with fix-ups (hence PIC requirements).
In macOS and Windows, there are heirarchies.
dlopen permits you to control whether exported (externally visible) functions in a module become available to satisfy link dependencies in the application, such as subsequent module loads. See the dlopen flags RTLD_GLOBAL and RTLD_LOCAL.
dlmopen is for controlling the visibility of shared library dependencies pulled in by dlopen'd modules, whether RTLD_GLOBAL or RTLD_LOCAL, which only effect the immediate symbols in the module and not symbols from automatically loaded shared library dependencies. If you link the main application with OpenSSL (-lssl -lcrypto), or a prior module you dlopen'd pulled in OpenSSL as a dependency, then those OpenSSL symbols become available to satisfy requirements for subsequent dlopen'd modules. dlmopen allows you to create an entirely different symbol namespace for a module or modules, where symbols dependencies are only ever satisfied from that namespace, and exported (global) symbols, whether pulled in by dlopen or transitively via a shared library, are never visible outside that namespace.
None of these options directly map to the behavior of DLLs. DLLs fundamentally use different semantics, AFAIU. The closest behavior to DLLs might be DT_RUNPATH + dlmopen, but dlmopen use is explicit so not really the same thing. You could use ELF symbol versioning (maybe in combination with DT_SONAME and DT_RUNPATH) to accomplish the same thing as DLLs by effectively renaming all the symbols in a library (e.g. attaching a version component), but there aren't any tools around to help automate that, AFAIK; you'd have to generate linker scripts and it'd be a complex build. Much easier to just static link at that point.
What systems don't support a version gcc that compiles c99 at this point?
Some hacks are quite crazy.
E.G: there is this very old navy broacasting protocol, NMEA (https://en.wikipedia.org/wiki/NMEA_0183), that was designed so it could be transmitted through old fashion radio waves. You'll find it in some sonars, water sensors or AIS beacons. For this reason, despite that it looks more like a layer 4 protocol, it embeds its own packet format and checksum, all in ASCII, that clients are expected to parse.
Now of course, a lot of devices are still emitting their data in NMEA, and it's not uncommon to be able to just telnet or netcat (if UDP) into one to see the data flowing.
But after a while, people started to aggregate those data from their numerous sources into one single router, and expose this router for convenience through... HTTP over TCP/IP.
And now you have those all those old computer towers (some still rocking a CRT screens or windows xp) doing long polling to get broadcasting data over a protocol that was made for request/reponse, to read the payload that is another protocol that was meant for radio equipment and hence requires manual consistency checks, that is transported by yet another protocol that is doing its best to preserve packets.
And they say the spirit of hacking is dead :)
(sometimes I feel IoT or domotic stacks look the same honestly)
Of course, somewhere in there, there is a curl call. The question is therefore does curl author want to support a potential upgrade path for such twisted use case or not. I would say "nahhhh", but maybe curl had precisely the success it did because the author was ready to support it in crazy settings.
I think nobody really does any serous development on such old machines.
Again, not sure that it means curl should endorse such niche situations, but the modes of failure are numerous.
TL;DR: the world is complicated
Whether this is feasible in a economic sense is another question. But technical it should be possible.
Why isn't it possible that the updates aren't as good as the people who wrote them thought they were?
VLAs aren't a problem. VLA in C are.
But that's not because there is any issue with VLAs. The issue is that C concepts are stupid, but nobody fixes the roots of the issues.
If you put something of variable length (a VLA) into something with limited static length (the "stack") it will explode. That's nothing new and nothing special or exclusive to VLAs. Actually, exactly this is one of the main issues with the bad C design since inception: It does not do bound checks (especially no static ones; as this would require proper depended typing for safe VLAs). Out of bound access will just "explode" as always in C (likely leaving a nice security crater).
To be honest I don't get why we're still stuck with the stack / heap nonsense. There is not stack (or heap). There is only memory.
What would be much more interesting would be direct control over the caches… Instead we still use the pure fantasy products "stack" and "heap" which are actually irrelevant (as they don't exist in the end).
There is nothing "special" about "stack" memory. That's just a very primitive region based automatic memory allocator backed into the C runtime!
The whole "using registers" thingy in context of "stack" is also just fake by now. You don't use HW registers—but some virtualization of them presented to you by the VM that runs inside the CPU. So you don't control register allocation anyway! So this could be made completely transparent without any impact. (The VM inside the CPU does the actual register allocation fully automatic. Presenting the "faked virtual ISA registers" to the outside world just to make "legacy" code happy).
Do you mean the removal of VLAs from Linux? Why do you think that was misguided?
That's not what the post said. He didn't say that C99 offered no advantages to anyone. He said no one could come up with benefits to the curl project that would be gained by moving to C99, therefore the risk introduced by doing so was not worth it for now (my paraphrase, obviously).
It sounds to me like a perfectly good reason to stay with the current standard for that project.
[edited to fix a typo]
It's still deeply depressing how costly/impractical it is to apply improved technologies in the long tail of environments that isn't "linux on amd64" and similar, but it's not really a language design question in my opinion. We didn't get it "exactly right", we got it "good enough", and upgrade costs are prohibitively high for the general case.
My opinion is C is already way too rich and complex. I would stick to c89 with benign bits of c99 and c11. The benchmark being "one average system developer coding a naive and real life C compiler in a reasonable amount of time and effort".
That said, I know that my "next" C compiler will probably be a RISC-V assembler with a very conservative usage of a macro preprocessor.
If Curl already has its own "decent and functional replacement" for `snprintf()` that's used extensively throughout the codebase, or if they just don't need that functionality (I haven't checked) then I guess that's not an issue. But that would be the big selling point as far as I'm concerned.
Like 1980s Internet protocol features the rationale for weird things in C is more often "That's how Unix works" than "This is actually a clever safety feature".
... AND where you want to pad the remaining space with zero bytes, so that you don't leak uninitialized memory onto the disk, or network.
The null byte padding behavior of strncpy makes it clear what the intended use was.
Also, the way C initializes character arrays from literals has strncpy-like behavior, because the entire aggregate is initialized, so the extra bytes are all zero:
char a[4] = "a"; // like strncpy(a, "a", 4);
char b[4] = "abcd"; // like strncpy(a, "abcd", 4);
the compiler could literally emit strncpy calls to do these initializations, so we might say that strncpy is a primitive that is directly relevant for run-time support for a C declaration feature.If you look at this function assuming it's a safety feature, that's a huge surprise, and indeed if you were skimming you might miss what it does because (in the context of "it's a safety feature") this is an insane choice. "Why would you do that?". Well, because it's not a safety feature.
The perf cost isn't what you'd expect from a safety feature either. Suppose we have a 1MB buffer, and we strncpy "THIS" into it using n = 1024. That's just four bytes right? Nope. strncpy() will write "THIS" and then 1020 zero bytes.
Not only that, but (because of its actual purpose) it also fills the buffer with nuls, which is a complete waste of resources.
So yes, the original intent of strncpy() germane to the GPs comment, because it makes strncpy actively dangerous and complete shit when working with C strings.
Calling a bespoke byte-sequence data structure a "string" is inaccurate. Treating strncpy() as a string function is erroneous and can easily lead to memory corruption.
It was completely clear, and can easily be inferred from its specified behaviour.
It's just completely useless nowadays, because its purpose is essentially obsolete, because the data type it works with is almost never used anymore.
strncpy works with fixed-size nul-padded fields as you'd find in e.g. mainframe-type software. That is why it:
- fills the destination buffer with NULs if the source is shorter
- does not nul-terminate if the source is the same size or longer than the destination
strncpy is essentially equivalent to zero-ing a buffer of size `n` then copying the first `n` bytes of src (up to the first nul) in the target
(I don't use strncpy and don't defend its functionality, just want to know what was intended when it was introduced.)
[1] "UNIX Implementation" by Ken Thompson, _The Bell System Technical Journal_, July-August 1978, Vol 57, No 6, Part 2, pg 1942.
This is mostly relevant to write into existing memory or memory-mapped records.
And on embedded devices you are often memory constrained so you might reuse an existing structure.
Some typo on the buffer limits and the same hazards as always.
Only fix is hardware memory tagging.
What I meant by optimization is not treating them same as other structs, but e.g. guaranteeing pass-by-register like other primitive types, spelled out explicitly in the ABI. The choice of length field type would be size_t, obviously.
Systems which do/did include Burroughs Large Systems (now Unisys ClearPath MCP), IBM System/38 and AS/400 and IBM i (the RISC versions of which used PowerPC AS Tagged Memory Extensions), ARM MTE, SPARC ADI, and CHERI/ARM Morello.
no-nul is not the goal of strncpy, it's the effect of strncpy.
strncpy is designed to work on fixed-size, nul-padded strings. That's why it fills the destination buffer with nuls if the source is too short, and it doesn't guarantee nul-termination (if the source is exactly the size of the destination).
snprintf(dest, sizeof dest, source); // BAD code do not repeat
that looks great at first sight, just another size-checked way of copying strings, but remember that the third argument to `snprintf()` [1] is of course a `printf()`-style formatting string. So if that `source` argument contains any percent symbols, there's gonna be a party in your computer and both Undefined and Behavior are going to show up. You don't want that.Instead, if you want to use `snprintf()` for this, remember to do:
snprintf(dest, sizeof dest, "%s", source);
[1]: https://linux.die.net/man/3/snprintfI mean, it's a trivially replicated function, you can just copy it over from an existing codebase. So it's not exactly a major feature.
char *stecpy(char *d, const char *s, const char *e)
{
if (e) e--;
while (d < e && *s)
*d++ = *s++;
if (d)
*d = '\0';
return d;
}
main() {
char buf[64];
char *ptr, *end = buf+sizeof(buf) ;
ptr = stecpy(buf, "hello", end);
ptr = stecpy(ptr, " world", end);
}
As discussed here https://twitter.com/hyc_symas/status/1382298601641152513The point is that the end of the buffer is invariant, there's no reason to screw around recalculating the length of remaining space after each copy into the buffer. This also fixes the nonsense of strcpy/strcat returning the same dst pointer that was passed in. By returning the pointer to where copying ended, you don't need a separate strcat function any more, nor do you have the Shlemiel The Painter problem with strcat.
Functions can always be implemented with additional header files.
They aren't actual extensions to the language.
> Like improving the test suite, increasing test coverage, making sure more code is exercised by the fuzzers.
I wonder if using a more recent version of C might draw in more developers willing to contribute (in a similar vein to Linux kernel introducing Rust).
struct point p = { getx(obj), gety(obj) }; /* C90 error, C99 OK */
GCC had this as an extension before C99. Coding around that one can get ugly, and it's just a syntactic limitation.Also, the C99 preprocessor is more powerful, with variadic support.
The macros I implemented in the cppawk project (awk with C preprocessor) would be impossible without C99. I was able to make a multi-clause loop macro, with user-definable clauses.
But in retrospect I'm surprised to see how many features we used liberally would have been unavailable in pure C89. snprintf first and foremost. But also // comments, __func__, and stdint.h
void foo(size_t n) { int a[n][n]; }
does not have to be supported outside C99 but void bar(size_t n, void *p) { int (*pa)[n][n] = p; }
has to be in both C99 and C23 (though not in C11 or C17).Is there risk associated with changing? Probably yes in terms of security and limiting compatibility.
It probably doesn't make sense to change now unless there are specific reason(s) that will lead to impactful improvements.
> ... risk opening the flood gates for people rewriting things...
This sounds like "we feel that changes are needed but we will not be able control it".
Only if a C99 compiler is available. As the post points out:
> The slowest of the “big compilers” to adopt C99 was the Microsoft Visual C++ compiler, which did not adopt it properly until 2015 and added more compliance in 2019. A large number of our users/developers are still stuck on older MSVC versions so not even all users of this compiler suite can build C99 programs even today, in late 2022.
(emphasis mine)
[1] Also because the C++ committee ignored Herb and just forced it into later C++ revs from a library perspective.
Naturally without the stuff that got made optional in C11, no need to spend development effort on legacy features.
> [...] we would have to go gently and open up for allowing new C99 features slowly
> A challenge with that approach, is that it is hard to verify which features that are allowed vs used as existing tooling normally don’t have that resolution.
> The question has also been asked that if we consider bumping the requirement, should we then not bump it to C11 at once instead of staying at C99?
The motivation ultimately boils down to:
> Ultimately, not a single person has yet been able to clearly articulate what benefits such a C flavor requirement bump would provide for the curl project
The syntax isn't as nice, and you have to remember to do it (maybe a linter can help with that) but it ends up more of a "nice to have".
The boolean primitive data type is a standard in modern languages. Take any language: Rust, Go, Zig... and C99 (kind of).
In C89, a project may use enums. Another macros. Another ints.
Isn't cURL susceptible to this too? (Not to mention i always find that style much nicer :)
What features of C99 isn't this true of today?
IIRC, MSVC stubbornly refuses to add support for variable-length arrays (which AFAIK are required for full C99 support). I don't know if there's anything else on that list of C99 features that MSVC doesn't support yet.
You are correct that VLAs are mandatory for C99, but they turned out to be such a bad idea that they are optional in C11 onwards.
> I don't know if there's anything else on that list of C99 features that MSVC doesn't support yet.
IIRC, MSVC's `snprintf` and friends are broken (returns incorrect values and/or interprets the size parameter incorrectly). I think all of the annexure K stuff is broken in MSVC.
Doesn’t annex k only exist in MSVC because it’s a bunch of crap MS got the committee to add and no one else wanted to implement?
And isn’t it a C11 thing?
Yes, but in a perverse twist of fate, Microsoft's implementation does not conform.
> And isn’t it a C11 thing?
I stand corrected, it is a C11 thing.
Google has paid the development effort for the Linux kernel to get rid of all VLA occurrences.
> Google has paid the development effort for the Linux kernel to get rid of all VLA occurrences.
IIRC, what they were really interested in getting rid of was not the normal VLA we're talking about, but something even more esoteric called VLAIS (variable length arrays in structures), which clang didn't support. See https://lwn.net/Articles/441018/ for a discussion about that.
I dunno how to probe the stack in standard C, and so I never used VLAs, nor allowed them because (I thought that) there was no way to determine if a VLA declaration would cause a stack overflow.
They were made optional in C11. It was a good decision because it's a dangerous misfeature.
That's how it's implemented in all the major compilers. Anyway, even a hypothetical heap-based implementation would be bad because there's no way to report allocation errors.
// VLA
int foo(int n) { int array[n]; }
// flexible array member
struct s { int n; double d[]; }; struct s s1 = malloc(sizeof (struct s) + (sizeof (double) 8));
I don't need a lecture. I have actually lectured on C before.
I do have them bucketed similarly in my mind. They are both c99 features about variable sized arrays. MS has not implemented either one iirc.
Flexible array members were implemented a long time ago. The oldest Visual C++ I have at hand is 2005 and it already has them (although only in C mode).
There is a Blog from Raymond on this[1]. The TL;DR: is that zero length arrays and FAMs weren't legal until C99. Not sure where the support from MSVC factors in. But that's the story and AFAIK he's sticking too it.
[1] https://devblogs.microsoft.com/oldnewthing/20040826-00/?p=38...
[0] https://learn.microsoft.com/en-us/cpp/preprocessor/preproces...
Cool.